This paper outlines the changing functions of language testing and assessment from the era of the White Australia policy to the present day. Key contributors to the academic and professional status of the field are identified along with influential tests and assessment frameworks emerging from particular policy contexts. Attention is drawn to the expanding scope of the discipline, which now encompasses diverse domains of language use and demonstrates a growing social consciousness. Australian accomplishments in language testing and assessment are now recognized both locally and internationally as central to the field of applied linguistics and as playing a critical contribution to fairness and equity in the broader educational and political context.
The importance of input from occupational experts in defining valid criteria to assess performance on English for specific purposes (ESP) tests is widely acknowledged. However, few studies have described the process of collecting indigenous criteria and establishing their suitability for a language testing context. The paper reports on this process with specific reference to the writing sub-test of the Occupational English Test (OET) for overseas-trained health professionals. The OET writing task requires candidates, in the role of health professional, to draw on a set of written prompts to write a letter of referral to a colleague on an aspect of patient care. With the aim of expanding the test construct as reflected in the criteria, health professionals from hospitals were asked to judge the adequacy of patient records obtained from health settings. Based on their feedback, an elaborate checklist was developed to reflect the qualities of the written documents that they considered critical. The paper discusses the challenges involved in the process, including the inevitable construct shrinkage evident in the final version of the checklist indicators due to the constraints of the testing situation. The study has implications for our understanding of authenticity in ESP assessments.
Researchers have recommended involving domain experts in the design of scoring rubrics of language for specific purpose tests by eliciting profession-relevant, indigenous criteria and applying these to test performances (see, e.g., Douglas, 2001; Jacoby, 1998; Pill, 2016). However, these indigenous criteria, derived as they are from people outside the assessment field, may be difficult to apply by the non-domain expert raters typically employed to rate performances on language tests. This paper addresses this question with reference to the writing component of the Occupational English Test (OET), a test designed to assess the English communication skills of overseas-trained health professionals. The paper describes the development of a set of professionally-relevant writing descriptors and then explores how well language-trained raters (N = 15) were able to apply these to a set of OET writing samples. All raters were interviewed and the rating data were analysed statistically. The findings show that while the statistical properties of the score data were generally satisfactory, some of the raters felt that they were not able to apply the scale confidently due to their perceived lack of medical knowledge. The study has implications for scale design, rater training and the use of professionally-relevant rating scales for LSP testing purposes.
Recent estimates indicate that French is spoken as a first or additional language by over 220 million people. French is an official language in 29 countries and in many organizations such as the United Nations and the Red Cross. French is also, after English, the most widely taught language in educational systems around the world, with an estimated 120 million students and 500,000 teachers. It is hardly surprising then that there is a strong international demand for official certification of French competence and that a range of tests are on offer to meet this goal. Among the recognized tests available for this purpose are the DELF (Diplôme d’études en langue française) and DALF (Diplôme approfondi de langue française). These are official qualifications awarded by the French Ministry of Education to certify the French competence of non-French citizens or of French citizens from non-francophone countries who have not completed a French secondary or higher education diploma. There are six independent diplomas: three for children or adolescents (DELF Prim, DELF Junior and DELF Scolaire) and three for adults (DELF tout public, a general proficiency qualification for those over 16 years of age, DELF Pro, a work-related test for those seeking initial employment opportunities or promotion, and DALF for higher level candidates). Each test is oriented to the CEFR scale with DELF Prim pitched at the preA1–A2 levels for immigrants with limited literacy backgrounds, the other DELF tests spanning the A1 to B2 levels and the DALF assessing proficiency at the more advanced C1 and C2 levels. Each test battery covers the four skill components of Listening, Speaking, Reading and Writing.
Models of communicative competence in a second language invoked in defining the construct of widely used tests of communicative language ability have drawn largely on the work of language specialists. The risk of exclusive reliance on language expertise to conceptualize, design and administer language tests is that test scores may carry meanings that are misaligned with the values of non-language specialists, that is, those without language expertise but perhaps with expert knowledge in the domain of concern. Neglect of the perspective of lay (i.e., non-linguistic) judges on language and communication is a serious validity concern, since they are the ultimate arbiters of what matters for effective communication in the relevant context of language use. The paper reports on three research studies exploring the validity of rating scales used to assess speaking performance on a number of high-stakes English-language tests developed for professional or general proficiency assessment purposes in Korea, Australia, China, and the UK. Drawing on Jacoby and McNamara's (1999) notion of "indigenous assessment", each project attempted to identify the values underlying non-language specialists' judgements of spoken communication as they rated test performance or participated in focus-group workshops where they viewed and commented on video- or audio-recorded samples of performance in the relevant real-world domain. The findings of these studies raise the question of whether language can or should be assessed as object independently of the content which it conveys or without regard for the goal and context of the communication. The studies' findings also cast doubt on the notion that the native speaker should always serve as benchmark for judging communicative effectiveness, especially with tests of language for specific purposes, where native speakers and second-language learners alike may lack the requisite skills for the kind of effective interaction demanded by the context. (C) 2017 Elsevier Ltd. All rights reserved.
This chapter highlights the role of English proficiency in academic study and the associated language assessment issues that emerge in the higher education environment. It considers validity issues surrounding the design and use of the tests used for (i) establishing minimum English entry requirements, (ii) identifying language support needs postentry and/or making English course placement decisions, (iii) establishing readiness to teach academic content through the medium of English, (iv) assessing the adequacy of English proficiency in the context of mainstream academic assignments, and (v) gauging the English standards achieved at exit from the university, including, by implication, students’ linguistic readiness to enter the workforce. It is argued that while research and development initiatives instigated by powerful testing agencies have contributed greatly to our thinking about language and have shaped the field of language assessment as we know it today, many problems remain. There are still uncertainties about how best to define and capture the academic language proficiency construct for testing purposes, in ways which can serve highly diverse student populations in contexts which are increasingly internationalized and technology mediated. Current assessment activities focus too much on English standards at university entry and too little on policies and practices geared to monitoring and fostering students’ language development throughout the course of their academic study. Future research and actions which might address some of these challenges are proposed.
The aim of this paper is to investigate from a discourse analytic perspective task authenticity in the speaking component of the Occupational English Test (OET), an English language screening test for clinicians designed to reflect the language demands of health professional–patient communication. The study compares the OET speaking sub-test roleplay performances of 12 doctors who were successful OET candidates with practice Objective Structured Clinical Examination (OSCE) roleplay performances of 12 international medical graduates (IMGs) preparing for the Australian Medical Council clinical examination. The premise for the comparison is that the OSCE roleplays can represent communication practices that are valued within the medical profession; therefore a finding of similarity in the discourse structure across the OET and the OSCE roleplays could be taken as supporting the validity of the OET as a tool for eliciting relevant communication skills in the medical profession. The study draws on genre theory as developed in Systemic Functional Linguistics (SFL) in order to compare the roleplay discourse structure and the linguistic realizations of the two tasks. In particular, it examines the role relationships of the participants (i.e. the tenor of the discourse), and the ways in which content is represented (i.e. the field of the discourse) by roleplay participants. The findings reveal some key similarities but also important differences. Although both tests inevitably fall short in terms of authentic representation of real world interactions, the findings suggest that the OET task, for a range of reasons including time allowances, training of test interlocutors, and the limits of contextual information provided to candidates, constrains candidate topic exploration and treatment negotiation, compared to the OSCE format. The paper concludes with proposals for mitigating these limitations in the interests of enhancing the OET’s capacity to elicit more professionally relevant language and communication skills.
Passing a test of English communication skills is mandatory for overseas-trained health professionals seeking entry to clinical practice in many countries: for example in Finland, in Germany, and in many English-speaking countries. An ongoing challenge facing those responsible for testing these skills has been the need to link the language performance standards required for entry into clinical settings with the actual communication demands of the clinical context. The term used to denote such linkage between a language test and the relevant nontest setting is authenticity, although the notion is somewhat contested in our field. Early contributors to this journal have claimed that the quest for authenticity in language testing is chimerical (Spolsky, 1985; Stevenson, 1985) since tests are a necessarily decontextualized and indirect means of predicting performance in the real world. Nevertheless, it is now generally agreed that for tests developed within the communicative paradigm some degree of authenticity is worth striving for, these constraints notwithstanding. Early discussions of authenticity were focused mainly on input materials and the question of whether these should, in the interests of naturalness, be used in undoctored form on a test. Widdowson (1979) was perhaps the first to point out that it was not the nature of the texts that was at issue but rather whether they were put to use in a manner consonant with the author’s intentions. This notion of authenticity as involving an interaction between the language users and the text or task was elaborated for testing purposes by Bachman (1990). Authenticity is, however, only one of several (potentially competing) components in Bachman’s framework of test usefulness and may not necessarily be the prime consideration in all testing situations. In LSP (language for specific purposes) testing, on the other hand, it is paramount. Representation of the target context in a way that will engage the specific abilities
This paper explores the underlying construct of both the English proficiency test for pilot and air traffic controller radiotelephony communication developed and administered in Korea and the ICAO language proficiency testing policy on which the test in Korea is based. It does so by canvassing the opinions of Korean airline pilots and air traffic controllers through 400 questionnaire and 22 interview sources. Results reveal a lack of fit between the policy construct and the reality through which the goals and objectives of the policy are accomplished and strong disapproval of the ICAO’s espoused construct and the associated Korean English test from language users in the target domain. This study confirms the importance of eliciting views from such stakeholders (i.e., domain experts) who are well-placed to determine what really matters for communicative success in the context of concern.
Gaining insights from domain experts into how they view communication in real world settings is recognized as an important authenticity consideration in the development of criteria to assess language proficiency for specific academic or occupational purposes. These “indigenous” criteria represent an articulation of the test construct and should therefore reflect what is germane to the particular domain of language use rather than general language-focused criteria familiar from other language tests. The methodological question of how to elicit such insights is, however, complex and has been addressed by various researchers using different methodological and theoretical frameworks. The paper draws on data from a larger research project to explore the affordances and constraints of more or less direct approaches to eliciting domain experts’ perspectives on what matters for effective communication in the workplace. The domain experts in this case were physiotherapy educators and supervisors. The study offers a qualitative comparison of expert feedback gathered from three different sites. Two were in the workplace where the communication skills of physiotherapy students in training were assessed routinely and the feedback given to them was naturally occurring rather than elicited. The third was a more artificial workshop setting in which video-recorded interactions between student and patients or simulated patients (i.e., actors role-playing a patient) were shown to two groups of expert informants who were then asked by the researcher to comment on the strengths and weaknesses of each performance. A qualitative analysis revealed that the nature of expert feedback differed significantly at each site, with the routinely occurring feedback containing scant and vague reference to language and communication aspects. The workshop setting, although it was less authentic, yielded much richer insights into the physiotherapists’ views about workplace communication. The implications of our findings for the development of relevant language test criteria are considered.
In order to broaden the scope of the discussion, this chapter presents four case studies from higher education contexts outside of Australia and New Zealand, where post-entry language assessments of incoming university students have been designed or implemented in rather different ways. In presenting each case we sketch the policy context as well as giving information on key features of test design, delivery, reporting, and use of test results, with specific examples where possible. The four cases have been selected to illustrate particular issues germane to post-entry language assessment and academic language enhancement.
The pedagogical focus of many genre studies in the field of applied linguistics has produced a wealth of materials designed to raise students’ awareness of the purposes, rhetorical structures, linguistic features, and contexts associated with particular educational genres. The desire to pin down the key characteristics of these genres has also resulted in a conceptualization of genres as rather more stable and constraining/normative than is the case in other disciplines such as literary studies and linguistic anthropology. In this chapter, we report on a rhetorical genre-based analysis of a spoken classroom event in the discipline of architecture - an event that was identified in the current study as both recurrent and patterned. As in many genre studies in the field of applied linguistics, we sought to characterize of the genre for teaching and learning purposes. Less usual was the case study approach adopted here, focusing on one teacher and his use of this classroom genre. A case study approach allowed us to explore the pattern and variability in the teacher’s improvisational pedagogical style. More generally, we want to argue that a study of particularity (in this case of one teacher’s use of a classroom genre) has the potential to contribute to a broader understanding of genre and generic boundaries. The chapter concludes by discussing the pedagogical implications of individual variation as well as underlining the need for a concept of genre in applied linguistics that allows a space to consider the tension between stability and creativity in language use.
This study investigates the impact of raters' language background on their judgements of the speaking performance in the College English Test-Spoken English Test (CET-SET) of China, by comparing the rating patterns of nonnative English-speaking (NNES) teacher raters, who are currently employed to assess performance on the CET-SET, with those of 'ideal' norm-owning native English-speaking (NES) teacher raters. Many-facet Rasch measurement and content analysis were applied to analyse the scores and stimulated recall data collected from the two rater groups. The results indicate that, although NES and NNES raters have somewhat different approaches to rating, the outcomes of the rating process are broadly similar, as are the categories that inform their judgements. We discuss the implications of these results for using raters from different language backgrounds for scoring high-stakes speaking tests, for the debate on NS norms for language testing in general and for the validity of the CET-SET rating scale in particular.
In this book, the authors describe the development and validation of a web-based test of second language pragmatics for learners of English.
In line with expanded conceptualizations of validity that encompass the interpretations and uses of test scores in particular policy contexts, this report presents results of a comparative analysis of institutional understandings and uses of 3 international English proficiency tests widely used for tertiary selection—the TOEFL iBT® test, the International English Language Testing Service (IELTS; Academic), and the Pearson Test of English (PTE)—at 2 major research universities, 1 in the United States and the other in Australia. Adopting an instrumental case study approach, the study investigated levels of knowledge about and uses of test scores in international graduate student admissions procedures by key stakeholders at Purdue University and the University of Melbourne. Data for the study were gathered via a questionnaire eliciting fixed‐choice responses and supplemented with qualitative interview data querying the basis for participants' beliefs, understandings, and practices. The study found that the primary use of language‐proficiency test scores, whether TOEFL®, IELTS, or PTE, by those involved in the admissions process at both institutions was often limited to determining whether applicants had met the institutional cutoff for admission. Beyond this focused and arguably narrow use, language‐proficiency test scores had little impact on admissions decisions, which largely depended on other required elements of applicants' admissions files. In addition, and despite applicants having submitted test scores that met the required cutoffs, survey respondents and interviewees often indicated dissatisfaction with enrolled students' levels of English‐language proficiency, both for academic study and for other roles within the university and in subsequent employment. A slight majority at both institutions indicated that they believed the institutional cutoffs represented adequate proficiency, while the remainder indicated that they believed the cutoffs represented minimal proficiency. The tension created by users' limited use of language‐proficiency scores beyond the cut, uncertainty about what cutscores represent, the assumption on the part of many respondents that students should be entering with language skills that allow success in graduate studies, and subsequent dissatisfaction with enrolled students' actual language proficiency may contribute to a perception that English‐language proficiency test scores are of questionable value; that is, perceived problems reside with the tests, rather than with how test scores are used and interpreted by those involved in the admissions process. At the same time, respondents at both institutions readily acknowledged very limited familiarity with or understanding of the English‐language tests that their institutions had approved for admissions. Owing to this lack of familiarity, a substantial majority at both institutions indicated no preference for either the TOEFL or the IELTS, counter to our expectation that score users in a North American educational context would prefer the TOEFL, while those in an Australian educational context would prefer the IELTS. The study's findings enhance understandings of test attitudes and test use. Findings may also provide insight for ETS and other language test developers about the context‐sensitive strategies that could be needed to encourage test score users to extend their understandings and use of language‐proficiency test scores.