
Despite the general consensus attributing an advantage to female readers, there are indications of their unsatisfactory performance on some reading tests. Adopting a cognitive approach, this article attempts to determine the underlying skills/attributes of and possible gender-based differences in a high-stakes nationwide English as a Foreign Language (EFL) reading comprehension test. Through two study phases, we utilized more recent, rigorous methodologies/frameworks of Cognitive Diagnostic Assessment (CDA) to scrutinize the existence of any gender-related differential item functioning (DIF) and/or differential attribute functioning (DAF) under the generalized deterministic input, noisy "and" gate (G-DINA) model. After our panel of expert applied linguists confirmed the existence of six reading attributes, the empirical evidence indicated that the Q-matrix is valid for the item response data of the EFL reading comprehension test. The results of the tests also showed that three test items and two attributes (i.e. sub-skills) of the reading comprehension test have large DIF and DAF against female examinees. The study discusses and concludes with the implications of the findings (including application of CDA and DIF/DAF methodology) for the teaching, testing, and researching of reading skills in general and high-stakes EFL reading examinations in particular.
In language testing research, fluency is typically operationalized through single, isolated temporal features. However, in natural speech, fluency (or disfluency) is often characterized by the clustering of multiple temporal features, collectively revealing the speaker's automaticity in speech production. In this study, we compare disfluency co-occurrence features with commonly used, single disfluency features in their prediction of language proficiency. We examined 31 (dis)fluency features of IELTS speaking responses by 278 L2 English speakers across three L1 backgrounds (Mandarin Chinese, Punjabi, and Arabic). These features were subjected to principal component analysis and ordinal regression analysis to predict IELTS band scores. The findings indicate that both single and co-occurrence disfluency features can effectively explain oral language proficiency, although disfluency co-occurrence features add nuances to the interpretation and implications for L2 speaking assessment research.
In Iran, a shift in educational policy was introduced 15 years ago, placing a stronger emphasis on assessment for learning (AfL). However, this shift has not been fully embraced by teachers of English as a Foreign Language (EFL) because of various factors including limited training in assessment and exam-oriented practices. This study explores how a professional development program with a strong emphasis on integrating theory and practice and teacher co-construction of knowledge can support teacher development in assessing literacy and implementing AfL in the classroom. The one-month professional development program included four 90-minute online workshops and teacher interactions in the forum, where they responded to the intervention assignments and reacted to their colleagues' responses. We traced changes in teachers' understanding of classroom assessment and how they could merge AfL and assessment of learning (AoL) in their classroom. Data included pre- and post-intervention open-ended questionnaires and teachers' forum interactions. We present the unique developmental trajectories of two EFL teacher participants to build an argument for praxis, a dialectical partnership between researchers/teacher educators and teachers, which allows these trajectories to emerge. The implications of the findings of this study are discussed with reference to teacher training programs and in-service teacher training.
Assessment quality (AQ) refers to key factors such as validity, reliability, fairness, and trustworthiness, all of which are essential for making accurate inferences and fair decisions about individuals in both large-scale and classroom assessments. While perspectives on AQ in large-scale assessments have significantly evolved and expanded over the past century, an ongoing debate revolves around the extent to which AQ criteria, originally developed in and for large-scale assessment contexts, apply to the context of classroom assessments. This paper reviews various perspectives surrounding this question and explores how the distinctive characteristics of classroom assessment can impact the conceptualization and application of AQ. The paper also reviews the limited research on how teachers conceptualize and enact AQ in classroom assessment and highlights the need for more research to explore when, how, and why second language (L2) teachers consider and address AQ-related questions during the design, selection, and use of assessments in the L2 classroom. It is suggested that such research has the potential to contribute significantly to the development of a framework of AQ specifically tailored to L2 classroom assessment practices.
This study investigates teachers' informal formative assessment (IFA) signalled by negative evidence in classroom interactions. In this single-case study, we examined classroom recordings of a language teacher and analysed micro-moments of the interaction using Conversation Analysis (CA). Concentrating on detailed excerpts of classroom interactions, we discuss how negative evidence signals the teacher's IFA process. We provide evidence that this process paves the way for maximising learning opportunities and shaping learner contributions in the immediate/subsequent turns of speech following a teacher's assessment and feedback, thus underscoring the concept of Classroom Interactional Competence (CIC). By presenting insights into the interactional patterns of IFA in second language (L2) classes, our findings contribute to the growing body of research integrating assessment and classroom interaction.
The Test of Chinese as a Heritage Language, or Huawen Shuiping Ceshi (HC), is a newly developed proficiency test by a university in Guangzhou, China. It defines "Heritage Language" as a family-transmitted language that is non-dominant in broader society and often incompletely acquired. The reading test is one of its three subtests. To investigate the cognitive patterns of learners of Chinese as a heritage language and to provide diagnostic assessment for test takers and educational institutions, the Rule Space Model (RSM) was implemented to conduct a diagnostic research of 236 test takers' responses on the reading test (Level 3). The research results indicate that: (1) the reading attributes of Chinese as a Heritage Language (Reading) were relatively accurately determined, and the hierarchical relationships of attributes were supported by the empirical data; (2) there were 12 mastery patterns of reading attributes for learners of Chinese as a heritage language. The RSM was successfully implemented in the diagnostic assessment of the reading test. This study can enhance the understanding of Chinese reading ability, advance the application of cognitive diagnosis in Chinese reading tests, and offer more detailed insights for reading instruction.
It is inevitable that second language (L2) learners often resort to test-wiseness (TW) strategies in reading tests to obtain higher scores. Previous research has reported mixed findings on whether TW strategies enhance L2 reading performance, but such strategies are often treated as a single, undifferentiated category. The aim of this study was to categorize TW strategies into distinct types and investigate their effects on L2 reading performance. Participants were 531 Taiwanese university students who completed a Test of English for International Communication (TOEIC) reading test and a TW strategy questionnaire to self-evaluate strategy use frequency. In the questionnaire, we divided TW strategies into four categories: time-using, error avoidance, deductive reasoning, and cue-using strategies. We then used Item Factor Analysis and structural equation modeling to examine the relationship between different types of TW strategies and students' TOEIC reading scores. Results showed that error avoidance strategies were the only type with a positive effect on reading performance. While deductive reasoning strategies could provide shortcuts to potentially determine the correct answer, our results indicated that over-reliance on such strategies may negatively impact performance. The implications for L2 teaching and future research are also discussed.
The validity of speaking tests is often defended by demonstrating their cognitive validity - the extent to which test tasks elicit mental processes mirroring real-world speaking. Traditional cognitive models of speech production, derived from psycholinguistics, emphasise individual processing and overlook the interactive nature of dialogue. This commentary explores an alternative framework that more directly integrates comprehension and production. Drawing on conversation analysis and recent cognitive modelling, it highlights the importance of predictive mental mechanisms in turn-taking. It advocates for greater interdisciplinary research to develop more comprehensive models of cognition for interaction.
This study investigated the effectiveness of using a collaborative problem solving-scenario to measure the examinees' situated English language proficiency. Examinees worked collaboratively with simulated peers online to build understandings to deliver an evidence-based oral pitch. The protocol was developed using Purpura and Turner's (2018) LOLA framework. The scenario provided opportunities to measure the display and development of second and foreign language proficiency. The paper presented evidence for four claims. The first examined the functionality of the measures within the scenario, showing that each displayed adequate psychometric properties, thus supporting substantively meaningful score interpretations about the constructs. The second targeted the validity of the culminating speaking task by demonstrating through MG-theory and MFRM analyses that the task is a generalizable and dependable measure of situated SFL speaking ability. The third addressing scenario interrelatedness showed through path analysis that building reading and listening understandings predicted the ability to share these understandings orally. The fourth targeted the assumption that the SBLA could be "instructional" by showing significant examinee gains in topical knowledge across administrations. Finally, the results suggested that for those interested in measuring SFL proficiency within a rich sociocultural context with teaching and learning embedded in the design, SBLA is a promising technique for achieving this.
This paper examines the use of scenario-based language assessment (SBLA) to measure situated Korean language proficiency in a Korean as a foreign language (KFL) setting. Specifically, this study investigated how a Korean SBLA (K-SBLA) serves as a valid and reliable measure of test-takers' ability to use Korean to build, consolidate, and share knowledge within a goal-oriented scenario. With 51 participants, the results suggest that the K-SBLA reliably elicited performances indicative of situated Korean proficiency. The study also explored performance differences between heritage and non-heritage Korean learners on the culminating scenario goal task: a proposal-pitching speaking task. These findings highlight the potential of SBLA as an innovative approach to assessing situated language proficiency in KFL contexts.
Scenario-based language assessments (SBLAs) measure second and foreign language proficiency through real-world tasks that build critical 21st-century skills. However, their complexity poses challenges to scalability, task authenticity, and validity. This article proposes using the Learning-Oriented Language Assessment (LOLA) framework to guide a partial interpretation argument for integration of artificial intelligence (AI) technologies into a B2 level SBLA. It explores how AI technologies, including generative AI, automated item generation platforms, adaptive feedback systems, and virtual agents can support five LOLA dimensions: proficiency, elicitation, social-interactional, socio-cognitive, and technological. This approach could support the scalability, validity, and pedagogical effectiveness of AI-enhanced SBLA through ongoing validation research.
This paper explores the use of a scenario-based academic speaking test to investigate strategic competence as a key dimension of L2 academic speaking ability and its relationship with performance on a complex academic speaking task. Drawing on a sociocognitive perspective, strategic competence in this study is defined as the use of cognitive resources involved in processing L2 information. The study involved 155 test-takers who completed a scenario-based language assessment (SBLA) simulating an 'intro to journalism' online class, where the culminating activity required an oral post to an online discussion forum. Throughout the test, test-takers were asked to build, consolidate, and share understanding about a journalism topic by engaging with audiovisual materials and performing a series of coherently sequenced speaking and strategy tasks. The strategy tasks were designed to elicit and assess test-takers' ability to use cognitive and metacognitive strategies (i.e. strategic competence). The study findings support the psychometric functionality of the SBEST along with evidence of learning gains. The results show that strategic competence is integral to L2 academic speaking ability and is also a significant contributor to L2 speaking performance. Finally, the paper offers insights into how strategic competence can be operationalized and scored in L2 assessments.
The present article employs eye tracking methodology to investigate the influence of item type (explicit and implicit) on test-taker attention to visual cues in video L2 listening tests. The findings reveal that implicit items direct test takers' attention toward nonverbal communicative cues related to affect more than explicit items, and that explicit items may direct more attention to illustrative gestures and objects. These findings raise questions about the construct addressed by implicit questions in multimodal video listening tests and call for careful consideration of test purpose when designing L2 listening tests.
In the Editorial of this special issue - Multimodality in second language assessment, we will introduce the concept of multimodality in language education, in order to set the scene for the special issue. We will then summarise the key findings of the nine papers included in the special issue and highlight the key findings and the research methods and features of multimodal tasks used in these studies, to identify gaps in research and practice. We will also reflect on what we have achieved through this special issue and propose future directions for implementing and researching multimodality in second language (L2) assessment.