
This paper explores automatic assessment of written language proficiency. We have trained a CEFR-classifier for texts that have been written by Finnish as a second or foreign language (L2) learners. The aim of the study is to investigate to what extent we can model human ratings with the available L2 data for Finnish, whether the accuracy of the predictions varies at the different CEFR levels, and whether we can explain the discovered misclassifications. The FinBERT model was trained on the largest available CEFR-annotated datasets for L2 Finnish: ICLFI, LAS2, CEFLING, and TOPLING, which represent different kinds of genres, first language backgrounds, ages, and genders. The results are promising, with an F1-score of 72.7% and a .86 Pearson correlation between machine predictions and human assessors. Learners’ gender was not related to classification accuracy but learners’ L1 background may have some effect. However, text length seemed to cause misclassifications as unusually short or long samples were often assessed lower / higher than expected. While more annotated data is needed to train a more accurate model for higher-stakes assessment purposes, the open-source model developed in the study is likely useful for formative feedback purposes and paves the way for further work in the future.
This study is part of the DD-Lang project (see Leontjev et al., 2024) aimed at enhancing Finnish upper secondary school students’ reading in English and developing their understanding of their reading processes by bringing together two approaches that support language learning: diagnostic assessment (DiagA) and dynamic assessment (DA). The project designed online reading exercises for English based on retired Matriculation Examination items that implement graduated support (DA based mediation) for learners who struggle to complete the reading tasks. The tasks also include a chatbot that the learners can query when taking the exercises. The study reported here explores how this AI-powered chatbot might complement the standardised mediation accompanying the reading tasks. In the study, five students completed the online reading tasks, received mediation and interacted with the chatbot when needed. Their experiences were then discussed in a group session with their teacher and a researcher. Findings show that the students’ reactions to standardised mediation were varied: while some thought it changed the way they read, others considered it repetitive and quite general. Although limited in scale, this research also suggests that integrating an AI-based chatbot into DA can enhance learners’ reading comprehension processes and inform classroom practices.
This paper presents the integration of AI features into the language-teaching platform, Revita. The system is an intelligent online tutor, developed to support learners from lower-intermediate toward advanced levels, in several languages. Target skills currently include grammar, vocabulary, aural comprehension, and pronunciation. Based on authentic texts uploaded by the learners themselves, the system creates a rich variety of exercises that are tailored to the individual learner’s level of proficiency. Revita’s main guiding principle is personalization, motivated by current theories from educational science, notably Vygotsky’s concept of the Zone of Proximal Development and dynamic assessment, as well as the principles of diagnostic assessment. The linguistic foundation for the system comes from Construction Grammar, the goal being to build a complete inventory of constructions in the target language, as the basis for judging the correctness of the learners’ responses to the exercises. Revita is enhanced with AI tools from natural language processing, machine learning, and educational data mining.
In the paper we describe the theoretical and methodological basis of a multidisciplinary project, Aasis, focusing on automatic assessment of spoken interaction in L2 Finnish. The aim is to describe the project corpora that have been built for developing ASA (Automatic Speaking Assessment), not to study detailed empirical research questions as such even though we present some preliminary results from the data. The goals of the project are novel in three ways. Firstly, we aim to develop an ASA (Automatic Speaking Assessment) system to assess oral proficiency in dialogic rather than monologic speech. Secondly, our approach includes automatic assessment of nonverbal features in interaction, extending beyond the more conventionally used ASA tools. Thirdly, the language to be assessed is Finnish, a language with scarce previous studies on ASA.
Previous research on training and assessment of oral skills has mainly focused on English as L2, but since languages and learning contexts vary, it is imperative to study automatic speaking assessment (ASA) in other languages with local relevance as well. This paper summarizes a project which set out to develop a prototype tool to support training and assessment of oral skills in two low-resourced languages, Finnish and Swedish. This project addressed the applicability of automated assessment to measure multiple features of monologue speech, the accuracy of human ratings used for training a system based on automatic speech recognition (ASR) and the technical conditions of providing individualized feedback to improve student learning. Encouraging results for both Finnish and Swedish were gained when adapting a big pre-trained wav2vec2.0 speech model that was fine-tuned first with a larger L1 dataset, and then with an L2 dataset collected in the project. The results suggested that the most suitable features for automatic analysis were quantifiable fluency measures and vocabulary range. Machine and human estimates were most consistent for assessing fluency, range and accuracy, while the results were more controversial for pronunciation features other than fluency. The prototype will be further developed, fine-tuned and adjusted to address the needs of adult learners preparing for the final test of integration training in L2 Finnish.
Current Aptis Speaking rubrics reflect Common European Framework of Reference (CEFR) Phonological Control descriptors, which were revised in 2018 to better emphasize the intelligibility of second language (L2) speech. The present study investigates the validity of Aptis Speaking scores and related CEFR-based interpretations by comparing official Aptis Speaking scores to laypersons' assessments of intelligibility, a measure of accuracy of understanding using orthographic transcription (0-100%), and comprehensibility, a measure of ease of understanding using ratings on a scale of 1 to 9. Additionally, as Aptis Speaking tasks feature several target performance levels, we considered relationships between task complexity and speakers' intelligibility and comprehensibility. Archived speaking performances from 50 Aptis examinees were assessed by layperson listeners for intelligibility (n = 562) and comprehensibility (n = 567). Comprehensibility was generally a stronger predictor of Aptis Speaking scores than intelligibility for both overall and task-level scores. Segmented regressions revealed specific breakpoints in which predictive power was greatest for each dimension (intelligibility up to 70% in transcription accuracy; comprehensibility >= 2). In sum, results from the current study generally provided support for the current Aptis Speaking rubrics, though the way in which comprehensibility is described at the upper levels may benefit from more nuance.
This longitudinal, descriptive-exploratory case-study examined Iranian EFL learners' writing complexity through the lenses of Dynamic Systems Theory (DST). One hundred and twenty independent essays written by 12 intermediate to advanced female EFL learners in a TOEFL iBT preparation course over six months constituted the 43,478-word learner-generated corpus of this study. L2 Syntactic Complexity Analyzer was employed to analyze the length of production, sentence complexity, subordination, coordination, and particular structures. Moreover, three lexical analysis software programs including Coh-Metrix, Lexical Complexity Analyzer, and VocabProfile were employed to measure lexical density, diversity, and sophistication. The results of repeated measures analysis of variance (ANOVA) indicated significant differences between time and mean scores in five out of 14 syntactic indices. Correlational analyses among syntactic indices revealed positive relationships among the measures of the same sub-dimension of syntactic complexity. Meanwhile, particular structures enjoyed a positive correlation with both coordination and length of production. The analysis of syntactic and lexical relationships revealed that mean length of sentence, mean length of T-unit and mean length of clause closely corresponded with only lexical diversity. However, these syntactic indices revealed no significant correlations with both lexical density and sophistication. The findings suggest that different syntactic and lexical dimensions interactively comprise L2 writing complexity.
High-stakes test design must be informed by test takers' views so that the test takers can benefit from the tests (Fox & Cheng, 2015; Hamid et al., 2019; Jin, 2023; Kang, Miao, & Hirschi, 2024; O'Sullivan, 2012). High-stakes English tests are being administered in Francophone West Africa, yet little research has investigated test takers' experiences with them. Thus, we investigated 64 Malian test takers' views after taking a standardized English test, the Duolingo English Test (DET). Across 10 afternoon focus groups, we asked test takers about their perceptions of the test (which they had taken in the morning) and what they believed designers should do to make the test better for them. Emergent themes, found through analyses using MAXQDA, centered on the test's features/areas, technology, inclusivity, and professionalization opportunities. Test takers additionally discussed the test-taking anxieties they experienced and suggested recommendations for the test designers. In sum, local culture, infrastructure, and educational contexts interplayed with exam experiences. Test takers expressed a desire for more African-centric exam content. Overall, the test takers enjoyed participating, yet discussed that for them, the test would be more accessible if it cost less, was mobile-phone based, and could be paid for through local banking systems.
Extensive reading (ER) in a foreign language (L2) has been regarded as incompatible with language assessment primarily because ER aims to read for pleasure (Day & Bamford, 2002), which is believed to be discouraged by assessment. However, in reality, ER programs implemented in the L2 classroom often necessitate assessment, mainly to fulfill the institutional requirement of student evaluation (i.e., summative purposes of the assessment use). Less is discussed as to how assessment can be used to promote language learning through ER (i.e., formative purposes of the assessment use). The present paper aims to discuss how assessment should be conceptualized and implemented in ER practice from the perspective of Turner and Purpura's (2016) working framework of learning-oriented assessment (LOA). Based on the discussion, this paper emphasizes the necessity for further research on the effectiveness of LOA in language learning in ER programs.
This exploratory study investigated the assessment conceptions of university English as a foreign language (EFL) teachers and their reading and writing assessment practices. An online questionnaire about teachers' assessment conceptions and practices (i.e., why, how, and when they assess reading and writing) was developed, piloted, and then administered to about 100 university EFL teachers in English and non-English departments (e.g., engineering, medicine) across Tunisia. The findings indicated that the participants' assessment practices varied in relation to teacher qualification, experience, and context. Additionally, cluster analysis revealed two sub-groups that differed in terms of their assessment conceptions. The largest group viewed assessment as a tool to make students accountable and to improve education, while a smaller group viewed assessment as irrelevant. The first group reported using alternative assessment methods and assessing reading and writing for formative purposes significantly more frequently than did the second group, who reported using traditional assessment methods and assessing reading and writing for summative purposes (i.e., grading) more often. The paper reports the findings and discusses their implications for research on teacher assessment conceptions and practices.
Despite general recognition that, in human-led item writing, training is key to producing good-quality tests, there is little empirical research on what constitutes good practice in item-writing training for language test development. This has led to a lack of evidence-based guidance for those who (plan to) organise item-writing training. To help address this research gap, this study explored the perceptions of 25 novice item writers on the usefulness of a three-month, online induction item-writing training course. Views were collected via four feedback questionnaires administered at fixed points throughout the course and in semi-structured interviews conducted on course completion. Findings showed that participants particularly valued a clear bite-size course structure, extensive item-writing practice, timely and detailed tutor feedback, and regular opportunities for peer collaboration. Combining language testing theory with item-writing practice was also viewed as beneficial for learning. Participants held mixed views, however, on the platforms used for course delivery. Based on the findings, practical recommendations are proposed for how training for human-led itemwriting can be usefully structured and delivered.
This article reports a study that investigated typical utterance fluency features' development across three oral proficiency levels (A2-B2) as measured by a local application of the Common European Framework of Reference for languages. The speakers were 60 teenaged learners of English from Finland. Approximately 20 seconds of their semi-spontaneous speech was analysed for eleven speech features related to speed, breakdown, and repair fluency. Results reveal that the speakers' tempo, number of words produced and length of uninterrupted speech between pauses increased along with proficiency. Also, the number of silent pauses in unconventional (mid-clause) positions and length of silent pauses decreased along with increased proficiency. As an implication, second-language learners could benefit from fluency practice such as focus on speed of delivery and pausing. Moreover, these aspects could be considered more consistently in language proficiency scales.
This study examined the relationships between English language proficiency (ELP) test scores and academic success at the University of Hawai'i at M & amacr;noa (UHM) and considered whether academic outcomes differed for students who entered on the basis of different tests. Building on Isaacs et al. (2023), this study notably includes data on the use of Duolingo English Test (DET) in admissions decisions. More broadly, it fills a gap by examining outcomes in a new context (a large, public, less-selective U.S. university), using data from Fall 2022 and later when instruction was delivered primarily in-person after COVID, and covering a wider range of ELP scores than is typically represented in such research. In addition to GPA as an indicator of success, this study considered proportions of students who faced a negative academic action (academic probation or withdrawal) in relation to test submitted and made comparisons to international students who were exempt from submitting ELP scores. Students admitted unconditionally with higher ELP scores were also compared to those admitted conditionally with lower scores and further English language instruction requirements. Findings are relevant to valid use of the DET, IELTS, and TOEFL in admissions, and advance research on the topic by incorporating academic outcomes indicators beyond GPA.
This paper explores the innovative use of Book Creator as an emergency formative assessment tool in a higher education context. Book Creator, a platform for creating digital books, portfolios, and interactive multimedia projects, was adapted by the authors to urgently maintain students' progress and conduct formative assessment in the TEFL course during emergency remote teaching. The study involved 32 third-year students majoring in Secondary Education and four university teachers who specialised in training pre-service teachers. The authors employed a mixed-methods approach, combining quantitative analysis of student performance data with qualitative analysis of teacher reflections to evaluate the effectiveness of Book Creator as an emergency formative assessment tool within the TEFL course. Student performance data encompassed average scores for the generated TEFL books and students’ final exam scores, which were scrutinised using descriptive statistics, Krippendorff's alpha to ascertain interrater reliability, and Pearson correlation analysis to establish a relationship between TEFL book scores and final exam scores. The end-of-term outcomes in the TEFL course revealed a substantial correlation between students' proficiency in creating books and their academic performance. The study concludes that Book Creator is a potent formative assessment tool that fosters students' progress and autonomy in constructing pedagogical content knowledge in an emergency.
The COVID-19 pandemic revealed significant disparities in education, with vulnerable student populations—such as those with special needs, English language learners, and socioeconomically disadvantaged students—experiencing substantial learning losses. These challenges were particularly evident in writing skills, which received less attention during remote learning compared to reading and math. The present study examined the impact of the BalanceAI tutoring program, a technology-enhanced, community-based program aimed at improving writing and self-regulation skills in young students affected by the pandemic. Utilizing an integrative mixed methods analysis combining qualitative and quantitative data analyses, the study explored the interactions between tutors and students during writing tasks and their impact on skill development. Grounded in dynamic assessment (DA), the intervention focused on formative feedback and scaffolded learning, promoting self-regulated learning (SRL). Findings revealed that while the intervention did not yield statistically significant improvements in students' writing performance, qualitative analysis indicated growth in metacognitive skills and self-regulation, particularly in goal setting and self-reflection. The study highlights the complex nature of learning recovery post-pandemic and the critical role of mediated feedback in supporting skill development, emphasizing the need for inclusive, adaptive educational interventions to address the diverse needs of students.
The purpose of this paper was to investigate how English teachers employed formative assessment practices in emergency remote teaching (ERT) during the COVID-19 pandemic, and their perceived training needs in formative assessment. The data were collected from a questionnaire and interviews with teachers from Finland (n = 33) and Germany (n = 91). The results suggest that the teachers employed feedback, self-assessment and peer assessment in diverse ways, such as using them along with portfolios and online presentations. However, many teachers faced challenges with the technical aspects of providing feedback. The teachers also mentioned several areas in which they needed further training, such as identifying students’ needs and learning new techniques and formats of formative assessment. The results pave the way for reshaping formative assessment practices in post-pandemic language education.
The COVID-19 pandemic perhaps produced the largest disturbance our educational systems have ever seen. Assessment was particularly impacted as systems and educators had to identify the most optimal and adequate accommodations to meet both external and classroom-based mandates. When in-person teaching was possible again, various assessment practices that emerged to meet the pandemic challenges were and, still, are in use. This study aims to add to the exploration of practices and competences of Higher Education teachers in the context of second language (L2) classroom assessment. Based on Schumpeter's theory of creative destruction, the study attempts to unpack the abrupt albeit creative ways that teachers used when moving into a formative language assessment orientation in their practice, and explain how this worked, what the challenges were and which of these are still relevant and implemented in university systems. The paper contributes to the discussion of the tendency observed among teachers to resort mainly to formative assessment paradigms to address the challenges imposed during the pandemic and what the field of language assessment has learned from it.