
This study presents the development and validation of a scale to assess L2 students’ self-perceived paraphrasing competence. Recognizing the limitations of existing assessment tools that primarily focus on surface-level linguistic transformations, this research establishes a multidimensional framework that captures both the mechanical and rhetorical dimensions of effective paraphrasing. The scale was developed through establishing the item pool, refining items, determining the scale format, teacher reviewing, and student piloting and revision, then validated with a sample of 780 L2 students in the Chinese context. The resulting instrument demonstrated robust psychometric properties based on the evidence gleaned from the exploratory factor analysis (EFA) (N = 250) and confirmatory factor analysis (CFA) (N = 530), which encompasses four core dimensions of paraphrasing competence: Understanding Paraphrasing Purposes, Understanding Paraphrasing Techniques, Understanding Paraphrasing Criteria, and Understanding Content Recontextualization, with 30 items. This research findings provide valuable insights into the theorization of paraphrasing. The scale serves as a diagnostic tool to help L2 students identify their strengths and weaknesses across the paraphrasing competence dimensions. The diagnostic feedback can inform their understanding of paraphrasing competence and guide them in setting learning goals to enhance their paraphrasing competence.
While metacognition is widely recognized as a critical factor in enhancing instructional effectiveness and student learning, its conceptualization and measurement within academic writing remain limited. In particular, there is a lack of valid instruments specifically designed to assess learners’ metacognitive knowledge of academic writing, and the underlying components of this construct remain unclear. To address these gaps, the present study developed and validated a multidimensional scale for assessing learners’ metacognitive knowledge of academic writing. Drawing on a sample of 331 Chinese university students, the study employed exploratory and confirmatory factor analyses to develop a 15-item instrument comprising three dimensions: Linguistic Knowledge, Genre Knowledge, and Cognitive Knowledge. The results demonstrated strong reliability and validity of the scale, aligning with rigorous psychometric properties. The instrument not only clarifies the construct of metacognitive knowledge of academic writing, but also provides empirical support for integrating metacognitive principles into academic writing instruction and assessment.
Artificial intelligence (AI) has increasingly been used to assess the academic writing of English-as-a-foreign-language (EFL) university students. This scoping review synthesized empirical studies comparing AI-assisted and human scoring of EFL academic writing from perspectives of Critical Language Testing (CLT) and Critical Digital Pedagogy (CDP). Following PRISMA-guided procedures, this review searched Web of Science, ERIC, and PubMed for empirical studies published between 2015 and 2025. Twelve studies met the inclusion criteria of peer-reviewed empirical studies comparing AI-assisted and human scoring in EFL university learners’ academic writing and were further analyzed through thematic synthesis guided by theoretical frameworks of CLT and CDP. The results revealed that AI-generated scores generally showed moderate alignment with human-created scores, particularly for surface-level linguistic features, but demonstrated lower agreement for higher-level writing aspects. AI tools were also frequently reported to be stricter than human raters and to be affected by prompting design. From a CLT perspective, fairness-related issues were underexplored, while assessment transparency and social implications were unevenly illustrated across the reviewed studies. From a CDP perspective, most studies reflected critical awareness of using AI and emphasized human oversight over AI-assisted scoring but rarely explained how teachers can practically and critically use AI to support dialogic assessment processes. The findings highlight the need for future studies to integrate technical validation of AI-assisted writing assessments with considerations of fairness, transparency, human agency, the critical use of AI, and dialogic pedagogies.
Academic and scientific writing in the first language (L1) is cognitively demanding, yet few Spanish-language computational tools support it through well-grounded pedagogical frameworks. This study presents PEUMO (Plataforma para la Escritura Universitaria con Mediación Online), an AI-based writing support tool grounded in Corpus Linguistics (CL) and Genre-Based Pedagogy (GBP), designed to enhance disciplinary writing in engineering education. Unlike most automated writing evaluation systems, which focus on micro-level feedback or L2 contexts, PEUMO integrates corpus-informed rhetorical analysis with NLP and transformer-based classification (BETO) to support genre-specific writing in L1 Spanish. This study describes the tool and provides multi-source empirical evidence (product, process, and perception) to illustrate its potential as a formative writing support system. Product evidence is based on a comparison of the textual quality of thesis introductions between two groups of engineering students. Process evidence draws on screen recordings to characterize students’ interactions with the tool, while perception data are derived from a survey addressing usability dimensions. The findings suggest that PEUMO contributes to current discussions on AI in writing assessment by illustrating how genre-based and corpus-informed approaches can be operationalized in intelligent systems to support disciplinary academic writing.
Summary writing is a common write-to-learn strategy in undergraduate STEM education, particularly for pre-class learning. Yet producing summaries that demonstrate deep comprehension is demanding and often requires instructional support. In response, Automated Summary Evaluation (ASE) tools have been developed to provide formative assessment and feedback on student summaries. This study examines the effectiveness of a generative AI-powered ASE tool designed to scaffold student engagement in pre-class summarization. We investigated whether AI support enhances editing behaviors and promotes concept learning while interacting with learner background characteristics. We analyzed 1081 revision attempts across seven topics over seven weeks from 49 undergraduates in an Introductory Physics course at a large public university. A longitudinal analytic approach using Linear Mixed-Effects Models was employed to address the research questions. Findings indicate that AI-powered formative feedback fostered revision behaviors associated with higher concept learning scores. Effective revisions occurred when students added concepts in response to AI feedback while avoiding careless deletions and surface-level sentence changes. Engagement and performance varied more by assignment than by tool proficiency, with AI scaffolds especially beneficial for students historically underperforming in STEM. These results underscore the importance of personalized feedback strategies that promote targeted revision across diverse learners.
Source-based writing requires integrating information across multiple documents, yet most Automated Writing Evaluation (AWE) systems evaluate only the final written product, overlooking the comprehension processes that precede the writing. We propose a comprehension-sensitive AWE framework that makes the reading-to-writing pipeline visible through knowledge graphs (KGs) and large language model (LLM) inference. The framework comprises three layers (artifact representation, inferential linking, and idea tracing) that together trace how students’ ideas move and transform across source texts, process artifacts, and final essays. As a proof of concept, we applied the framework to archival data from 132 students completing a source-based writing task. Integrated KGs were constructed for each student, tracing idea flow across four source texts, constructed responses, and final essays. Graph-based metrics correlated significantly with human rubric scores across multiple measurement configurations, providing evidence that the representational structures capture systematic variation in writing quality. A diagnostic case analysis demonstrates how graph representations make specific patterns of source engagement and idea transfer visible, revealing meaningful differences between two students who received identical holistic scores. Cohort-level analyses further reveal systematic variation in how students engage with and transform source ideas. The framework is designed to adapt across genres, task structures, and educational contexts, and offers a foundation for diagnostic feedback that is sensitive to the comprehension processes source-based writing demands.
Citation is an essential feature of source use in integrated writing. However, scarcely any research has investigated test-takers’ citation performance on timed integrated writing tests. This study addressed this gap by investigating undergraduate students’ citation performance on a timed integrated reading-to-write argumentative essay test. The study recruited 204 undergraduate participants, who completed an integrated reading-to-write essay based on five source texts. Their writing samples were manually coded to generate statistics on citation density, accuracy, distribution, formats, and functions, which were compared across three performance levels, source texts, and part genres (i.e., introduction, body, and conclusion) to identify features of citation patterns in timed integrated writing tests. The study provides valuable information for understanding citations, source use, and their potential as performance indicators of writing performance in integrated writing tests and research.
Previous studies on L2 phrase complexity predominantly focused on noun phrases, while the verb phrase complexity are few and mostly simple verb phrase structures, they have limited expressive capacity. Complex verb phrase structures (e.g., n-grams, VACs) also have limitations. The present study proposes a set of 41 Chinese complex verb phrase structures and are meaningful. A total of 246 verb phrase complexity measures are calculated from the dimensions of account, frequency, diversity, and density. The ability of these measures to predict writing quality is compared with that of 10 large-grained syntactic complexity measures. The results show that large-grained measures (Average sentence length) and verb phrase complexity measures (two verb-object structures and three adverbial-head structures) can respectively explain 14.4% and 41.9% of variance in writing scores. Our results illustrate the importance of Chinese complex verb phrase structures in assessing L2 writing quality.
As automated writing evaluation (AWE) systems proliferate, it is important to assess them for the extent to which they serve an authentic formative assessment purpose. We developed eRevise as an AWE system to engage upper elementary students in text-based writing, receive automated feedback, and revise their essay based on that feedback. In past work, we presented a validity argument for a response-to-text formative assessment. Here, we replicate evidence for the mediational processes we identified with a second response-to-text formative assessment. Beyond replication, we expand the validity argument in several ways: We examine multiple writing outcomes to understand whether the relationships in the data generalize. We also explore patterns of relationships to better understand which students are most likely to benefit from eRevise. Furthermore, we expand the investigation from feature score improvement alone to also consider students’ conceptual development over time. Our findings provide further evidence for sociocultural mechanisms supporting students’ conceptual development of evidence use in writing. We discuss the implications of our findings for future design of eRevise and AWE systems, in general. We also discuss the need for a recursive process where our findings also contribute to refined theory development and design of future AWE systems.(1). Formative Assessment; (2) Argumentative Writing; (3) Adaptive Expertise; (4) Conceptual Change; (5) Validity Argument.
Effective source integration is a complex but essential skill in academic writing; yet it remains difficult to teach and evaluate. Assessing source integration has important implications for formative feedback and instructional practice, but existing approaches face limitations, particularly in automated systems. The purpose of this study was to develop and validate an automated measure of source integration by extracting linguistic features from student essays and modeling them as a latent construct with confirmatory factor analysis. Predictive validity with human scores and generalizability across datasets using measurement invariance was tested and examined. The source integration construct included linguistic features related to citation, quotation, plagiarism, and semantic overlap with the source text. Results indicated strong alignment with human ratings (beta = .81, R2 = .65) and evidence of structural consistency across a new dataset with novel prompts and sources. Predictive utility analyses showed that the latent construct improved machine learning models and enhanced agreement with human ratings when paired with BERT embeddings. GPT-5.2 produced interpretable justifications but lower scoring reliability. These findings suggest that a source integration construct grounded in linguistic features can complement modern AI methods, providing a foundation for formative feedback on source-based writing that addresses issues of fairness and interpretability.
This article critically examines the advantages and limitations of Turnitin’s AI writing detector in second language (L2) writing assessment. Although Turnitin and similar detectors (e.g., GPTZero, Copyleaks, Writer, and the OpenAI Classifier) claim high accuracy in identifying AI-generated text, recent evidence reveals systemic biases against multilingual writers, whose syntactic regularity and formulaic phrasing are frequently misclassified as AI-generated. This article draws on recent empirical studies and argues that overreliance on such tools in some institutional contexts may contribute to epistemic injustice, false accusations, and the reinforcement of deficit ideologies in L2 contexts. This article draws on recent empirical studies and argues that overreliance on such tools in some institutional contexts may contribute to epistemic injustice, false accusations, and the reinforcement of deficit ideologies in L2 contexts. Rather than treating detection scores as verdicts of misconduct, we propose a dialogic, process-oriented framework that repositions Turnitin’s output as a formative cue, one that, when triangulated with drafts, reflective commentaries, and oral explanations, can encourage authorship transparency and pedagogical trust. The article thus advances a linguistic approach to L2 assessment that acknowledges AI’s presence while safeguarding student agency, integrity, and equity in writing evaluation.
GenAI feedback can support L2 writing, but its benefits vary across classrooms. This study examines when GenAI-mediated formative feedback may support learning within classroom assessment routines. Using a QUAL-dominant embedded mixed-methods multiple-case design in two university EFL writing classes, we examined how assessment enactment was associated with students’ trust calibration and verification practices. Across cases, clearer norms, learning-oriented accountability, and teacher positioning of AI as contingent input were associated with more frequent verification, deeper revisions, and stronger transfer. Under higher assessment pressure and weaker verification routines, students more often accepted AI feedback uncritically, showed less revision reasoning, and demonstrated weaker transfer, patterns consistent with what we term learning displacement. Using RAFE as a preliminary analytic framework, the study traces how classroom conditions, assessment enactment, trust calibration, and appropriation trajectories were related across the two cases. The findings suggest that verification-oriented norm design is an important condition for more sustainable GenAI-supported L2 writing development.
Feedback is a common pedagogical approach to enhance students’ writing performance. Research on feedback emotion has recently gained growing interest and sparked a series of empirical studies. To timely update research progress and guide future research in this promising area, this study provides a scoping review of 61 empirical investigations published in academic journals and ProQuest Dissertations & Theses from 2005 to 2025. Iterative content and thematic analyses of these publications were conducted to scrutinize the conceptualizations of emotions, research scopes, contexts, and methodological characteristics regarding approaches, instruments, reporting of methodological rigor, and study durations. The findings reveal that emotions in writing feedback are fundamentally context-sensitive and dynamic. Furthermore, the literature shows a concentrated interest in describing participants’ emotions within single-source, written feedback situations, while relying heavily on self-report instruments across qualitative, quantitative, and mixed-methods studies. This review provides several empirically grounded suggestions for future research based on the results.
This editorial introduces the 2026 Tools & Tech Forum, which examines how generative AI writing technologies are transforming writing assessment and pedagogy. Across the contributions, a central theme emerges: AI systems do not merely assist writers but actively shape what becomes recognizable as assessable and valued writing. This raises important questions about validity, authorship, language bias, and assessment justice. Collectively, the reviews invite readers to consider AI tools as rhetorical and assessment infrastructures that reframe the constructs and consequences of writing assessment in contemporary educational contexts.
This study investigated how young L2 learners engaged with an AI-assisted writing assessment feedback tool that employed the GPT-4o model. In particular, we focused on students' interactions with the AI chatbot to seek feedback and their subsequent use of the chatbot's responses. Conducted as part of a prototype development and usability testing efforts, the study involved eight teachers and their EFL/ESL students (N = 206) from upper elementary and middle school grades in Hong Kong, South Korea, Turkiye, and the United States. Students completed three Opinion writing tasks using the tool. Students' chat messages were coded, and textual analyses were conducted to examine students' revisions. Results indicated considerable variation in students' usage with the chatbot. During the outlining stage, students primarily sought support with content development and translation; during revision, they focused on correction and content-related feedback. Students' incorporation of AI feedback was evidenced by increased text length and more token additions than deletions or replacements. Both teachers and students rated the tool's translation (L1 support) and personalized, interactive feedback features as highly useful. Implications for L2 writing assessment and future research are discussed.
This study investigates L2 writers’ engagement with an Automated Writing Evaluation (AWE) system through the lens of psychological contract theory. Using a mixed-methods approach, this research examined the development of transactional and relational psychological contracts between students and an AWE system over a semester, and analyzed their impact on behavioral, cognitive, and emotional engagement. Through a latent profile analysis of 153 students’ submission patterns, we identified six students as focal participants for in-depth interviews and iterative draft analyses. The findings reveal that transactional contracts, particularly expectations of score improvement, were frequently fulfilled, motivating repeated revisions. However, this fulfillment often prompted strategic system “gaming” rather than genuine writing development. Conversely, relational contracts were routinely breached. Faced with unmet expectations, students adapted their strategies by reducing engagement with the AWE. They supplemented this system with Gen-AI tools and human feedback. The study underscores the dynamic and often fragile nature of learner-AWE relationships, highlighting that engagement is strongly influenced by the perceived expectations and obligations between learners and technologies. Implications include the need for AWE designers to enhance feedback transparency and support for higher-order writing skills, and for instructors to complement AWE use with multi-source feedback to sustain engagement and meaningful writing development.
While writing self-efficacy has long been acknowledged as a determinant of L2 writing performance, the mechanisms through which it translates into meaningful differences in learners’ writing achievements remain underexplored. To bridge this gap, our research investigated the synergistic relationship between L2 learners’ writing self-efficacy and writing performance, with particular attention to the serial mediating effects of writing anxiety and motivation. A total of 327 Iranian L2 learners were recruited from various universities through a maximum variation sampling technique. Participants first responded to three validated self-report questionnaires and then took the CUNY assessment test in writing (CATW), a standardized instrument that evaluates preparedness for college-level study through an integrated writing task. A structural equation modeling (SEM) was conducted to examine the interactions of the variables under investigation. The SEM results revealed a significant positive correlation between writing self-efficacy and performance. More importantly, writing anxiety and motivation were found to mediate this relationship both independently and sequentially. The study underscores the multifaceted essence of writing development by integrating psychological, affective, and motivational factors into a single framework. It highlights the importance of viewing writing performance not as a static outcome but as the result of interactive psycho-emotional processes that evolve in tandem.