Effective source integration is a complex but essential skill in academic writing; yet it remains difficult to teach and evaluate. Assessing source integration has important implications for formative feedback and instructional practice, but existing approaches face limitations, particularly in automated systems. The purpose of this study was to develop and validate an automated measure of source integration by extracting linguistic features from student essays and modeling them as a latent construct with confirmatory factor analysis. Predictive validity with human scores and generalizability across datasets using measurement invariance was tested and examined. The source integration construct included linguistic features related to citation, quotation, plagiarism, and semantic overlap with the source text. Results indicated strong alignment with human ratings (beta = .81, R2 = .65) and evidence of structural consistency across a new dataset with novel prompts and sources. Predictive utility analyses showed that the latent construct improved machine learning models and enhanced agreement with human ratings when paired with BERT embeddings. GPT-5.2 produced interpretable justifications but lower scoring reliability. These findings suggest that a source integration construct grounded in linguistic features can complement modern AI methods, providing a foundation for formative feedback on source-based writing that addresses issues of fairness and interpretability.
更多
查看译文
关键词
Source integration,Automated writing evaluation,Natural language processing,Validity,Formative feedback