Test items using open-ended response formats can increase an instrument's construct validity. However, traditionally, their application in educational testing requires human coders to score the responses. Manual scoring not only increases operational costs but also prohibits the use of evidence from open-ended items to inform routing decisions in adaptive designs. Using machine learning and natural language processing, automatic scoring provides classifiers that can instantly assign scores to text responses. Although optimized for agreement with manual scores, automatic scoring is not perfectly accurate and introduces an additional source of error into the response process, leading to a misspecification of the measurement model used with the manual score. We propose two joint models for manual and automatic scores of automatically scored open-ended items. Our models extend a given model from Item Response Theory for the manual scores by a component for the automatic scores, accounting for classification errors. The models were evaluated using data from the Programme for International Student Assessment (2012) and simulated data, demonstrating their capacity to mitigate the impact of classification errors on ability estimation compared to a baseline that disregards classification errors.
BackgroundLearning analytics dashboards (LAD) have been developed as feedback tools to help students self-regulate their learning (SRL) by using the large amounts of data generated by online learning platforms. Despite extensive research on LAD design, there remains a gap in understanding how learners make sense of information visualised on LADs and how they self-reflect using these tools.ObjectivesWe address this gap through an experimental study where a LAD delivered personalised SRL feedback based on interactions and progress to a treatment group, and minimal feedback based on the average scores of the lecture to a control group.MethodsAfter receiving feedback, students were asked to write down how they planned to adjust their study habits. These reflection texts are the target of this study. Three human coders analysed 1251 self-reflection texts from 417 students at three different times, using a coding system that categorised learning strategies, metacognitive strategies and learning materials.Results and ConclusionsOur results show that learners who received personalised feedback intend to focus on different aspects of their learning in comparison to the learners who received minimal feedback and that the content of the LAD influences how students formulate their self-reflection texts. Furthermore, the extent to which students incorporated suggested behavioural changes into their reflections was predicted by state measures like perceived helpfulness of the feedback. Our findings outline areas where support is needed to improve learners' sense-making of feedback on LADs and self-reflection.
Learning in asynchronous online settings (AOSs) is challenging for university students. However, the construct of learning engagement (LE) represents a possible lever to identify and reduce challenges while learning online, especially, in AOSs. Learning analytics provides a fruitful framework to analyze students' learning processes and LE via trace data. The study, therefore, addresses the questions of whether LE can be modeled with the sub-dimensions of effort, attention, and content interest and by which trace data, derived from behavior within an AOS, these facets of LE are represented in self-reports. Participants were 764 university students attending an AOS. The results of best-subset regression analysis show that a model combining multiple indicators can account for a proportion of the variance in students' LE (highly significant R2 between 0.04 and 0.13). The identified set of indicators is stable over time supporting the transferability to similar learning contexts. The results of this study can contribute to both research on learning processes in AOSs in higher education and the application of learning analytics in university teaching (e.g., modeling automated feedback).
Learning Analytics Dashboards (LAD) have been developed as feedback tools to help students self-regulate their learning (SRL), using the large amounts of data generated by online learning platforms. Despite extensive research on LAD design, there remains a gap in understanding how learners make sense of information visualised on LADs and how they self-reflect using these tools. We address this gap through an experimental study where a LAD delivered personalised SRL feedback based on interactions and progress to a treatment group, and minimal feedback based on the average scores of the class to a control group. Following the feedback, students were asked to state in writing how they would change their study behaviour. Using a coding scheme covering learning strategies, metacognitive strategies and learning materials, three human coders coded 1,251 self-reflection texts submitted by 417 students at three time points. Our results show that learners who received personalised feedback intend to focus on different aspects of their learning in comparison to the learners who received minimal feedback and that the content of the dashboard influences how students formulate their self-reflection texts. Based on our findings, we outline areas where support is needed to improve learners' sense-making of feedback on LADs and self-reflection in the long term.
By tailoring test forms to the test-taker's proficiency, Computerized Adaptive Testing (CAT) enables substantial increases in testing efficiency over fixed forms testing. When used for formative assessment, the alignment of task difficulty with proficiency increases the chance that teachers can derive useful feedback from assessment data. The application of CAT to formative assessment in the classroom, however, is hindered by the large number of different items used for the whole class; the required familiarization with a large number of test items puts a significant burden on teachers. An improved CAT procedure for group-based testing is presented, which uses simultaneous automated test assembly to impose a limit on the number of items used per group. The proposed linear model for simultaneous adaptive item selection allows for full adaptivity and the accommodation of constraints on test content. The effectiveness of the group-based CAT is demonstrated with real-world items in a simulated adaptive test of 3,000 groups of test-takers, under different assumptions on group composition. Results show that the group-based CAT maintained the efficiency of CAT, while a reduction in the number of used items by one half to two-thirds was achieved, depending on the within-group variance of proficiencies.
The NAEP EDM Competition required participants to predict efficient test-taking behavior based on log data. This paper describes our top-down approach for engineering features by means of psychometric modeling, aiming at machine learning for the predictive classification task. For feature engineering, we employed, among others, the Log-Normal Response Time Model for estimating latent person speed, and the Generalized Partial Credit Model for estimating latent person ability. Additionally, we adopted an n-gram feature approach for event sequences. Furthermore, instead of using the provided binary target label, we distinguished inefficient test takers who were going too fast and those who were going too slow for training a multi-label classifier. Our best-performing ensemble classifier comprised three sets of low-dimensional classifiers, dominated by test-taker speed. While our classifier reached moderate performance, relative to the competition leaderboard, our approach makes two important contributions. First, we show how classifiers that contain features engineered through literature-derived domain knowledge can provide meaningful predictions if results can be contextualized to test administrators who wish to intervene or take action. Second, our re-engineering of test scores enabled us to incorporate person ability into the models. However, ability was hardly predictive of efficient behavior, leading to the conclusion that the target label's validity needs to be questioned. Beyond competition-related findings, we furthermore report a state sequence analysis for demonstrating the viability of the employed tools. The latter yielded four different test-taking types that described distinctive differences between test takers, providing relevant implications for assessment practice.
Consulting Editors John Barnard EPEC, Australia G. Gage Kingsbury Psychometric Consultant, U.S.A. Juan Ramón Barrada Universidad de Zaragoza, Spain Wim J. van der Linden Pacific Metrics, U.S.A. Kirk A. Becker Pearson VUE, U.S.A. Alan D. Mead Illinois Institute of Technology, U.S.A. Barbara G. Dodd University of Texas at Austin, U.S.A. Mark D. Reckase Michigan State University, U.S.A. Theo H. J. M. Eggen Cito and University of Twente, The Netherlands Barth Riley University of Illinois at Chicago, U.S.A. Andreas Frey Friedrich Schiller University Jena, Germany Bernard P. Veldkamp University of Twente, The Netherlands Kyung T. Han Graduate Management Admission Council, U.S.A. Wen-Chung Wang The Hong Kong Institute of Education Matthew D. Finkelman, Tufts University School of Dental Medicine, U.S.A. Steven L. Wise Northwest Evaluation Association, U.S.A.
The shadow testing approach (STA; van der Linden & Reese, 1998) is considered the state of the art in constrained item selection for computerized adaptive tests. The present paper shows that certain types of constraints (e.g., bounds on categorical item attributes) induce a matroid on the item bank. This observation is used to devise item selection algorithms that are based on matroid optimization and lead to optimal tests, as the STA does. In particular, a single matroid constraint can be treated optimally by an efficient greedy algorithm that selects the most informative item preserving the integrity of the constraints. A simulation study shows that for applicable constraints, the optimal algorithms realize a decrease in standard error (SE) corresponding to a reduction in test length of up to 10% compared to the maximum priority index (Cheng & Chang, 2009) and up to 30% compared to Kingsbury and Zara's (1991) constrained computerized adaptive testing.
The assessment of a person’s traits is a fundamental problem in human sciences. Compared to traditional paper & pencil tests, computer based assessments not only facilitate data acquisition and processing but also allow for adaptive and personalized tests so that competency levels are assessed with fewer items. We focus on speeded tests and propose a mathematically sound framework in which latent competency skills are represented by belief distributions on compact intervals. Our algorithm updates belief based on directional feedback; adaptation rate and difficulty of the task at hand can be controlled by user-defined parameters. We provide a rigorous theoretical analysis of our approach and report on empirical results on simulated and real world data, including concentration tests and the assessment of reading skills.
Item Tree Analysis (ITA) can be used to mine deterministic relationships from noisy data. In the educational domain, it has been used to infer descriptions of student knowledge from test responses in order to discover the implications between test items, allowing researchers to gain insight into the structure of the respective knowledge space. Existing approaches to ITA are computationally intense and yield results of limited accuracy, constraining the use of ITA to small datasets. We present work in progress towards an improved method that allows for efficient approximate ITA, enabling the use of ITA on larger data sets. Experimental results show that our method performs comparably to or better than existing approaches.
The assessment of a person’s traits such as ability is a fundamental problem in human sciences. We focus on assessments of traits that can be measured by determining the shortest time limit allowing a testee to solve simple repetitive tasks, so-called speed tests. Existing approaches for adjusting the time limit are either intrinsically nonadaptive or lack theoretical foundation. By contrast, we propose a mathematically sound framework in which latent competency skills are represented by belief distributions on compact intervals. The algorithm iteratively computes a new difficulty setting, such that the amount of belief that can be updated after feedback has been received is maximized. We provide theoretical analyses and show empirically that our method performs equally well or better than state of the art baselines in a near-realistic scenario.
The assessment of a person’s traits such as ability is a fundamental problem in human sciences. Compared to traditional paper and pencil tests, computer based assessment not only facilitates data acquisition and processing, but also allows for real-time adaptivity and personalization. By adaptively selecting tasks for each test subject, competency levels can be assessed with fewer items. We focus on assessments of traits that can be measured by determining the shortest time limit allowing a testee to solve simple repetitive tasks (speed tests). Existing approaches for adjusting the time limit are either intrinsically non-adaptive or lack theoretical foundation. By contrast, we propose a mathematically sound framework in which latent competency skills are represented by belief distributions on compact intervals. The algorithm iteratively computes a new difficulty setting, such that the amount of belief that can be updated after feedback has been received is maximized. We rigorously prove a bound on the algorithms’ step size paving the way for convergence analysis. Empirical simulations show that our method performs equally well or better than state of the art baselines in a near-realistic scenario simulating testee behaviour under different assumptions.
Ulf Brefeld合作论文数Institute of Information Systems, Leuphana University of Lüneburg6