When reading long and complex texts, students may disengage and miss out on relevant content. In order to prevent disengaged behavior or to counteract it by means of an intervention, it is ideally detected an early stage. In this paper, we present a method for early disengagement detection that relies only on the classification of scrolling data. The presented method transforms scrolling data into a time series representation, where each point of the series represents the vertical position of the viewport in the text document. This time series representation is then classified using time series classification algorithms. We evaluated the method on a dataset of 565 university students reading eight different texts. We compared the algorithm performance with different time series lengths, data sampling strategies, the texts that make up the training data, and classification algorithms. The method can classify disengagement early with up to 70% accuracy. However, we also observe differences in the performance depending on which of the texts are included in the training dataset. We discuss our results and propose several possible improvements to enhance the method.
The NAEP EDM Competition required participants to predict efficient test-taking behavior based on log data. This paper describes our top-down approach for engineering features by means of psychometric modeling, aiming at machine learning for the predictive classification task. For feature engineering, we employed, among others, the Log-Normal Response Time Model for estimating latent person speed, and the Generalized Partial Credit Model for estimating latent person ability. Additionally, we adopted an n-gram feature approach for event sequences. Furthermore, instead of using the provided binary target label, we distinguished inefficient test takers who were going too fast and those who were going too slow for training a multi-label classifier. Our best-performing ensemble classifier comprised three sets of low-dimensional classifiers, dominated by test-taker speed. While our classifier reached moderate performance, relative to the competition leaderboard, our approach makes two important contributions. First, we show how classifiers that contain features engineered through literature-derived domain knowledge can provide meaningful predictions if results can be contextualized to test administrators who wish to intervene or take action. Second, our re-engineering of test scores enabled us to incorporate person ability into the models. However, ability was hardly predictive of efficient behavior, leading to the conclusion that the target label's validity needs to be questioned. Beyond competition-related findings, we furthermore report a state sequence analysis for demonstrating the viability of the employed tools. The latter yielded four different test-taking types that described distinctive differences between test takers, providing relevant implications for assessment practice.
In this explorative study, we investigate how sequences of behaviour are related to success or failure in complex problem-solving (CPS). To this end, we analysed log data from two different tasks of the problem-solving assessment of the Programme for International Student Assessment 2012 study (n = 30,098 students). We first coded every interaction of students as (initial or repeated) exploration, (initial or repeated) goal-directed behaviour, or resetting the task. We then split the data according to task successes and failures. We used full-path sequence analysis to identify groups of students with similar behavioural patterns in the respective tasks. Double-checking and minimalistic behaviour was associated with success in CPS, while guessing and exploring task-irrelevant content was associated with failure. Our findings held for both tasks investigated, from two different CPS measurement frameworks. We thus gained detailed insight into the behavioural processes that are related to success and failure in CPS.
As Internet sources provide information of varying quality, it is an indispensable prerequisite skill to evaluate the relevance and credibility of online information. Based on the assumption that competent individuals can use different properties of information to assess its relevance and credibility, we developed the EVON (evaluation of online information), an interactive computer-based test for university students. The developed instrument consists of eight items that assess the skill to evaluate online information in six languages. Within a simulated search engine environment, students are requested to select the most relevant and credible link for a respective task. To evaluate the developed instrument, we conducted two studies: (1) a pre-study for quality assurance and observing the response process (cognitive interviews of n = 8 students) and (2) a main study aimed at investigating the psychometric properties of the EVON and its relation to other variables (n = 152 students). The results of the pre-study provided first evidence for a theoretically sound test construction with regard to students' item processing behavior. The results of the main study showed acceptable psychometric outcomes for a standardized screening instrument with a small number of items. The item design criteria affected the item difficulty as intended, and students' choice to visit a website had an impact on their task success. Furthermore, the probability of task success was positively predicted by general cognitive performance and reading skill. Although the results uncovered a few weaknesses (e.g., a lack of difficult items), and the efforts of validating the interpretation of EVON outcomes still need to be continued, the overall results speak in favor of a successful test construction and provide first indication that the EVON assesses students' skill in evaluating online information in search engine environments.
In large-scale assessments, performance differences across different groups are regularly found. These group differences (e.g., gender differences) are often relevant for educational policy decisions and measures. However, the formation of these group differences usually remains unclear. We propose an approach for investigating this formation by considering behavioral process measures as mediating variables between group membership and performance on the 2012 Programme for International Student Assessment complex problem solving (CPS) items. We found that across all investigated countries interactive behavior can fully explain gender differences in CPS, but cannot explain differences between students with and without a migration background. However, in some countries these results differ from the cross-country results. Our results indicate that process measures derived from log data are useful for further investigating and explaining performance differences between girls and boys and students with and without migration background. Educational Impact and Implications Statement The study suggests that the higher performance of boys compared to girls in complex problem solving seems to stem from gender-specific interaction with the problem space, while performance differences by migration status cannot be explained by behavioral differences. Specifically, the amount of exploration behavior seems to have a huge impact on complex problem solving performance. The study demonstrates how the formation of performance differences between different groups of students can be explained.
Complex problem solving (CPS) is a highly transversal competence needed in educational and vocational settings as well as everyday life. The assessment of CPS is often computer-based, and therefore provides data regarding not only the outcome but also the process of CPS. However, research addressing this issue is scarce. In this article we investigated planning activities in the process of complex problem solving. We operationalized planning through three behavioral measures indicating the duration of the longest planning interval, the delay of the longest planning interval and the variance of intervals between each two successive interactions. We found a significant negative average effect for our delay indicator, indicating that early planning in CPS is more beneficial. However, we also found effects depending on task and interaction effects for all three indicators, suggesting that the effects of different planning behaviors on CPS are highly intertwined.