This systematic review synthesizes empirical research on factors associated with student engagement in higher education blended learning. Drawing on 81 peer-reviewed studies published since 2016, we distinguished between four factors affecting student engagement: individual, instructional, interactional, and learning-environment factors. The synthesis shows that instructional factors, particularly learning support, assessment and feedback, course delivery, and learning materials, and individual factors, particularly intrapersonal competencies, were most frequently examined across studies. Interactional and learning environment factors were less frequently studied. In addition, behavioral, emotional, and cognitive engagement were the most commonly investigated dimensions, whereas social and agentic engagement received comparatively limited attention. We also identified substantial reliance on one-time self-report measures, with relatively limited use of longitudinal or digital-trace data to capture the dynamic nature of engagement. These findings highlight the need for more nuanced research on social and agentic engagement, improved operationalization of factors using longitudinal and digital-trace approaches, and greater attention to how adaptable instructional factors interact with less adaptable individual differences when designing blended learning environments.
In a pre-registered experiment, we investigated the effect of retrieval practice on the likelihood of choosing to review over taking a break (i.e., volitional task continuation). Participants studied lists of foreign language vocabulary word pairs under one of three conditions. In the study-once condition, participants first studied a filler list before studying a critical (to-be-tested) list of word pairs. In the restudy condition, participants studied the critical list twice. In the retrieval condition, participants studied the critical list once and then they received a retrieval practice test. Before taking a final cued recall test, all participants were first presented with a choice: Either review the vocabulary list one more time, or take a break. We found that prior retrieval practice increased participants’ subsequent likelihood of choosing to review. Exploratory analysis suggested that the increased likelihood to continue learning following retrieval practice could be the result of (a combination of) enhanced metacognitive monitoring and achievement motivation.
Background. Retrieval practice is a potent learning task that can enhance long-term retention, and potentiate learning during subsequent restudy. However, little is known about the effects that retrieval practice can have on students’ motivation. Aims. In two pre-registered experiments, we investigated the effect of retrieval practice on motivation. Sample. Sixty-eight, and 165 undergraduate Psychology students participated in Experiment 1 and 2, respectively. Methods. Participants studied a list of foreign language vocabulary word pairs under one of three conditions. In the study-once condition, participants studied the list and then did not engage in further study (Experiment 1), or they studied a filler list first followed by the criterion list (Experiment 2). In the restudy condition, participants studied the list twice. In the retrieval condition, participants studied the list once and then they received a test. Subsequently, all participants were provided the choice to spend the next three minutes before a final test on either reviewing the vocabulary list or taking a break. Results and Conclusions. In Experiment 1, we found no effect of learning task on participants’ willingness to further engage in the learning activity when given the choice. In Experiment 2, using a comparable design, a larger sample size, and equating for time on task, we found that retrieval practice during learning increased participants’ likelihood of choosing to review learning materials during a free period. Exploratory analysis of participants’ reasons for choosing to review suggested that the increased likelihood to persist in learning was the result of enhanced (achievement) motivation.
Adaptive learning technologies (ALTs) provide teachers with student data in teacher dashboards (TDs). However, there is substantial variation in dashboard use among teachers, and many find it difficult to draw conclusions based on student data. Teachers' skills, knowledge, and contextual conditions are believed to be essential in effective dashboard use. In this study, 26 primary school teachers using dashboards daily were interviewed about their perspectives on the skills, knowledge, and contextual conditions needed to facilitate dashboard use. Results indicated that teachers require a combination of skills, knowledge, and contextual conditions to make well-informed decisions while using a dashboard, emphasizing data literacy skills and pedagogical knowledge. In addition, teachers addressed other competencies, such as skills and knowledge related to the ALT curriculum and general computer proficiency. Also, contextual conditions at the school and technology level were found to be necessary. Based on these results, we propose various factors to explain teacher dashboard use.
Retrieval practice is a highly effective learning strategy that enhances long-term retention by encouraging the active recall of information. However, the optimal question format for maximizing knowledge retention remains uncertain. In this study, we compared the effect of very short answer (VSAQ) versus multiple-choice question (MCQ) practice tests on students’ knowledge retention. By analyzing these two formats, we aim to identify the most effective approach to retrieval practice, thereby helping to optimize its implementation and improve learning outcomes. In this randomized within-subjects study, students (n = 45) practiced with both VSAQs and MCQs in an extracurricular lifestyle course, without receiving feedback. The final retention test consisted of identical questions in both formats. A 2 × 2 repeated measures ANOVA was used to determine the effect of question format in practice testing and final test on final test score. Additionally, digital questionnaires were used to explore students’ test-taking experiences. The VSAQs were answered incorrectly more frequently on the practice tests and final test. There was no main effect of practice question format on final test performance, and no interaction effect between question format on the practice and final test. Regardless of question format, most students thought the practice tests were beneficial for learning. We found no evidence indicating that either MCQ or VSAQ is more effective for knowledge retention during retrieval practice. The lower initial retrieval success in the VSAQs, indicated by the higher degree of incorrect answers on the practice tests, might have limited their effectiveness during retrieval practice. To optimize the use of VSAQs in retrieval practice, it seems important to improve initial retrieval success to maximize learning outcomes. Not applicable.
Multiple choice questions (MCQs) offer high reliability and easy machine-marking, but allow for cueing and stimulate recognition-based learning. Very short answer questions (VSAQs), which are open-ended questions requiring a very short answer, may circumvent these limitations. Although VSAQ use in medical assessment increases, almost all research on reliability and validity of VSAQs in medical education has been performed by a single research group with extensive experience in the development of VSAQs. Therefore, we aimed to validate previous findings about VSAQ reliability, discrimination, and acceptability in undergraduate medical students and teachers with limited experience in VSAQs development. To validate the results presented in previous studies, we partially replicated a previous study and extended results on student experiences. Dutch undergraduate medical students (n = 375) were randomized to VSAQs first and MCQs second or vice versa in a formative exam in two courses, to determine reliability, discrimination, and cueing. Acceptability for teachers (i.e., VSAQ review time) was determined in the summative exam. Reliability (Cronbach's α) was 0.74 for VSAQs and 0.57 for MCQs in one course. In the other course, Cronbach's α was 0.87 for VSAQs and 0.83 for MCQs. Discrimination (average Rir) was 0.27 vs. 0.17 and 0.43 vs. 0.39 for VSAQs vs. MCQs, respectively. Reviewing time of one VSAQ for the entire student cohort was ±2 minutes on average. Positive cueing occurred more in MCQs than in VSAQs (20% vs. 4% and 20.8% vs. 8.3% of questions per person in both courses). This study validates the positive results regarding VSAQs reliability, discrimination, and acceptability in undergraduate medical students. Furthermore, we demonstrate that VSAQ use is reliable among teachers with limited experience in writing and marking VSAQs. The short learning curve for teachers, favourable marking time and applicability regardless of the topic suggest that VSAQs might also be valuable beyond medical assessment.
Introduction Multiple choice questions (MCQs) offer high reliability and easy machine-marking, but allow for cueing and stimulate recognition-based learning. Very short answer questions (VSAQs) may circumvent these limitations. We investigated VSAQ reliability, discriminative capability, acceptability, and knowledge retention compared to MCQs.Methods Dutch undergraduate medical students (n=375) were randomised to a formative exam with VSAQs first and MCQs second or vice versa in two courses, to determine reliability and discrimination. Next, acceptability (i.e., VSAQ review time) was determined in the summative exam. Knowledge retention at 2 and 5 months was determined by comparing score increase on the three-monthly progress test (PT) between students tested with VSAQs and students from previous years tested without VSAQs.Results Reliability (Cronbach’s α) was 0.74 for VSAQs and 0.57 for MCQs in one course. In the other course, Cronbach’s α was 0.87 for VSAQs and 0.83 for MCQs. Discrimination (Rir) was 0.27 vs. 0.17 and 0.43 vs. 0.39 for VSAQs vs. MCQs, respectively. Reviewing time of one VSAQ for the entire student cohort was ±2 minutes on average. No clear effect on knowledge retention after 2 and 5 months was observed.Discussion We found increased reliability and discrimination of VSAQs compared to MCQs. Reviewing time of VSAQs was acceptable. The association with knowledge retention was unclear in our study. This study supports and extends positive results of previous studies on VSAQs regarding reliability, discriminative capability, and acceptability in Dutch undergraduate medical students.### Competing Interest StatementThe authors have declared no competing interest.### Funding StatementThis study did not receive any funding.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:Educational Research Review Board of the Leiden University Medical Centre gave ethical approval for this work.I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesAll data in the present study are available upon reasonable request to the corresponding author.
Retrieval practice improves long-term retention. However, it is currently debated if this testing effect can be further enhanced by overtly producing recalled responses. We addressed this issue using a standard cued-recall testing-effect paradigm with verb–noun action phrases (e.g., water the plant) to prompt motor actions as a specifically powerful response format of recall. We then tested whether motorically performing the recalled verb targets (e.g., ?–the plant) during an initial recall test ( enacted retrieval) led to better long-term retention than silently retrieving them ( covert retrieval) or restudying the complete verb–noun phrases ( restudy). The results demonstrated a direct testing effect, in that long-term retention was enhanced for covert retrieval practice compared to restudy practice. Critically, enactment during retrieval further improved long-term retention beyond the effect of covert memory retrieval, both in a congruent noun-cued recall test after 1 week (Experiment 1) and in an incongruent verb-cued recall test of nouns after 2 weeks (Experiment 2). This finding suggests that successful memory retrieval and ensuing enactment contribute to future memory performance in parts via different mechanisms.
Taking a test on studied materials results in better delayed recall performance than restudying (a.k.a. thetesting effect). A common finding in testing effect research is that the effect depends on test format: the magnitude of the testing effect differs between free-recall, cued-recall, and recognition testing. This is explained by the effortful retrieval hypothesis: effortful successful retrieval results in better memory for an item than less effortful successful retrieval. However, the assumption that successful retrieval on different types of tests requires different levels of effort has not yet been tested. To test this assumption, we measured perceived mental effort on different test formats. Participants indicated free-recall was more effortful than cued-recall, and cued-recall more effortful than recognition. Furthermore, cued and free-recall yielded better cued-recall performance on a one-week delayed test than restudy or recognition. The results support the assumption that different practice test formats require different levels of mental effort.
In many innovations in technology and education in secondary schools, teachers are the crucial agents of these innovations. To select, match and support groups of teachers for particular school projects, school principals could be supported with insights into teachers’ beliefs about teaching, learning and technology. A teacher typology has been developed based on an online questionnaire completed by 1602 teachers from 59 Dutch secondary schools. Teachers are grouped on the basis of their beliefs about learned-centered teaching and attitudes towards technology, which underlie the school innovations that form the context of the current research. Five teacher types are distinguished: 1) Learner-centered teachers with technology, 2) Teachers critical of technology use in school, 3) Teachers uncomfortable with technology, 4) Teachers uneasy with learned-centered teaching and 5) Teachers critical of a clear-cut stance. This classification of teachers into these five types could be used to select or match the right group of teachers to a particular intervention or to organize different professional development activities for different types of school teachers.
In 2 experiments we investigated the efficacy of self-paced study in multitrial learning. In Experiment 1, native speakers of English studied lists of Dutch-English word pairs under 1 of 4 imposed fixed presentation rate conditions (24 × 1 s, 12 × 2 s, 6 × 4 s, or 3 × 8 s) and a self-paced study condition. Total study time per list was equated for all conditions. We found that self-paced study resulted in better recall performance than did most of the fixed presentation rates, with the exception of the 12 × 2 s condition, which did not differ from the self-paced condition. Additional correlational analyses suggested that the allocation of more study time to difficult pairs than to easy pairs might be a beneficial strategy for self-paced learning. Experiment 2 was designed to test this hypothesis. In 1 condition, participants studied word pairs in a self-paced fashion without any restrictions. In the other condition, participants studied word pairs in a self-paced fashion but total study time per item was equated. The results showed that allowing self-paced learners to freely allocate study time over items resulted in better recall performance.
Research has shown that testing during learning can enhance the long-term retention of text material. In two experiments, we investigated the testing effect with a fill-in-the-blank test on the retention of text material. In Experiment 1, using a coherent text, we found no retention benefit of testing compared to a restudy (control) condition. In Experiment 2, text coherence was disrupted by scrambling the order of the sentences from the text. The material was subsequently presented as a list of facts as opposed to connected discourse. For the incoherent version of the text, testing slowed down the rate of forgetting compared to a restudy (control) condition. The results suggest that the connectedness of materials can play an important role in determining the magnitude of testing benefits for long-term retention. Testing with a completion test seems most beneficial for unconnected materials and less so for highly structured materials.
The effect of repeated testing on delayed relearning of paired associates was investigated. Participants learned two lists of Lithuanian-Dutch word pairs until reaching the criterion of one correct recall from long-term memory. In one condition, items subsequently received three post-retrieval study trials and in the other condition items received three post-retrieval test trials. Participants returned one week later for delayed recall and relearning. Post-retrieval test trials resulted in better delayed recall performance than post-retrieval study trials. Moreover, we found that the items that were repeatedly studied or tested one week prior to relearning were relearned faster than a new set of similar (not previously presented) items. Most importantly, items were relearned faster when they had previously been learned under conditions of post-retrieval testing than items learned under conditions of post-retrieval study. Taken together, the results indicate that the benefits of repeated testing are not just limited to conscious recall on a delayed test. Repeated testing during initial learning is also a very effective strategy to enhance delayed relearning.
The present study examined the effect of presentation rate on foreign-language vocabulary learning. Experiment 1 varied presentation rates from 1 s to 16 s per pair while keeping the total study time per pair constant. Speakers of English studied Dutch–English translation pairs (e.g., kikker–frog) for 16 × 1 s, 8 × 2 s, 4 × 4 s, 2 × 8 s, or 1 × 16 s. The results showed a nonmonotonic relationship between presentation rate and recall performance for both translation directions (Dutch → English and English → Dutch). Performance was best for intermediate presentation rates and dropped off for short (1 s) or long (16 s) presentation rates. Experiment 2 showed that the nonmonotonic relationship between presentation rate and recall performance was still present after a 1-day retention interval for both translation directions. Our results suggest that a presentation rate in the order of 4 s results in optimal learning of foreign-language vocabulary.
In the present study we investigated the effect of repeated testing on item selection, retention, and delayed relearning of paired associates. Participants learned both related (easy) and unrelated (difficult) word pairs under conditions of repeated study and repeated testing. A retention test was given after both a 5-min and a 1-week interval. Following the 1-week retention test, participants received a relearning task. During the initial learning phase of the experiment, more related word pairs were successfully recalled on the practice tests compared to unrelated word pairs. Also, long-term retention benefits were found for items that were repeatedly tested compared to items that were repeatedly studied, regardless of item difficulty. The results suggest that the testing benefit following conditions of repeated testing cannot be attributed to mere item selection. Secondly, we found that delayed relearning was faster for previously restudied items compared to previously tested items. However, at the end of the relearning phase, repeated study and repeated testing 1 week prior to relearning resulted in comparable levels of recall performance. The results suggest that repeated testing can enhance delayed recall performance with little additional cost in terms of delayed relearning.