Large data sets from a state reading assessment for third and fifth graders were analyzed to examine differential item functioning (DIF), differential distractor functioning (DDF), and differential omission frequency (DOF) between students with particular categories of disabilities (speech/language impairments, learning disabilities, and emotional behavior disorders) and students without disabilities. Multinomial logistic regression was employed to compare response characteristic curves (RCCs) of individual test items. Although no evidence for serious test bias was found for the state assessment examined in this study, the results indicated that students in different disability categories showed different patterns of DIF, DDF, and DOF, and that the use of RCCs helps clarify the implications of DIF and DDF.
Abstract Some students are less accurately measured,by typical reading tests than other students. By asking teachers to identify students whose performance,on state reading tests would likely underestimate their reading skills, this study sought to learn about characteristics of less accurately measured,students while also evaluating how well teachers can make such judgments. Twenty students identified by eight teachers participated in structured interviews and completed brief assessments matched to characteristics their teachers said impeded,the students’ test performance. Researchers found information from evidence provided by teachers, teacher and student interviews, and student assessments that confirmed teacher judgments for some students and information that failed to confirm or was at odds with teacher judgments for other students. Along with observations about student characteristics that affect assessment accuracy, recommendations,from the study include suggestions for working with teachers who are asked to make judgments about test accuracy and procedures for confirming teacher judgments. Identifying Less Accurately Measured Students 3 Identifying Less Accurately Measured,Students Test scores provide imperfect estimates of students’ knowledge,and skills. Statistics such as reliability coefficients, standard errors of measurement, and confidence intervals reflect random,measurement,error that affects all students. Any single test score may over- or underestimatea student’s knowledge and skills. In a sense, then, most students are to some extent inaccurately measured. But some students are less accurately measured,than others. In addition to random measurement error that clouds the picture for all students, some students’ test scores also include systematic error that further distorts the picture of their knowledge,and skills. Some systematic error produces test results that give a misleadingly high impression of a student’s knowledge,and skills; some produces a misleadingly low impression. Some systematic error affects many,test takers uniformly; some affects certain test takers differentially. For example, uniformly inflated results may be produced by a test that included poorly designed items that gave away the answers. Differentially inflated results may occur when some test takers cheat. Uniformly suppressed results can arise from scoring errors that mark correct responses as incorrect. Differentially suppressed results can be caused by personal characteristics that impede test performance,more for some students than for others. Common,examples,of such
This review examines the measurement of academic motivation in college students. It distinguishes pencil-and-paper group-administered instruments according to their conceptions of academic motivation: academic motivation taken as a single general motivation, as single specific motivations, or as a complex of motivations. It evaluates these classes of instruments in terms of the interpretability and the utility of the information each type of instrument is likely to provide.