Email zilberax@jmu.edu Abstract Despite the extensive testing for federal accountability mandates, college students’ understanding of federal accountability testing (e.g., No Child Left Behind, Race to the Top, Spellings) has not been examined, resulting in a lack of knowledge regarding how such understanding (or lack thereof) impacts college students’ behavior on accountability tests in higher education contexts. This study explores college students’ understanding and misconceptions of federal accountability testing in K-12. To this end, we crafted nine multiple choice items with four distracters and piloted these items with two college student samples. The results indicated that college students tend to be moderately confident in their responses regardless of the accuracy of the response. These findings imply that educating students on the purpose and process of accountability testing will require not only imparting correct information, but also debunking misconceptions.
Many universities rely on data gathered from tests that are low stakes for examinees but high stakes for the various programs being assessed. Given the lack of consequences associated with many collegiate assessments, the construct-irrelevant variance introduced by unmotivated students is potentially a serious threat to the validity of the inferences that institutions can make from their assessments. Two approaches to evaluating examinee motivation are discussed in this article: a global paper-and-pencil self-report measure of students' motivation across all tests completed during the course of a testing session, and a computer-based method that non-intrusively measures the amount of time students spend on each item in a test. This study presents evidence that the two motivation filtering methods provide similar filtered aggregate test scores, although more data was removed using the global paper-and-pencil self-report technique. Consequently, those interested in motivation filtering may not need to employ computer-based testing techniques but might instead effectively filter data from unmotivated students using self-report measures.
This study examined the psychometric properties of scores from the University Attachment Scale, a measure that operationalizes group and member attachment as two separate dimensions of attachment to a university. A two-factor model was championed over a one-factor model providing evidence of a distinction between university attachment and member attachment. Relationships with external criteria provided further support for this distinction and construct validity evidence. As predicted, "involved" students had practically and statistically significantly higher group attachment than "noninvolved" students. Furthermore, transfer students had practically and statistically significantly lower member attachment than nontransfer students. Additionally, there was a statistically significant positive relationship between students' perceived cohesion to the university and both group and member attachment. Overall, the authors believe that this is a promising new measure of university attachment.
Despite the best eff orts of measurement specialists and assessment practitioners to make valid inferences about student learning from the instruments they develop and administer, these inferences are predicated on two important student behaviors: the students must show up to take the test for the student population to be accurately represented, and the students must give eff ort when responding for their responses to be valid. Th ese two behaviors are cause for serious concern in low-stakes testing contexts. Testing companies, researchers who pull from voluntary subject pools, and university assessment programs across the country often rely heavily on low-stakes tests. Despite the importance of the test scores to the programs being assessed, there are few, if any, repercussions for students who choose not to attend these low-stakes test administrations (Sundre & Kitsantas, 2004). Who are these students who do not attend lowstakes testing sessions, why do they choose not to attend, and most important, what are the implications of not including their data in program assessment? Th ese questions have implications for student learning both at a localized level (such as assessing a university general education program) and at a national level (such as assessing No Child Left Behind, or nclb ).
The effectiveness of a postsecondary strategic learning course for improving metacognitive awareness and regulation was evaluated through systematic program assessment. The course emphasized students' awareness of personal learning through the study of learning theory and through practical application of specific learning strategies. Students assessed personal gains through pretest and posttest assessments of both metacognitive awareness and regulation. Pretest-to-posttest gains were statistically significant with large, meaningful effect sizes for program participants, including students with disabilities. Evidence supports the effectiveness of the program and, by extension, the value and importance of learning strategies instruction as a powerful educational intervention for students with disabilities.
General education program assessment involves low-stakes testing, but students may not be motivated to perform optimally if they know the test results will not represent them personally. We propose a protocol for administering general education tests under low-stakes conditions and describe simple proctor strategies that engender effort and inhibit inattention.
We studied the scores of students who avoided low-stakes testing and those who attended required testing. "Avoiders" who did exert effort on makeup tests scored similarly to students who initially attended testing. Practitioners should consider effort and include scores from avoiders to ensure that results reflect the entire student population.
Institutions can assess the impact of general education programs using direct measures of student knowledge and ability. Often a direct measure is a test for which scores indicate achievement of one or more of the general education learning outcomes. If students do well on the test, the general education instructional strategies that faculty employ at that institution appear to be eff ective. Alternatively, poor performance on the test serves as evidence that the general education program requires restructuring or needs to be strengthened. Either way, scores on direct measures of student ability can help institutions evaluate the strengths and weaknesses of their general education program. When administering a test to assess a general education program, academic leaders need to determine the appropriate assessment context and consider the impact that context will have on the resultant test data. For example, if a test is newly developed or being piloted for a new use, administering the test in a high-stakes context is usually not defensible. High-stakes contexts occur when resultant test scores are associated with signifi cant consequences for individuals. Th e Standards for Educational and Psychological Testing (American Educational Research Association, American Psychological Association, & National Council of