: Efforts to improve the quality of education in the United States increasingly emphasize the need for high-stakes achievement testing. The availability of valid, reliable, and cost-effective measures of achievement is critical to the success of many reform efforts. This paper begins with a brief discussion of the context surrounding large-scale (e.g., statewide) achievement testing. We then describe a new approach to assessment that we believe holds promise for reshaping the way achievement is measured. This approach uses tests that are delivered to students over the Internet and are tailored (adapted) to each student's own level of proficiency. We anticipate that this paper will be of interest to policy makers, educators, and test developers who are charged with improving the measurement of student achievement. We are not advocating the wholesale replacement of all current paper-and-pencil measures with web-based testing. However, we believe that current trends toward greater use of high-stakes tests and the increasing presence of technology in the classroom will lead assessment in this direction. Indeed, systems similar to those we describe in this report are already operational in several U.S. school districts and in other countries. Furthermore, we believe that although web-based testing holds promise for improving the way achievement is measured, a number of factors may limit its usefulness or potentially lead to undesirable outcomes. It is therefore imperative that the benefits and limitations of this form of testing be explored and the potential consequences be understood. The purpose of this paper is to stimulate discussion and research that will address the many issues raised by a shift toward web-based testing.
In his 1997 book, King announced “A Solution to the Ecological Inference Problem”. This review discusses King’s method, and tests it on data where truth is known. In the test data, his method produces results that are far from truth, and diagnostics are unreliable. Ecological regression makes estimates that are similar to King’s, while the neighborhood model is more accurate. His announcement is premature.
The Collegiate Learning Assessment (CLA) program measures value added in colleges and universities, by testing the ability of freshmen and seniors to think logically and write clearly. The program is popular enough that it has attracted critics. In this paper, we outline the methods used by the CLA to determine value added. We summarize the criticisms, which revolve around the question of which students take the CLA tests. Typically, samples are not random, so that selection bias is a concern, as is confounding. We respond by showing that criticisms of CLA procedures are not supported by the data.
Assessment of learning in higher education is a critical concern to policy makers, educators, parents, and students. And, doing so appropriately is likely to require including constructed response tests in the assessment system. We examined whether scoring costs and other concerns with using open-end measures on a large scale (e.g., turnaround time and inter-reader consistency) could be addressed by machine grading the answers. Analyses with 1359 students from 14 colleges found that two human readers agreed highly with each other in the scores they assigned to the answers to three types of open-ended questions. These reader assigned scores also agreed highly with those assigned by a computer. The correlations of the machine-assigned scores with SAT scores, college grades, and other measures were comparable to the correlations of these variables with the hand-assigned scores. Machine scoring did not widen differences in mean scores between racial/ethnic or gender groups. Our findings demonstrated that machine scoring can facilitate the use of open-ended questions in large-scale testing programs by providing a fast, accurate, and economical way to grade responses.
can be found at: Evaluation Review Additional services and information for http://erx.sagepub.com/cgi/alerts Email Alerts: http://erx.sagepub.com/subscriptions Subscriptions: http://www.sagepub.com/journalsReprints.nav Reprints: http://www.sagepub.com/journalsPermissions.nav Permissions: http://erx.sagepub.com/cgi/content/refs/31/5/415 SAGE Journals Online and HighWire Press platforms): (this article cites 5 articles hosted on the Citations
The Collegiate Learning Assessment (CLA) is a computer administered, open-ended (as opposed to multiple-choice) test of analytic reasoning, critical thinking, problem solving, and written communication skills. Because the CLA has been endorsed by several national higher education commissions, it has come under intense scrutiny by faculty members, college administrators, testing experts, legislators, and others. This article describes the CLA's measures and what they do and do not assess, how dependably they measure what they claim to measure, and how CLA scores differ from those on other direct and indirect measures of college student learning. For instance, analyses are conducted at the school rather than the student level and results are adjusted for input to assess whether the progress students are making at their school is better or worse than what would be expected given the progress of "similarly situated" students (in terms of incoming ability) at other colleges.
The Collegiate Learning Assessment (CLA) is a computer administered, open-ended (as opposed to multiple-choice) test of analytic reasoning, critical thinking, problem solving, and written communication skills. Because the CLA has been endorsed by several national higher education commissions, it has come under intense scrutiny by faculty members, college administrators, testing experts, legislators, and others. This article describes the CLA's measures and what they do and do not assess, how dependably they measure what they claim to measure, and how CLA scores differ from those on other direct and indirect measures of college student learning. For instance, analyses are conducted at the school rather than the student level and results are adjusted for input to assess whether the progress students are making at their school is better or worse than what would be expected given the progress of “similarly situated” students (in terms of incoming ability) at other colleges.
There is confusion about what kind of assessment is appropriate in higher education. There is also misunderstanding over the relationship between assessment and accountability. Even when institutions and state policymakers use similar assessment information they do so for different purposes. Faculty are focused on improving educational programs while state leaders are interested in holding their public higher education institutions accountable for their performance.
A three-year study finds that students who had been exposed to more reform-oriented teaching performed slightly better in both math and science than those who had experienced less. Evidence also suggests that reform-oriented teaching may help improve students' problem-solving skills.
Presents the findings of a multiyear study of the effectiveness of reform-oriented science and mathematics teaching (instructional practices for engaging students actively in their own learning and enhancing the development of complex cognitive skills) — specifically, whether such practices are associated with higher student achievement and whether that association is sensitive to the aspects of achievement that are measured. (CD-ROM enclosed.)
This study examines (1) the extent to which student engagement is associated with experimental and traditional measures of academic performance, (2) whether the relationships between engagement and academic performance are conditional, and (3) whether institutions differ in terms of their ability to convert student engagement into academic performance. The sample consisted of 1058 students at 14 four-year colleges and universities that completed several instruments during 2002. Many measures of student engagement were linked positively with such desirable learning outcomes as critical thinking and grades, although most of the relationships were weak in strength. The results suggest that the lowest-ability students benefit more from engagement than classmates, first-year students and seniors convert different forms of engagement into academic achievement, and certain institutions more effectively convert student engagement into higher performance on critical thinking tests.
In a three-year study, RAND researchers examined the relationship between reform-oriented instruction and student performance in mathematics and science. At the end of the study, students who had been exposed to more reform-oriented teaching performed better in both math and science than those who had experienced less, but the differences in scores were small. The relationship between instructional practice and student performance was stronger when performance was measured on items that required problem-solving skills rather than procedural skills, an outcome that is generally consistent with the goals of reform-oriented teaching. These results illustrate the importance of matching performance measures to reform goals in evaluating instructional innovations. In science, students might be expected to generate scientifi c questions, plan research, collect data, analyze relationships, and write reports. More generally, reform-oriented instruction encourages inquiry-based activities and the skills of intellectual conversation: asking questions, discussing alternative approaches to problems, presenting reasons for answers to questions, and making connections between old and new knowledge and between superfi cially disparate topics. Th e study of reform-oriented instruction described here, called Mosaic II, was an extension of Mosaic I, an earlier investigation in which a team of RAND researchers found “a weak but positive relationship” between reform-oriented instructional practices and students’ scores on standardized tests. Intrigued by these results, they undertook a new study focused on the same question, but designed to yield a more detailed picture of this relationship. Th e new study involved three key methodological features: a longitudinal research design, multiple measures of instructional practices, and multiple measures of student performance. Student performance on standardized tests was measured over a three-year period At the outset of the study, the RAND team selected fi ve cohorts of elementary and middle-school students in three school districts. A cohort consisted of all the students in a selected grade in one of the two subject areas (for example, third-grade mathematics students or sixth-grade science students).1 Each cohort was followed for three years, and their performance on standardized tests was measured at the end of each year (current-year performance), as well as aggregated at the end of the three-year period (cumulative performance). Multiple measures were used to assess instructional practice To obtain data regarding teachers’ choices and classroom activities as they occur, the Mosaic II team used the following methods to assess instructional practices: • Teacher Surveys: Th e surveys gathered information about teachers’ educational background, experience, and classroom practices. • Classroom Vignettes: Teachers were presented with written vignettes, consisting of hypothetical classroom situations, and asked to indicate how likely it was that they would choose each of a list of possible actions that could be taken in response. Th e actions refl ected a range of reform-oriented and non-reform-oriented alternatives. Unlike the survey, which required teachers to recall their actions over a wide range of situations over a long period, this method provided teachers with a realistic context in which to situate their responses, potentially off ering insights into their preferred teaching style and increasing the generalizability of their responses. • Teacher Logs: Each day during a twoto fi ve-day period, teachers were asked to fi ll out a log by checking off items on a list describing specifi c activities that occurred during mathematics or science lessons. Th e logs focused on both teacher and student activities each day during the data-collection period. Th e immediacy of this procedure was intended to increase accuracy in reporting of classroom activities. • Classroom Observations: Th e RAND team observed 52 teachers, recording behaviors on a protocol designed to capture important elements of reform-oriented instruction. Because the results of these observations were independent of teacher reports, they were free of any biases due to forgetting or other reporting limitations. Both multiple-choice and open-ended measures were used to assess student performance on standardized tests Two measures—scores on both multiple-choice and openended items on standardized tests—were used as indicators of student learning in mathematics and science. In addition, performance on items that require problem-solving skills was diff erentiated from performance on items that required only the application of procedural skills.
Over the past decade, state legislatures have experienced increasing pressure to hold higher education accountable for student learning. This pressure stems from several sources, such as increasing costs and decreasing graduation rates. To explore the feasibility of one approach to measuring student learning that emphasizes program improvement, we administered several open-ended tests to 1365 students from 14 diverse colleges. The strong correspondence between hand and computer assigned scores indicates the tests can be administered and graded cost effectively on a large scale. The scores were highly reliable, especially when the college is the unit of analysis; they were sensitive to years in college; and they correlated highly with college GPAs. We also found evidence of “value added” in that scores were significantly higher at some schools than at others after controlling on the school’s mean SAT score. Finally, the students said the tasks were interesting and engaging.
Wherein the authors reply to Toenjes, Vol. 13 No. 36 "What do Klein et al. tell us about test scores in Texas?"