For several decades research into the student evaluation of teaching has periodically found an association between how well students like an instructor and the evaluations. The association has been largely ignored, being seen as an indicator of bias, or as a statistical or procedural artifact. However, these interpretations may be obscuring a more fundamental hypothesis. It is possible that the evaluations, instead of being a measure of 'good' or 'effective' teaching as commonly conceived, are actually a measure of a student-perceived construct similar to likability. This study looks directly at the influence of likability on the student evaluation of teaching. Knowing nothing about an instructor or how a class was taught, students' perception of likability accounted for two-thirds of the total variance of the evaluations. The student evaluation of teaching could be replaced with a single likability measure, with little loss of predictability.
Students were asked to rank instructors, who differed by age, gender and political leaning, by their expected helpfulness, and how much a student expected to learn. Students selected older instructors as those from whom they would learn the most, but chose young instructors as the most helpful. Overall, male instructors were preferred over female instructors, especially when emphasis was placed on learning. The political leaning of the instructor was a discriminating factor in humanities classes, with liberal instructors preferred over conservatives. The preferred age, gender and political leaning patterns were distinctly different for instructors who were helpful, and from those from whom students thought they would learn the most, indicating a dichotomy between perceived helpfulness and learning. The stereotypic images of instructors did not differ significantly by the students' own gender and academic major, except for male students ranking conservative instructors higher than females. Students do have stereotypical images of instructors based on the instructor's age, gender and political leaning.
The student evaluation of teaching process is generally thought to produce reliable results. The consistency is found within class and instructor averages, while a considerable amount of inconsistency exists with individual student responses. This paper reviews these issues along with a detailed examination of common measures of reliability that are utilised with the instruments. While inter-item consistency of the evaluations has been shown to be high, the agreement between students was shown to be no better than what would be expected by chance, indicating that students do not agree on what they are being asked to evaluate. The reliability measures generated by the student evaluations of teaching are an insufficient foundation for establishing validity. Further, the pattern of reliability indicates that the instruments are generally providing information about students, not instructors.
An exploratory study was conducted to see how types of validity influence students' perception of the fairness of exams. Face validity was essential for an exam to be perceived as fair. Content validity had the highest fairness rating and was least affected by other variables. In general, simpler concepts of validity had a greater impact on fairness than did more complex forms. Implications are discussed.
The research literature shows that students respond to both expected and deserved grades. The explanations for these effects, however, lead to opposite conclusions when the expected grades exceed the deserved grades. A grade/evaluation hypothesis suggests that the higher the expected grade, the higher the evaluation, while a fairness hypothesis suggests that any deviation of expected grades from deserved grades would lead to lower evaluations.