Measurement error is a common theme in classical measurement models used in testing and assessment. In classical measurement models, the definition of measurement error and the subsequent reliability coefficients differ on the basis of the test administration design. Internal consistency reliability specifies error due primarily to poor item sampling. Rater reliability indicates error due to inconsistency among raters. For estimates of test-retest reliability, error is attributed mainly to changes over time. In alternate-forms reliability, error is assumed to be due largely to variation between samples of items on test forms. Rasch models can also compute reliability estimates of scores under different test situations. The authors therefore present the Rasch perspective on calculating reliability (measurement error) and present Rasch measurement model programs to compute the various reliability estimates.
The Rasch measurement model using dichotomous scoring of item response data from a newly created Mobility Scale administered to elderly independent living individuals is presented. The dichotomous scoring model, item calibration, person calibration, logit scale, normative scale score, reliability, and validity are explained. Results indicated that additional activity statements need to be written and tested to improve the Mobility Scale instrument.
The study investigated five factors which can affect the equating of scores from two tests onto a common score scale. The five factors were: (a) item distribution type (i.e., normal versus uniform; (b) standard deviation of item difficulty (i.e.,.68,.95,.99); (c) number of items or test length (i.e., 50, 100, 200); (d) number of common items (i.e., 10, 20, 30); and (e) sample size (i.e., 100, 300, 500). SIMTEST and BIGSTEPS programs were used for the simulation and equating of 4,860 item data sets, respectively. Results from the five-way fixed effects factorial analysis of variance indicated three statistically significant two-way interaction effects. Simple effects for the interaction between common item length and test length only were interpreted given Type I error rate considerations. The eta-squared values for number of common items and test length were small indicating the effects had little practical importance. The Rasch approach to equating is robust with as few as 10 common items and a test length of 100 items.
Differential item functioning (DIF) detection rates were compared between logistic regression and analysis of variance for dichotomously scored items. These two DIF methods were compared using simulated binary item response data sets of varying test length (20, 40, and 60 items), sample size (200, 400, and 600 examinees), discrimination type (fixed and varying), and relative underlying ability (equal and unequal) between groups under conditions of uniform DIF, nonuniform DIF, combination DIF, and false positive errors. These test conditions were replicated 100 times. For both DIF detection methods, a test length of 20 items was sufficient for satisfactory DIF detection with detection rate increasing as sample size increased. With the exception of uniform DIF, the logistic regression method had higher mean detection rates than the analysis of variance method. Because the type of DIF present in real data is rarely known, the logistic regression method is recommended for most practical applications.
Many-facet Rasch analysis provides the bases for making fair and meaningful decisions from individual ratings by judges on tasks. The typical measurement design employed in a many-facet Rasch analysis has judges crossed with other facets or conditions of measurement. A nested design does not permit facets to be compared. However, a mixed design can be used to achieve a common vertical ruler when the frame of reference permits commensurate measures to be linked. Examples of crossed, nested, and mixed designs are presented to illustrate how a many-facet Rasch analysis can be modified to meet the connectivity requirement for comparing facet measures.
A Monte Carlo study was conducted using simulated dichotomous data to determine the effects of guessing on Rasch item fit statistics (weighted total, unweighted total, and unweighted between fit statistics) and the Logit Residual Index (LRI). The data were simulated using 100 items, 100 persons, three levels of guessing (0%, 25%, and 50%), and two item difficulty distributions (normal and uniform). The results of the study indicated that no significant differences were found between the mean Rasch item fit statistics for each distribution type as the probability of guessing the correct answer increased. The mean item scores differed significantly with uniformly distributed item difficulties, but not normally distributed item difficulties. The LRI was more sensitive to large positive item misfit values associated with the unweighted total fit statistic than to similar values associated with the weighted total fit or unweighted between fit statistics. The greatest magnitude of change in LRI values (negative) was observed when the unweighted total fit statistic had large positive values greater than 2.4. The LRI statistic was most useful in identifying the linear trend in the residuals for each item, thereby indicating differences in ability groups, i.e. differential item functioning.
Throughout the mid to late 1970's considerable research was conducted on the properties of Rasch fit mean squares. This work culminated in a variety of transformations to convert the mean squares into approximate t-statistics. This work was primarily motivated by the influence sample size has on the magnitude of the mean squares and the desire to have a single critical value that can generally be applied to most cases. In the late 1980's and the early 1990's the trend seems to have reversed, with numerous researchers using the untransformed fit mean squares as a means of testing fit to the Rasch measurement models. The principal motivation is cited as the influence sample size has on the sensitivity of the t-converted mean squares. The purpose of this paper is to present the historical development of these fit indices and the various transformations and to examine the impact of sample size on both the fit mean squares and the t-transformations of those mean squares. Because the sample size problem has little influence on the person mean square problem, due to the relatively short length (100 items or less), this paper focuses on the item fit mean squares, where it is common to find the statistics used with sample sizes ranging from 30 to 10,000.
As organizations begin to implement work teams, their assessment will ultimately reflect compensation strategies that move away from individual assessment. This will involve not only using multiple raters, but also the use of multiple criteria. Team assessment using multiple raters and multiple criteria is therefore necessitated; however, this can produce differences in ratings due to the leniency or severity of the individual team raters. This study analyzed the ratings of individual members on 31 different teams across 12 different criteria of team performance. Utilizing the many-facet Rasch model, statistical differences between the teams and 12 criteria were calculated.
Random-number generators are frequently used in Monte Carlo simulation studies to examine expected statistical outcomes. The purpose of this study was to investigate the normality of number distributions generated by various random-number generators. The research question focused on when the random-number generator reached a normal distribution (efficiency) and at what sample size. Findings suggest that when using a random-number generator in a Monte Carlo simulation study, the following steps should be observed: (a) normality of number distribution should be checked-means and standard deviations, (b) sample size should be large (n > 10,000), (c) starting seed values should be tested with regard to normality of resulting number distribution, (d) the length of periods before numbers repeat should be considered, (e) serial correlation between number sequences should be examined, and (f) uniformity of number distributions should be checked.
The different chi-square statistics reported in the many-faceted Rasch model analysis are presented and interpreted. In addition, other chi-square summary values are computed and presented for interpretation of facets. The chi-square values are useful for determining: (1) the significance of a facet in the Rasch model; (2) the significant contribution of facet main and interaction effects; (3) differences among facet elements; and (4) identifying the specific facet interaction adjustments to the subjects' calibrated logit ability measure.
The purpose of this study was to compare the results and interpretation of the data from a performance examination when four methods of analysis are used. Methods are 1) traditional summary statistics, 2) inter-judge correlations, 3) generalizability theory, and 4) the multi-facet Rasch model. Results indicated that similar sources of variance were identified using each method; however, the multi-facet Rasch model is the only method that linearized the scores and accounts for differences in the particular examination challenged by a candidate before ability estimates are calculated.