The primary purpose of this research is to examine the impact of estimation methods, actual latent trait distributions, and item pool characteristics on the performance of a simulated computerized adaptive testing (CAT) system. In this study, three estimation procedures are compared for accuracy of estimation: maximum likelihood estimation (MLE), expected a priori (EAP), and Warm's weighted likelihood estimation (WLE). Some research has shown that MLE and EAP perform equally well under certain conditions in polytomous CAT systems, such that they match the actual latent trait distribution. However, little research has compared these methods when prior estimates of. distributions are extremely poor. In general, it appears that MLE, EAP, and WLE procedures perform equally well when using an optimal item pool. However, the use of EAP procedures may be advantageous under nonoptimal testing conditions when the item pool is not appropriately matched to the examinees.
..............................................................1
The use of more performance items in large-scale testing has led to an increase in the research investigating the use of polytomously scored items in computer adaptive testing (CAT). Because this research has yet to be complimented with information pertaining to exposure control, the present research investigated the impact of using .ve different exposure control algorithms in two sized item pools calibrated using the generalized partial credit model. The results of the simulation study indicated that the a-stratified design, in comparison to a no-exposure control condition, could be used to reduce item exposure and overlap, increase pool utilization, and only minorly degrade measurement precision. Use of the more restrictive exposure control algorithms, such as the Sympson-Hetter and conditional Sympson-Hetter, controlled exposure to a greater extent but at the cost of measurement precision. Because convergence of the exposure control parameters was problematic for some of the more restrictive exposure control algorithms, use of the more simplistic exposure control mechanisms, particularly when the test length to item pool size ratio is large, is recommended.
An alternative to dichotomous scoring of multiple items anchored to a common stem is scoring these items as a single polytomous item (testlet scoring). This study systematically compared the partial credit model (PCM), the generalized partial credit model (GPCM), and the graded response model (GRM) in the context of testlet scoring. Data sets included a sample from the fall 1994 administration of the SAT I (N = 2,548) and a simulated data set. Theta estimation, information, and model fit were analyzed. Correlations among theta estimates ranged from 0.9748 to 0.9921. The relationship among the information functions of the PCM, GPCM and the GRM reflected the discrimination parameter estimates for the latter two models. Suggestions are made with regard to model selection.
The purpose of the present study was to investigate item and trait parameter recovery for Andrich's rating scale model using the PARSCALE computer program. The four factors upon which the simulated data matrices varied were (a) the distribution of the scale values for the items (skewed or uniform), (b) the number of category response options (4 or 5), (c) the distribution of known trait levels (normal or skewed), and (d) the sample size (60, 125, 250, 500, or 1,000). Each condition was replicated 10 times resulting in 400 data matrices. Accurate item and trait parameter estimates were obtained for all sample sizes examined. As expected, sample size seemed to have little influence on the recovery of trait parameters but did influence item parameter recovery. The distribution of known trait levels did not seriously impact the item parameter recovery. It was concluded that Andrich's rating scale model allows for the use of considerably smaller calibration samples than are typically recommended for other polytomous IRT models.
A simulation study was conducted to investigate the application of expected a posteriori (EAP) trait estimation in computerized adaptive tests (CAT) based on the partial credit model and compare it with maximum likelihood trait estimation (MLE). The performance of EAP was evaluated under different conditions: the number of quadrature points (10, 20,40, and 80) and the type of prior distribution (normal and uniform). The relative performance of MLE and the EAP estimation methods was assessed under two distributional forms of the latent trait (normal and negatively skewed). Results showed that, regardless of the latent trait distribution, MLE and EAP with a normal prior or a uniform prior using either 20, 40, or 80 quadrature points provided relatively accurate estimation in CAT based on the partial credit model. Also, increasing the number of quadrature points from 20 to 80 did not increase the accuracy of EAP estimation.
This study investigated parameter recovery for the partial credit model using the MULTILOG computer program. Factors studied were the sample size and the number of item parameters, which were manipulated by systematically varying the number of steps per item and the number of items. The findings suggest that the ratio of sample size to number of item parameters being estimated as a "rule of thumb" can be a more complete guideline when the number of steps per item is taken into account. Accurate estimation of ability can be obtained across all conditions, even with sample sizes as small as 250. With regard to estimation of step values, however, more caution is warranted. Accurate estimation of the step values of items which have more categories requires larger sample sizes for a given number of total parameters to be estimated.
A simulation study was conducted to investigate the effect of population distribution on maximum likelihood estimation (MLE) and expected a posteriori estimation (EAP) in computerized adaptive testing (CAT) based on Andrich's rating scale model. Comparisons were made among MLE and EAP with a normal prior distribution and EAP with a uniform prior distribution within two data sets: one generated using a normal trait distribution and the other using a negatively skewed trait distribution. Descriptive statistics, correlations, scattergrams, and accuracy indices were used to compare the different methods of trait estimation. EAP estimation with a normal prior or uniformprior yielded results similar to those obtained with MLE, even though the prior did not match the underlying trait distribution. An additional simulation study based on real data suggested that more work is needed to determine the optimal number of quadrature points for EAP in CAT based on the rating scale model. The choice between MLE and EAP for particular measurement situations is discussed.
In the present study, a procedure that has been used to select dichotomous items in computerized adaptive testing was applied to polytomous items. This procedure was designed to select the item with maximum weighted information. In a simulation study, the item information function was integrated over a fixed interval of ability values and the item with the maximum area was selected. This maximum interval information item selection procedure was compared to a maximum point information item selection procedure. Substantial differences between the two item selection procedures were not found when computerized adaptive tests were evaluated on bias and the root mean square of the ability estimate.
IRTINFO is a collection of SAS macros that compute item and test information for the graded response model (Samejima, 1969), the partial credit model (Masters, 1982), the generalized partial credit model (Muraki, 1992), the rating scale model (Andrich, 1978), the successive intervals model (Rost, 1988), and the three-parameter logistic model (~irnbaurn, 1968). Information is computed at each level of 0 in a range of 0 values specified by the user. The macros can handle items with differing numbers of categories when
Simulated data were used to investigate systematically the impact of various characteristics of the threshold values (number, symmetry, and distance between adjacent threshold values) and of the delta values on the distribution of item information in the successive intervals Rasch model. The results revealed that the shift in the peak of the information function away from the scale value of an item depended on the degree of asymmetry of the threshold values and the magnitude of the delta value of the item. The implications of the findings for computerized adaptive attitude measurement are discussed.
Simulated datasets were used to research the effects of the systematic variation of three major variables on the performance of computerized adaptive testing (CAT) procedures for the partial credit model. The three variables studied were the stopping rule for terminating the CATs, item pool size, and the distribution of the difficulty of the items in the pool. Results indicated that the standard error stopping rule performed better across the variety of CAT conditions than the minimum information stopping rule. In addition it was found that item pools that consisted of as few as 30 items were adequate for CAT provided that the item pool was of medium difficulty. The implications of these findings for implementing CAT systems based on the partial credit model are discussed.
The direct assessment of the writing ability of 2000 randomly selected secondary school students was performed using the partial credit model. The effects on item parameter estimates of rating scale and type of writing sample were investigated. The rating scales used were a sum and an averaging of the individual raters' ratings and the types of items were expository and narrative. Results showed that expository items tended to provide more information for high ability examinees than did narrative items and that the sum holistic rating method yielded greater information than did the average holistic rating scale. Additional analyses of the rating scales used as well as implications for test construction are discussed.
Individuals use a variety of drugs for a host of reasons and college students are no exception. Reasons provided for specific types of "recreational" drug use have included, but have not been limited to, medicinal, social, mood enhancement, and experimentation. In an attempt to discern the relationship between drug use and life satisfaction among college students, a slightly modified version of the National Institute on Drug Abuse (NIDA) Monitoring the Future Survey was administered to 683 students attending a major research university located in the southwestern United States. Based on the obtained study results a life satisfaction composite variable was created via factor analysis. Additionally, a polynomial multiple regression analysis was conducted to discern the association between the derived life satisfaction composite variable and drug use indices.
Real and simulated datasets were used to investigate the effects of the systematic variation of two major variables on the operating characteristics of computerized adaptive testing (CAT) applied to instruments consisting of poly- chotomously scored rating scale items. The two variables studied were the item selection procedure and the stepsize method used until maximum likelihood trait estimates could be calculated. The findings suggested that (1) item pools that consist of as few as 25 items may be adequate for CAT; (2) the variable stepsize method of preliminary trait estimation produced fewer cases of nonconvergence than the use of a fixed stepsize procedure; and (3) the scale value item selection procedure used in conjunction with a minimum standard error stopping rule outperformed the information item selection technique used in conjunction with a minimum information stopping rule in terms of the frequencies of nonconvergent cases, the number of items administered, and the correlations of CAT 0 estimates with full scale estimates and known 0 values. The implications of these findings for implementing CAT with rating scale items are discussed. Index terms:
Computerized adaptive testing (CAT) is a testing procedure that adapts an examination to an examinee's ability by administering only items of appropriate difficulty for the examinee. In this study, the authors compared Lord's flexilevel testing procedure (flexilevel CAT) with an item response theory‐based CAT using Bayesian estimation of ability (Bayesian CAT). Three flexilevel CATs, which differed in test length (36, 18, and 11 items), and three Bayesian CATs were simulated; the Bayesian CATs differed from one another in the standard error of estimate (SEE) used for terminating the test (0.25, 0.10, and 0.05). Results showed that the flexilevel 36‐ and 18‐item CATs produced ability estimates that may be considered as accurate as those of the Bayesian CAT with SEE = 0.10 and comparable to the Bayesian CAT with SEE = 0.05. The authors discuss the implications for classroom testing and for item response theory‐based CAT.
Huahua Chang (张华华)合作论文数Department of Educational Studies College of Education,Purdue University1