
Rasch/Guttman Scenario (RGS) scales produce scores that map onto interpretable descriptions of individuals at different levels of hierarchically progressive constructs. The unique scenario item format provides actionable and rich content-relevant feedback about (a) respondent status, (b) intervention design, and (c) longitudinal change on a construct. This article presents a seven-step methodological framework for the development of RGS scales. We also reflect on plausible challenges that may arise in the applications of RGS scale development and propose future research directions for the methodology.
Purpose: The purpose of this study was to explore the emotional work of diabetes during emerging adulthood and to explicate the validity of a newly developed measure of diabetes distress (DD) for use with emerging adults living with type 1 diabetes mellitus (T1DM), the Problem Areas in Diabetes—Emerging Adult version (PAID-EA). Methods: Young people ages 18 to 30 with T1DM were recruited online to complete a cross-sectional survey including measures of DD, depressive symptomology, and the PAID-EA. To evaluate content validity, 2 open-ended questions asked what was the most significant emotion or worry discussed in the survey items and what feelings were missed in those items. Responses were analyzed using directed qualitative content analysis. Results: A total of 254 (87%) participants responded to at least 1 of the 2 open-ended questions. Three themes and 1 subtheme were identified: (1) fear of the future with the subtheme of worry about the cost of diabetes, (2) acute worries about living with diabetes, and (3) challenges with finding support. More PAID-EA items corresponded with these themes than items on the original Problem Areas in Diabetes or Center for Epidemiologic Studies Depression Scale, supporting the validity of the PAID-EA and clarifying the developmental-stage-specific aspects of DD. Conclusions: Emerging adulthood is a period in which the future should hold infinite possibility, but young people with T1DM describe a staggering fear of the future with markedly limited possibilities, supporting the need to measure the developmental-stage-specific experience of DD as captured on the PAID-EA.
Mental toughness (MT) predicts outcomes across several high-stress contexts such as athletics, the military, and the workplace. Despite this, researchers have struggled to reach consensus regarding how best to conceptualize and measure MT. MT assessments have focused on measuring general MT rather than domain-specific MT. The current study proposed a measurement model of MT grounded in social-cognitive theory and introduced an assessment of MT within a situational judgment context relevant to the workplace. Participants completed the new MT measure as well as assessments to establish construct validity. Both exploratory and confirmatory factor analyses suggested a three-factor solution fit the data best, consisting of task persistence, emotional control, and utilization of feedback. Cross-structure analyses indicated that the new assessment avoided common-method bias in responding, evidenced by weak correlations with measures of other constructs. The results provided initial evidence to continue research on using a situational judgment test to measure MT.
Accurate parameter estimation in the Rasch model involves the assumption of conditional independence, also termed local independence. Conditional on ability, the responses to items A and B should be independent. Two types of conditional dependence are detailed in this pedagogical piece: trait dependency and response dependency. The bias in difficulty and reliability and the estimates of fit and correlated residuals resulting from these dependencies are compared and contrasted to results from using models that account for the dependency. Contrasts with results from a 2-parameter item response theory model are also briefly noted.
Streiner, Norman and Cairney (2015) "Health Measurement Scales: A practical guide to their development and use", now in its fifth edition, is one of the foundational texts of the health outcomes movement. It states that "the differences between scales constructed with IRT and CTT are trivial." (Streiner, Norman and Cairney, 2015, p. 299) This statement is representative of the view which emphasizes the equivalence of True-Score Theory (TST) (also known as Classical Test Theory [CTT]) and the Rasch Measurement Model [RMM]). This view is widely held and has been one factor in limiting the application of RMM in the development of health outcome measures. However, this equivalence view relies heavily on a paper by Fan (1998) which examined the item statistics derived from TST, IRT (Item Response Theory) and the RMM for a large educational dataset. While subject to a number of theoretical and practical criticisms from a RMM perspective this paper has not been replicated with a large sample. This paper by replicating and extending the paper by Fan (1998) challenges the finding that item difficulty indexes derived from high and low ability samples using TST techniques are invariant. They are not. On the other hand, item locations derived from the RMM have a high degree of invariance. This secondary data analysis, by working through the methods used by Fan (1998) also demonstrates that a reliance on the magnitude of correlational coefficients cannot be used to determine the invariance of item difficulty indexes. An investigation into the linearity of the correlations using scatter plots is also required. Finally, an item analysis derived from the item difficulty indexes which displays a picture of the test as a whole shows that, for this large sample, the differences between scales constructed with TST and the RMM are not trivial.
Residual-based fit statistics are among the most common indicators of fit to the Rasch model. There is considerable discussion in the literature of the efficacy of item fit statistics in detecting measurement disturbances. However, to date there has been no investigation of whether these fit statistics are robust to interactions between item discrimination and item difficulty. This study uses simulations to investigate whether interaction effects occur for fit statistics commonly used with the Rasch model. It is found that when the parameters are estimated with the Rasch model, the values of certain item fit statistics vary depending on the interaction between location and discrimination. Specifically, in the study, OUTFIT MNSQ and INFIT MNSQ provide a relatively consistent index of item discrimination across a range of item difficulties, whereas the t-statistic and the log transformed fit residual vary in a systematic fashion that depends on item location.
In this study we investigate whether transformations between different representations of mathematical objects constitute a suitable framework for the assessment of students' comprehension of fraction addition. Participants (N = 164) solved a set of 20 fraction addition problems constructed on the basis of Duval's (2017) theory of the role of representational transformations in mathematical comprehension. Using Rasch measurement theory and principal component analysis, we found that the items could be separated into three levels of difficulty based on the transformation involved. This large-scale structure was consistent across gender and across subgroups of preservice teachers and middle-grade students. On a finer scale, the production of diagrammatic representations, and the type of diagrammatic representation involved, constitute potential subdimensions of the instrument. We conclude that transformations between representations can be productive for the assessment of fraction addition comprehension as long as care is taken to curtail the potential effects of multidimensionality.
The present study developed and validated a short form of the Cross-Cultural (Chinese) Personality Assessment Inventory for adolescents (CPAI-A; Form B) focusing on the personality scales by unidimensional and multidimensional Rasch models. Multiple evidence from unidimensional Rasch models (item fit, DIF statistics, dimensionality, reliability indices, construct coverage) were evaluated in order to create a short scale with optimal psychometric properties. Further, multidimensional Rasch model, canonical analysis, and predictive validity were performed and evaluated to validate the CPAI-A-SF further. As a result, 65 of 277 items were selected in the short measure with a four-dimensional structure. The infit and outfit mean-squares (MNSQ) of the personality scale items ranged between .81 and 1.25. Good construct coverage was displayed on the item-person map, and all four dimensions demonstrate reasonable EAP/PV reliability ranging from .81 to .87. The personality scores of CPAI-A-SF predicted life satisfaction as well as the scores from the original inventory.
To understand the role of fit statistics in Rasch measurement is simple: applied researchers can only benefit from the desirable properties of the Rasch model when the data fit the model. The purpose of the current study was to assess the Q-Index robustness (Ostini and Nering, 2006), and its performance was compared to the current popular fit statistics known as MSQ Infit, MSQ Outfit, and standardized Infit and Outfit (ZSTDs) under varying conditions of test length, sample size, item difficulty (normal and uniform), and dimensionality utilizing a Monte Carlo simulation. The Type I and Type II error rates are also examined across fit indices. This study provides applied researchers guidelines the robustness and appropriateness of the use of the Q-Index, which is an alternative to the currently available item fit statistics. The Q-Index was slightly more sensitive to the levels of multidimensionality set in the study while MSQ Infit, Outfit, and standardized Infit and Outfit (ZSTDs) failed to identify the multidimensional conditions. The Type I error rate of the Q-Index was lower than the rest of the fit indices; however, the Type II error rate was higher than the anticipated beta = .20 across all fit indices.
Many assessment scales in the social sciences are composed of multiple items that form a subscale structure. They have this structure because more than one aspect of the variable is assessed and more than one item assesses each aspect. Nevertheless, generally, a single measurement is required from the scale. A characteristic of this measurement is that the greater the number of items, and categories within an item, that assess an aspect, the greater its influence on the final measurement. One way to control this influence is to include the desired relative number of items and categories to assess each aspect in the scale. However, there are circumstances where designing the required number of items and categories for each aspect is challenging. This paper shows a method of controlling the influence of the number of items and categories assessing each aspect by a-priori weighting of items at the person measurement stage with the Rasch model.
Most research on multistage testing (MST) uses simulated data. This study adds to the literature by using both operational test data and simulated data to compare two different MST designs with regard to proficiency estimation accuracy and module exposure rates and by investigating whether simulation studies and operational test studies yield similar results. Two MST designs (1-2 and 1-3-4 designs) from one state's sixth-grade summative mathematics assessment across two years were compared in this study. Both simulation and operational test studies demonstrate similar results: the two MST designs yield no significant performance differences with regard to estimation accuracy and module exposure. These results provide evidence that simulation studies can provide adequate results to inform decisions about MST designs.
Researchers and practitioners have used the Modern Language Aptitude Test (MLAT) to assess language aptitude and identify possible language learning deficiencies in examinees since the 1950s. However, researchers have not assessed its psychometric properties using modern measurement theory methods. We use the dichotomous Rasch model to explore the psychometric properties of the MLAT, including data-model fit indices, item difficulty and student ability calibrations, reliability of separation, and differences in achievement across gender subgroups based on a sample of undergraduate and graduate university students (N=204). Our findings suggest that the MLAT has acceptable psychometric properties such that it can be meaningfully interpreted as a measure of language proficiency. Our findings confirm previous research that language performance across gender groups significantly differs. We found no significant interactions between gender subgroups and the difficulty of the five domains of the assessment. We discuss these results in terms of their implications for research and practice.
Research using the National Teacher and Principal Survey (NTPS) has consistently demonstrated that teachers' reported working conditions are related to both intentions to leave the profession and attrition (Tickle, Chang, and Kim, 2011). However, limited research evaluates teacher appraisals of job-related demands and resources as an antecedent to job dissatisfaction. We tested for differential item functioning (DIF) using a partial credit model approach within a Rasch modeling context to examine whether elementary and secondary teachers with similar overall stress levels respond to the NTPS Demands and Resources items in similar ways. For the Demands items, seven of the items displayed differences that were negligible, four were intermediate, and three items indicated large DIF contrasts. For the Resources items, 10 items displayed differences that were negligible, two were intermediate, and zero items indicated large DIF contrasts. These results indicate elementary and secondary teachers exhibit different appraisal patterns, suggesting implications for the development and use of survey data in public school settings in general, and for the use of the NTPS data in particular.
The purpose of this study is to demonstrate an application of the many-facet Rasch model (MFRM) in evaluating the impromptu speech skills of pre-service principals in Taiwan. The findings showed that the topics of speech did not exhibit different difficulty measures. With respect to scoring criteria, time control was the most difficult aspect among the scoring criteria. Regarding gender difference in raters, female raters gave lower scores than male raters, but there was no statistical evidence for gender-related bias. However, raters exhibited statistically significant differences in rater severity. The results of this study demonstrates that the MFRM provides a scientific approach to assessment, which can reveal some useful diagnostic information from the original ordinal rating scores on impromptu speech.
Because modern, simultaneously estimated longitudinal Rasch models are unable to handle many timepoints, new methods of producing person and item estimates and evaluating test function are necessary. Longitudinal anchoring, in which a common scale of item parameters is used to estimate trait levels over multiple occasions, is a potential solution. With proper anchoring procedures, person and item estimates can be obtained without limiting the number of timepoints that can be analyzed. A simulation study examining the performance of six longitudinal anchoring methods (Floated, Racked, Time One, Mean, Random, and Stacked) was conducted. The Mean and the Stacked anchoring methods best recovered the population change over time, person and item estimates, and model fit. The Racked method could not produce reliable change estimates and should be avoided. Longitudinal anchoring is an easily implemented solution when analyzing large longitudinal datasets and shows promise as a low-computation method of producing latent trait estimates.
While several proofs exist that the number keyed (or number correct) score is a sufficient statistic to estimate person measure (or ability, beta) in the dichotomous Rasch model, there are few proofs about the direct mathematical link from beta to the number correct score. This manuscript proves that the estimation link going from score to beta is the test response function, which is the sum of the probabilities correct for all items given the difficulty (delta) values and beta.
Multidimensional pairwise comparison (MPC) items have been widely used to assess career interest, value and personality to avoid response bias in educational sectors. In reality, a statement in an MPC item may have different utilities for different groups, which is referred to as differential statement functioning (DSF). Few studies have been investigated DSF assessment. Based on a Rasch model for MPC items, this study adapts three methods to detect DSF for polytomous MPC items: the equal-mean-utility (EMU) method, the all-other-statement (AOS) method and the constant-statement (CS) method. Simulation study was conducted to evaluate the recovery of parameters as well as the performance of the proposed methods. Results showed that when the test contains DSF statement(s), the CS method where one or more DSF-free statements are chosen as an anchor will yield accurate estimates and perform well for DSF assessment. An empirical example of career interest assessment was provided. .
In previous studies, researchers have focused on the development and interpretation of measurement tools related to self-efficacy. However, researchers have seldom investigated whether these instruments demonstrate acceptable psychometric properties, including similar item interpretations between subgroups of respondents. The purpose of this study was to explore the extent to which a self-efficacy measure has a consistent interpretation for two self-reported gender subgroups. The researchers utilized Rasch analysis to explore differences in item difficulty between the subgroups. Results suggested differences in item difficulty ordering for certain self-efficacy items. Implications for research and practice are discussed.
The estimates of intraclass correlations are known to be biased, but there are few analytical ways to assess the amount of bias. The analytical approach requires the normality assumption to estimate bias. Bootstrap requires no such assumption and can, therefore, be used to estimate bias, regardless of the model assumption. We utilize cluster bootstrapping to calculate the bias in estimating the intraclass correlation. A well-known dataset is provided to illustrate the bias estimation in a typical study design of intraclass correlation, and its implications for other study designs are also discussed.