Multidimensional forced choice (MFC) test formats are commonly used as an alternative to traditional rating scale formats to reduce aberrant responding, especially faking in high-stakes settings. However, MFC remains susceptible to random responding, particularly in low-stakes settings where respondents may be insufficiently motivated and in high-stakes settings where some assessments may be viewed as less consequential. To ensure the validity of inferences drawn from MFC data, effective methods for detecting random responding are needed. This research contributes to the MFC literature on aberrant responding detection by evaluating the effectiveness of the item response theory (IRT)-based person fit statistic l z for detecting random responding in multi-unidimensional pairwise preference (MUPP)-based MFC tests, using optimal appropriateness measurement (OAM) as a theoretical benchmark. Two simulation studies were conducted. Study 1 compared l z with OAM, and Study 2 examined l z in a broader simulation design. Results showed that (1) higher proportions of randomly answered items, longer tests, and the use of empirical critical values were associated with greater detection power for l z , (2) the proportion of aberrant respondents did not affect l z performance, and (3) OAM outperformed l z only when the random responding model was correctly specified, a condition that can be realized in simulation but may not hold in applied testing contexts. Overall, the findings support the use of l z with empirical critical values as a practical method for detecting random responding in MUPP-based MFC tests.
The field of psychometrics has made remarkable progress in developing item response theory (IRT) models for analyzing multidimensional forced choice (MFC) measures. This study introduces an innovative method that enhances the latent trait estimation of the Multi-Unidimensional Pairwise Preference (MUPP) model by incorporating latent regression modeling. To validate the efficacy of the new method, we conducted a comprehensive simulation study. The results of the study provide compelling evidence that the proposed latent regression MUPP (LR-MUPP) model significantly improves the accuracy of the latent trait estimation. This study opens new avenues for future research and encourages further development and refinement of MFC IRT models and their applications.
Multidimensional forced-choice (MFC) testing has been proposed as an alternative to single-statement (SS) Likert-type measures to reduce response biases in noncognitive measurement. Research progress has been made on MFC computerized adaptive testing (CAT) to improve testing efficiency. CAT enhances efficiency by successively selecting items that are most informative at each respondent’s trait estimate. In MFC CAT, this causes some forced-choice items and the statements composing them to be frequently exposed while others are rarely used, which adversely affects test security and costs. This research developed an exposure control method for MFC CAT based on the multi-unidimensional pairwise preference model (MUPP; Stark et al. Applied Psychological Measurement, 29,184–203, 2005). Because the method was intended to prevent the overuse of the most informative items and statements, it tended to decrease overall measurement accuracy and precision. Thus, a second purpose of this research was to examine the extent to which these losses in accuracy and precision might be offset by incorporating collateral information. The effectiveness of the exposure control method and the incorporation of collateral information in MFC CAT were investigated in a Monte Carlo study that also manipulated test length and the correlation between dimensions. A byproduct of this research was an MFC CAT algorithm that improves test security and cost-effectiveness, while simultaneously maintaining measurement accuracy and precision of noncognitive constructs.
Multidimensional forced choice (MFC) measures are gaining prominence in noncognitive assessment. Yet there has been little research on detecting differential item functioning (DIF) with models for forced choice measures. This research extended two well-known DIF detection methods to MFC measures. Specifically, the performance of Lord's chi-square and item parameter replication (IPR) methods with MFC tests based on the Multi-Unidimensional Pairwise Preference (MUPP) model was investigated. The Type I error rate and power of the DIF detection methods were examined in a Monte Carlo simulation that manipulated sample size, impact, DIF source, and DIF magnitude. Both methods showed consistent power and were found to control Type I error well across study conditions, indicating that established approaches to DIF detection work well with the MUPP model. Lord's chi-square outperformed the IPR method when DIF source was statement discrimination while the opposite was true when DIF source was statement threshold. Also, both methods performed similarly and showed better power when DIF source was statement location, in line with previous research. Study implications and practical recommendations for DIF detection with MFC tests, as well as limitations, are discussed.
Multidimensional forced choice (MFC) formats have emerged as a promising alternative to traditional single statement Likert-type measures for assessing noncognitive traits while reducing response biases. As MFC formats become more widely used, there is a growing need for tools to support MFC analysis, which motivated the development of the fcirt package. The fcirt package estimates forced choice model parameters using Bayesian methods. It currently enables estimation of the Generalized Graded Unfolding Model (GGUM; Roberts et al., 2000)-based Multi-Unidimensional Pairwise Preference (MUPP) model using rstan, which implements the Hamiltonian Monte Carlo (HMC) sampling algorithm. fcirt also includes functions for computing item and test information functions to evaluate the quality of MFC assessments, as well as functions for Bayesian diagnostic plotting to assist with model evaluation and convergence assessment.
Applications of multidimensional forced choice (MFC) testing have increased considerably over the last 20 years. Yet there has been little, if any, research on methods for linking the parameter estimates from different samples. This research addressed that important need by extending four widely used methods for unidimensional linking and comparing the efficacy of new estimation algorithms for MFC linking coefficients based on the Multi-Unidimensional Pairwise Preference model (MUPP). More specifically, we compared the efficacy of multidimensional test characteristic curve (TCC), item characteristic curve (ICC; Haebara, 1980), mean/mean (M/M), and mean/sigma (M/S) methods in a Monte Carlo study that also manipulated test length, test dimensionality, sample size, percentage of anchor items, and linking scenarios. Results indicated that the ICC method outperformed the M/M method, which was better than the M/S method, with the TCC method being the least effective. However, as the number of items “per dimension” and the percentage of anchor items increased, the differences between the ICC, M/M, and M/S methods decreased. Study implications and practical recommendations for MUPP linking, as well as limitations, are discussed.
OBJECTIVE:Electronic cigarettes are the most commonly used tobacco products by young adults. Measures of beliefs about outcomes of use (i.e., expectancies) can be helpful in predicting use, as well as informing and evaluating interventions to impact use. METHODS:We surveyed young adult students (N = 2296, Mean age=20.0, SD=1.8, 64 % female, 34 % White) from a community college, a historically black university, and a state university. Students answered ENDS expectancy items derived from focus groups and expert panel refinement using Delphi methods. Factor Analysis and Item Response Theory (IRT) methods were used to understand relevant factors and identify useful items. RESULTS:A 5-factor solution [Positive Reinforcement (consists of Stimulation, Sensorimotor, and Taste subthemes, α = .92), Negative Consequences (Health Risks and Stigma, α = .94), Negative Affect Reduction (α = .95), Weight Control (α = .92), and Addiction (α = .87)] fit the data well (CFI=0.95; TLI=0.94; RMSEA=0.05) and was invariant across subgroups. Factors were significantly correlated with relevant vaping measures, including vaping susceptibility and lifetime vaping. Hierarchical linear regression demonstrated factors were significant predictors of lifetime vaping after controlling for demographics, vaping ad exposure, and peer/family vaping. IRT analyses indicated that individual items tended to be related to their underlying constructs (a parameters ranged from 1.26 to 3.18) and covered a relatively wide range of the expectancies continuum (b parameters ranged from -0.72 to 2.47). CONCLUSIONS:A novel ENDS expectancy measure appears to be a reliable measure for young adults with promising results in the domains of concurrent validity, incremental validity, and IRT characteristics. This tool may be helpful in predicting use and informing future interventions. IMPLICATIONS:Findings provide support for the future development of computerized adaptive testing of vaping beliefs. Expectancies appear to play a role in vaping similar to smoking and other substance use. Public health messaging should target expectancies to modify young adult vaping behavior.
Although the use of ideal point item response theory (IRT) models for organizational research has increased over the last decade, the assessment of construct dimensionality of ideal point scales has been overlooked in previous research. In this study, we developed and evaluated dimensionality assessment methods for an ideal point IRT model under the Bayesian framework. We applied the posterior predictive model checking (PPMC) approach to the most widely used ideal point IRT model, the generalized graded unfolding model (GGUM). We conducted a Monte Carlo simulation to compare the performance of item pair discrepancy statistics and to evaluate the Type I error and power rates of the methods. The simulation results indicated that the Bayesian dimensionality detection method controlled Type I errors reasonably well across the conditions. In addition, the proposed method showed better performance than existing methods, yielding acceptable power when 20% of the items were generated from the secondary dimension. Organizational implications and limitations of the study are further discussed.
Multidimensional forced-choice (MFC) testing has been proposed as a way of reducing response biases in noncognitive measurement. Although early item response theory (IRT) research focused on illustrating that person parameter estimates with normative properties could be obtained using various MFC models and formats, more recent attention has been devoted to exploring the processes involved in test construction and how that influences MFC scores. This research compared two approaches for estimating multi-unidimensional pairwise preference model (MUPP; Stark et al., 2005) parameters based on the generalized graded unfolding model (GGUM; Roberts et al., 2000). More specifically, we compared the efficacy of statement and person parameter estimation based on a "two-step" process, developed by Stark et al. (2005), with a more recently developed "direct" estimation approach (Lee et al., 2019) in a Monte Carlo study that also manipulated test length, test dimensionality, sample size, and the correlations between generating person parameters for each dimension. Results indicated that the two approaches had similar scoring accuracy, although the two-step approach had better statement parameter recovery than the direct approach. Limitations, implications for MFC test construction and scoring, and recommendations for future MFC research and practice are discussed.
Differential item functioning (DIF) analysis is one of the most important applications of item response theory (IRT) in psychological assessment. This study examined the performance of two Bayesian DIF methods, Bayes factor (BF) and deviance information criterion (DIC), with the generalized graded unfolding model (GGUM). The Type I error and power were investigated in a Monte Carlo simulation that manipulated sample size, DIF source, DIF size, DIF location, subpopulation trait distribution, and type of baseline model. We also examined the performance of two likelihood-based methods, the likelihood ratio (LR) test and Akaike information criterion (AIC), using marginal maximum likelihood (MML) estimation for comparison with past DIF research. The results indicated that the proposed BF and DIC methods provided well-controlled Type I error and high power using a free-baseline model implementation, their performance was superior to LR and AIC in terms of Type I error rates when the reference and focal group trait distributions differed. The implications and recommendations for applied research are discussed.
Collateral information has been used to address subpopulation heterogeneity and increase estimation accuracy in some large-scale cognitive assessments. The methodology that takes collateral information into account has not been developed and explored in published research with models designed specifically for noncognitive measurement. Because the accurate noncognitive measurement is becoming increasingly important, we sought to examine the benefits of using collateral information in latent trait estimation with an item response theory model that has proven valuable for noncognitive testing, namely, the generalized graded unfolding model (GGUM). Our presentation introduces an extension of the GGUM that incorporates collateral information, henceforth called Explanatory GGUM. We then present a simulation study that examined Explanatory GGUM latent trait estimation as a function of sample size, test length, number of background covariates, and correlation between the covariates and the latent trait. Results indicated the Explanatory GGUM approach provides scoring accuracy and precision superior to traditional expected a posteriori (EAP) and full Bayesian (FB) methods. Implications and recommendations are discussed.
Multilevel latent class analysis (MLCA) has been increasingly used to investigate unobserved population heterogeneity while taking into account data dependency. Nonparametric MLCA has gained much popularity due to the advantage of classifying both individuals and clusters into latent classes. This study demonstrated the need to relax the assumption in specifying the nonparametric MLCA: item response probabilities varied only across level-1 latent classes, but not level-2 latent classes. An empirical demonstration with data from the Trends in International Mathematics and Science Study (TIMSS) 2011 showed that item response probabilities could vary across both level-1 and level-2 latent classes. This relaxed MLCA yielded better model fit and provided more nuanced understanding of the heterogeneous response patterns. Monte Carlo simulation was conducted to evaluate class enumeration and assignment accuracy of the relaxed MLCA. Based on the simulation results, we recommended the use of AIC in class enumeration and highlighted the benefits of having larger cluster size.
Situational judgment tests (SJTs) have much to recommend their use for personnel selection, but because of their low reliability the role of SJTs in behavioural training is largely unexplored. However, research showing that SJTs cannot measure homogenous constructs very well is based exclusively on internal analyses, for example, alpha reliability and factor analysis. In this study, we investigated whether patterns of correlations with external criteria could be used to show that SJT dimension scores are homogenous enough for feedback purposes in leadership development. A multidimensional SJT was designed for 268 high potential leaders on a development programme and used in conjunction with a multisource feedback instrument that measured the same competency framework. The SJT was criterion keyed using against the multisource feedback instrument using an N-Fold cross validation strategy. Convergent and divergent correlations between the SJT scores and corresponding multisource dimension scores suggested that SJT scores can be constructed in a way that permits dimension level feedback that would be useful in leadership development. Running Head: SJTs IN LEADERSHIP DEVELOPMENT 3 Are Situational Judgment Tests Precise Enough for Leadership Development? Situational judgment tests (SJT) are a type of measurement method that can be used to assess a variety of managerial dimensions including social skill, conflict resolution style, or leadership capability (McDaniel, Morgeson, Finnegan, Campion, & Braverman, 2001; McDaniel & Nguyen, 2001, Weekley & Ployhart, 2006). In the personnel selection and development literature, SJTs are classified as low-fidelity work samples (Motowidlo, Dunnette, & Carter, 1990). Typical SJTs consist of several scenarios representing challenging work-related situations. The content of a specific scenario can be presented to respondents in a written, audio, or video format, although the written format is by far the most common. Once an item stem is presented, respondents are asked to choose the most effective and/or least effective response among a set of seemingly equally desirable alternatives. Each alternative typically describes an action that could be taken in response to the scenario situation and has an associated “effectiveness” value. Numerous authors have outlined the case for SJTs in selection context (e.g. (Clevenger et al, 2001, Cullen, Sackett, & Lievens, 2006). Although they can be costly to develop, SJTs are often still more affordable to develop and run than assessment centres or work shadowing programmes. They can also be relatively easily deployed via the Internet or a local area network within organizations, and require considerably less testing time than these other methods. SJTs can also be objectively scored in a manner more like maximum performance measures (e.g., assessment centre simulations or cognitive aptitude tests) than typical performance measures. This means they SJT questions are less susceptible to response distortion issues commonly associated with Likert-type self-report measures. In addition, they also lead to favourable candidate reactions (Anderson, Salgado, Hulsheger, 2013). The validity of the SJT measurement method also explains their use in applied settings. McDaniel et al. (2001) showed with meta-analysis that the average corrected SJTs in LEADERSHIP DEVELOPMENT 4 criterion validity of well-developed SJTs was .34 for predicting job performance. Mechanisms that have been proposed to explain the relationship by Motowidlo, Dunnette, & Carter (1990) and Ployhart & Ehrhart (2003) include a) that SJT scenarios reflect samples of behaviour, and scores correlate with future performance because past behaviour is a good predictor of future behaviour (behavioural consistency); b) that responses to SJT scenarios reflect respondent signalling about their intentions to behave in particular ways in future situations that are like the scenarios, and c) that responses reflect job knowledge required for effective performance, and individuals apply the knowledge they show on the SJT in subsequent situations in the workplace. Researchers have also noted an attractive feature of SJTs is their incremental validity over other assessment methods and low adverse impact against women and ethnic minorities (Chan & Schmitt, 2002; Clevenger et al, 2001; Motowidlo et al., 1990; Olson-Buchanan, Drasgow, Moberg, Mead, Keenan, & Donovan, 1998; Weekley & Jones, 1997, 1999). SJTs in leadership development While SJTs have traditionally been used in personnel selection contexts, there is reason to believe they could have useful applications in training programmes. Because our sample is comprised of leaders, we focus specifically on leadership development programmes. A crucial advantage of SJTs for leadership development is that, due to the ability to make items highly contextualized, they can be considered samples of work performance rather than signs of future work performance (Sackett & Lievens, 2008). The degree of contextualization of SJTs and other assessment methods is referred to as the fidelity of the assessment method (e.g. Lievens & Patterson, 2011). This increased opportunity for item contextualization with SJTs allows test designers to prepare items that are more reflective of the complex situations in which leaders are required to exert influence than traditional Likert style items allow. Before situational judgment tests can be used in the same fashion for development as assessment centres or multisource feedback, it is important to demonstrate that SJTs can be used to deliver precise feedback on specific dimensions where each dimension correlates with SJTs in LEADERSHIP DEVELOPMENT 5 meaningfully different work-related outcomes. We note that research showing such an effect would have implications for personnel selection and development. However, such a finding is not as critical in personnel selection contexts where individual dimension scores are not as emphasized as overall scores. On the other hand, in development settings, narrow dimension scores are as, if not more, important than overall scores. It is these narrow scores that tell candidates where to focus their development efforts. Moreover, our primary focus is on feedback in leadership development contexts because our sample was comprised of participants on a leadership development programme. Evidence from analyses of SJTs scores to date suggests that SJTs do not seem to be assessing homogeneous characteristics. On the contrary, they are known to be highly heterogeneous (Chan & Schmitt, 2006, Lievens, Peeters, & Shollaert, 2008; Weekley & Ployhart, 2006, Whetzel & McDaniel, 2009). To this point, however, attempts to measure constructs with SJTs have been based primarily around internal analyses such as factor analysis or internal consistency analyses. No research has examined whether SJT scores show meaningful patterns of correlations with external variables suggesting that SJT subscales are assessing distinct constructs. The central goal of this study is to examine whether SJT dimension scores are homogenous enough to predict distinct outcomes, as is required for feedback in leadership development, despite the fact that the results of internal analyses alone indicate that SJT dimension scores are highly heterogeneous. If this were the case, feedback on SJT dimension scores could be interpreted in terms of the candidate’s strengths and weaknesses. One research design that would address this issue is to examine the correlations between a multidimensional SJT of a given leadership model and multisource ratings of the same dimension model (i.e. isomorphic content alignment between predictors and criteria). This design would allow us to see whether the layperson assumption about validity holds. In psychometric parlance this can be considered an evaluation of convergent and divergent validity via a multi-trait-multi-method correlation matrix (Campbell & Fiske, 1959). If the SJTs in LEADERSHIP DEVELOPMENT 6 scores for the same dimensions across measurement methods could be shown to be related, the applied relevance of SJT scores for on-the-job behaviors would be more explicitly clear than has been shown to date. It is very important to note that an SJT and a multisource feedback instrument assessing the same competency model represent maximum and typical measurements of the same constructs. Whereas in a typical MTMM design it traits are measured by different methods, in the current design traits are being measured with one method (SJT) and performance related manifestations of these traits are being measured with another method (multisource feedback). Therefore, it would be unreasonable to compare the magnitude of the ‘convergent’ correlations between the same construct across methods against any other standard than the typical magnitude of SJT – job performance correlations. While corrected correlations with job performance have been reported at high as .35 (McDaniel et al., 2001), uncorrected correlations are often much lower. Lievens et al. (2006) for example made a case for the utility of SJT to performance correlations at low as .11. Hypothesis development In hypothesizing about why this expected pattern of relationships might hold between SJT dimension scores and corresponding multisource dimension ratings we considered three theoretical/conceptual perspectives. The first was Motowidlo and colleagues’ theory that SJTs represent past samples of behavior that predict subsequent behavior (Motowidlo, Dunnette, & Carter, 1990). By explicitly improving the point-to-point correspondence between SJT dimensions and performance outcomes by isomorphic alignment between the content models underpinning the predictors and criteria, the correlations between corresponding constructs assessed via different measurement methods would be expected to be stronger. While earlier work (e.g. Lievens, Buyse, & Sackett, 2005) has shown t
This research developed a new ideal point-based item response theory (IRT) model for multidimensional forced choice (MFC) measures. We adapted the Zinnes and Griggs (ZG; 1974) IRT model and the multi-unidimensional pairwise preference (MUPP; Stark et al., 2005) model, henceforth referred to as ZG-MUPP. We derived the information function to evaluate the psychometric properties of MFC measures and developed a model parameter estimation algorithm using Markov chain Monte Carlo (MCMC). To evaluate the efficacy of the proposed model, we conducted a simulation study under various experimental conditions such as sample sizes, number of items, and ranges of discrimination and location parameters. The results showed that the model parameters were accurately estimated when the sample size was as low as 500. The empirical results also showed that the scores from the ZG-MUPP model were comparable to those from the MUPP model and the Thurstonian IRT (TIRT) model. Practical implications and limitations are further discussed.
Although modern item response theory (IRT) methods of test construction and scoring have overcome ipsativity problems historically associated with multidimensional forced choice (MFC) formats, there has been little research on MFC differential item functioning (DIF) detection, whereitemrefers to a block, or group, of statements presented for an examinee's consideration. This research investigated DIF detection with three-alternative MFC items based on the Thurstonian IRT (TIRT) model, using omnibus Wald tests on loadings and thresholds. We examined constrained and free baseline model comparisons strategies with different types and magnitudes of DIF, latent trait correlations, sample sizes, and levels of impact in an extensive Monte Carlo study. Results indicated the free baseline strategy was highly effective in detecting DIF, with power approaching 1.0 in the large sample size and large magnitude of DIF conditions, and similar effectiveness in the impact and no-impact conditions. This research also included an empirical example to demonstrate the viability of the best performing method with real examinees and showed how a DIF and a DTF effect size measure can be used to assess the practical significance of MFC DIF findings.
There has been reemerging interest within psychology in the construct of character, yet assessing it can be difficult due to social desirability of character traits. Forced-choice formats offer one way to address response bias, but traditional scoring methods (i.e., ipsative) associated with this format makes comparing scores between people problematic. Nevertheless, recent advances in modeling item responding (Thurstonian IRT) enable scoring that recovers absolute standing on latent traits and allows for score comparisons between people. Based on recent work in character measurement (CIVIC), we developed a multidimensional forced-choice measure of character (CIVIC-MFC) and scored it using Thurstonian IRT. Initial validation using a sample of 798 participants demonstrated good support for factorial, convergent, and concurrent validity for scores on the CIVIC-MFC, although they did not demonstrate more faking resistance than scores on a Likert-type format version. Potential explanations are discussed.
Forced-choice (FC) measures are gaining popularity as an alternative assessment format to single-statement (SS) measures. However, a fundamental question remains to be answered: Do FC and SS instruments measure the same underlying constructs? In addition, FC measures are theorized to be more cognitively challenging, so how would this feature influence respondents' reactions to FC measures compared to SS? We used both between- and within-subjects designs to examine the equivalence of the FC format and the SS format. As the results illustrate, FC measures scored by the multi-unidimensional pairwise preference (MUPP) model and SS measures scored with the generalized graded unfolding model (GGUM) showed strong equivalence. Specifically, both formats demonstrated similar marginal reliabilities and test-retest reliabilities, high convergent validities, good discriminant validities, and similar criterion-related validities with theoretically relevant criteria. In addition, the formats had little differential impact on respondents' general emotional and cognitive reactions except that the FC format was perceived to be slightly more difficult and more time-saving.
Evidence suggests specific types of combat can influence different posttraumatic stress disorder (PTSD) symptom presentations. This study applied structural equation modeling to extend this line of research in a cross-sectional sample of 2,159 U.S. Navy Sailors deployed to Iraq and Afghanistan between 2007 and 2008. Generally, support was for a model which specified direct relationships between the latent factors of combat experience and PTSD (χ2(881) = 4,015.82, p < .001, CFI = .956, TLI = .953, RMSEA = .040, 90%CI = .038, .041). Specifically, exposure to a combat environment predicted only the reeexperiencing and hyperarousal dimensions of PTSD, close physical engagement predicted all but the hyperarousal dimension, and nearness to serious injury or death predicted only reeexperiencing. While exploratory in nature, these findings support the hypothesis that an individual’s unique combat experience may also produce unique expressions of PTSD symptoms, implying that PTSD prevention, diagnosis, and treatment may need to be tailored to specific types of trauma and combat experiences.