Abstract This chapter sets the stage by focusing on the concept of validity, documenting key changes over time in how the term is used and examining the specific ways in which the concept is instantiated in the domain of personnel selection. It then moves from conceptual to operational and discusses issues in the use of various strategies to establish what is termed the predictive inference, namely, that scores on the predictor measure of interest can be used to draw inferences about an individual’s future job behavior or other criterion of interest. Finally, the chapter addresses a number of specialized issues aimed at illustrating some of the complexities and nuances of validation.
Bobko et al. (2024) responded to Sackett et al.'s (2022) compilation of meta-analytic evidence for the validity of a wide variety of measures used as predictors of overall job performance, offering a set of alternative methodological choices which they term "considered estimation" to counter the Sackett et al. approach of "conservative estimation." Here we offer a rebuttal to Bobko et al. A primary concern is that Bobko et al. apply the label "conservative estimation" to the full range of methodological choices made by Sackett et al. Yet, we clarify the narrow and specific meaning of "conservative estimation," and note that that the bulk of Bobko et al.'s concerns are independent of the principle of conservative estimation. We also respond to Bobko et al.'s two key concerns, namely, comparing validity estimates when one is corrected for range restriction and one is not and comparing validity estimates for predictors reflecting psychological constructs and those reflecting measurement methods, and also briefly address a range of other critiques offered by Bobko et al.
Background Learner assessments of faculty are widespread in medicine, yet concerns are growing about possible biases in these assessments and their associations with gender disparities. Objective To investigate gender-based differences in how residents and fellows describe faculty (rater effect) and how faculty are described (ratee effect) in faculty assessments, and their associations with teaching effectiveness ratings. Methods We analyzed 2164 trainee assessments of University of Minnesota Medical School faculty from 2019 to 2023 with trainee and faculty gender information and narrative comments. Using natural language processing, we categorized words and 2-word groups (n-grams) into communal (eg, caring, kind), standout (eg, outstanding, amazing), and agentic/ability (eg, assertive, controlling) groups. We examined gender-based differences in n-grams used by trainees (rater effect) and received by faculty (ratee effect), and relationships between n-gram and teaching effectiveness ratings. Results Women trainees used more communal (rater effect, incidence rate ratio [IRR]=1.36; 95% CI, 1.27-1.47), standout (IRR=1.20; 95% CI, 1.08-1.34), and agentic/ability words (IRR=1.37; 95% CI, 1.26-1.49; P<.001) than men trainees. Women faculty received fewer agentic/ability words than men faculty (ratee effect, IRR=0.83; 95% CI, 0.77-0.90; P<.001). Women trainees used fewer communal words when describing women faculty (interaction effect, IRR=0.84; 95% CI, 0.73-0.98; P<.05). Teaching effectiveness ratings correlated with faculty n-gram word frequency in standout (men: rs =0.29, women: rs=0.28, P<.001) and communal categories (men: rs =0.23, P=.003; women: rs=0.22, P=.01). Conclusions Women trainees used more communal, standout, and agentic/ability descriptors, while women faculty had fewer agentic/ability descriptors. Women trainees used fewer communal words when describing women faculty. Standout and communal word frequency predicted teaching effectiveness ratings for both genders.
Most studies of student performance focus on the between-person measures, such as GPA, which ignore the within-person variability across a time period and meaningful differences in performance at the individual level. Operationalizing grade variability as the standard deviation of grades in a single course across a semester, we explored a number of personal predictors, including personality, test anxiety, procrastination, and self-efficacy, of students' performance consistency over time. The results indicate that industriousness and intellect aspects negatively predict grade variability and the openness aspect positively predicts grade variability. Therefore, students with a high level of industriousness and intellect tend to perform more consistently and those who are high on openness tend to exhibit a more inconsistent performance pattern across a semester. Our study taps into the temporal stability aspect of student performance and carries implications for potential intervention opportunities for students low on industriousness and intellect, and high on openness.
The question of what assessment centers measure has remained a controversial topic in research for decades, with a recent increase in studies that (1) use generalizability theory and (2) acknowledge the effects of aggregating post-exercise dimension ratings into higher-level assessment center scores. Building on these developments, we used Bayesian generalizability theory and random-effects meta-analyses to examine the variance explained by assessment center components such as assessees, exercises, dimensions, assessors, their interactions, and the interrater reliability of AC ratings in 19 different assessment center samples from various organizations (N = 4,963 assessees with 272,528 observations). This provides the first meta-analytic estimates of these effects, as well as insight into the extent to which findings from previous studies generalize to assessment center samples that differ in measurement design, industry, and purpose, and how heterogeneous these effects are across samples. Results were consistent with previous trends in the ranking of variance explained by key AC components (with assessee main effects and assessee-exercise effects being the largest variance components) and additionally emphasized the relevance of assessee-exercise-dimension effects. In addition, meta-analytic results suggested substantial heterogeneity in all reliable variance components (i.e., assessee main effect, assessee-exercise effect, assessee-dimension effect, and assessee-exercise-dimension effect) and in interrater reliability across assessment center samples. Aggregating AC ratings into higher-level scores (i.e., overall AC scores, exercise-level scores, and dimension-level scores) reduced heterogeneity only slightly. Implications of the findings for a multifaceted assessment center functioning are discussed.
Sexual harassment and sexual assault are increasingly an area of employer concern. Although employers commonly use noncognitive personnel screening measures to reduce counterproductive work behaviors (CWB), the potential for such measures to reduce perpetration of sexual assault and sexual harassment, specifically, has received little attention. The current paper describes two studies to evaluate the potential value of including both domain-general (overt integrity test admissions, academic biodata, self-report personality) and domain-specific measures (explicitly referencing attitudes toward gender and relationships) in employee screening. Study 1 demonstrates that domain-general and domain-specific measures correlated with anonymous admissions of prior sexual coercion and sexual harassment intent among both males and females. Study 2 demonstrates that domain-specific measures (self-report attitudes towards women and depersonalized relationships) are also correlates of intentions to engage in broader CWB, even when individuals are directed to present themselves in a way they believe would maximize chances of personnel selection. Overall results support the use of such personnel screening measures as part of an organizational strategy to address sexual harassment, sexual assault, and other workplace deviance.
Given the centrality of the job performance construct to organizational researchers, it is critical to understand the reliability of the most common way it is operationalized in the literature. To this end, we conducted an updated meta-analysis on the interrater reliability of supervisory ratings of job performance (k = 132 independent samples) using a new meta-analytic procedure (i.e., the Morris estimator), which includes both within- and between-study variance in the calculation of study weights. An important benefit of this approach is that it prevents large-sample studies from dominating the results. In this investigation, we also examined different factors that may affect interrater reliability, including job complexity, managerial level, rating purpose, performance measure, and rater perspective. We found a higher interrater reliability estimate (r = .65) compared to previous meta-analyses on the topic, and our results converged with an important, but often neglected, finding from a previous meta-analysis by Conway and Huffcutt (1997), such that interrater reliability varies meaningfully by job type (r = .57 for managerial positions vs. r = .68 for nonmanagerial positions). Given this finding, we advise against the use of an overall grand mean of interrater reliability. Instead, we recommend using job-specific or local reliabilities for making corrections for attenuation.
Conceptualizations of workplace aggression converge in treating intent to harm others as a necessary feature of aggression. However, inspection of workplace aggression scales suggests that many items do not specify intent to harm. In a series of three studies, we examined the effect of inclusion of intent to harm on workplace aggression's psychometric properties. Study 1 found that existing workplace aggression scales do not consistently specify or imply intent to harm. Study 2 found that inclusion of intent to harm has substantial implications for aggression's occurrence rate. Prior research that does not assess intent to harm overestimates the frequency of aggression. Study 3A found that workplace aggression's correlations with external variables were also overestimated when failing to include intent to harm. We found that aggression measured without specifying intent is highly correlated with counterproductive work behavior (CWB), whereas aggression measured with intent specified is empirically distinguished from CWB. In Study 3A, a construct-valid workplace aggression scale was created, called the Intentional Workplace Aggression Scale (IWAS). Study 3B showed that the IWAS displayed relationships with affective constructs, such as trait anger and emotional stability, as well as with situational variables, such as job satisfaction and organizational justice perceptions. In applied contexts, workplace aggression is a valued criterion that has an array of negative consequences for workers. Research converges in defining "intent to harm" others as a necessary feature of workplace aggression. Workplace aggression scales do not sufficiently capture intent to harm. Consequently, past research has overestimated aggression's base rate and external correlates. A new, construct-valid scale is created that can be used as a criterion measure in applied settings for assessing aggressive workplace experiences and designing selection interventions to address aggression.
Twenty years ago, Rotundo et al. (2001) meta-analyzed the gender differences in sexual harassment (SH) perception. They found an overall d of 0.30: Women are more likely than men to label certain behaviors as SH. Much has changed since then, including the increased social awareness and the prevalence of SH training. Given the prevalence of SH in the workplace and the importance of SH perception in SH research, we conducted a mixed-methods research program to explore possible changes in the gender gap. In Study 1 (k = 72, N = 27,767), we meta-analyzed the perceptual gender differences to compare with those in Rotundo et al. and examined several moderators of the differences. We found an overall mean d of 0.33, implying a similar gender gap in SH perception as 20 years ago, yet none of the moderators examined in this study showed significant results. In Study 2, we empirically examined gender differences in mean levels of SH perception using the same measurement scales used in two older studies and compared with the differences found in these two studies. We found higher levels of SH perception for both men and women, but no difference in the mean d between men and women, suggesting that no change over time in mean d does not mean no change in SH perception. The implications of our findings are discussed.
A large body of literature has studied the effect of stereotype threat and stereotype lift on cognitive test performance. Research on stereotype threat (ST) examines whether the awareness of a negative stereotype can decrease stereotyped group members' test performance. A less commonly studied influence of stereotypes is stereotype lift (SL), defined as an increase in a group's test performance due to not being part of a negative stereotype. For example, men might perform better on math tests if they are primed on the stereotype that men are better than women at math. Walton and Cohen (2003) previously meta-analyzed the impact of SL on cognitive tests, finding an overall d = 0.24. We report an updated meta-analysis on SL with more samples and moderator analyses. We then meta-analyzed between-group effects (majority-minority group differences both in the presence and absence of SL and ST) to compare their relative contributions to subgroup mean differences on cognitive tests. Our results indicate that SL has a small influence on cognitive test performance (d = 0.09, SDres = 0.19), and that subgroup mean differences result largely from between-group effects rather than from the effects of ST and SL.
The relationship between general cognitive ability (GCA) and overall job performance has been a long-accepted fact in industrial and organizational psychology. However, the most prominent data on this relationship date back more than 50 years. This meta-analysis examines the relationship between GCA and overall job performance using studies from the current century. Results across 153 samples and a total sample size of 40,740 show a mean observed validity of .16, with a residual SD of .09. Correcting for unreliability in the criterion and correcting predictive studies for range restriction produces a mean corrected validity of .22 and a residual SD of .11. While this is a much smaller estimate than the .51 value offered by Schmidt and Hunter (1998), that value has been critiqued by Sackett et al. (2022), who offered a mean corrected validity of .31 based on integrating findings from prior meta-analyses of 20th century data. We obtain a lower value (.22) for 21st century data. We conclude that GCA is related to job performance, but our estimate of the magnitude of the relationship is lower than prior estimates.
PURPOSE:This study aimed to develop an instrument to measure medical trainees' perceptions of justice in clinical learning environments. METHOD:Between 2019 and 2023, the authors conducted a multiyear, multi-institutional, multiphase study to develop a 16-item justice measure with 4 dimensions: interpersonal, informational, procedural, and distributive. The authors gathered validity evidence based on test content, internal structure, and relationships with other variables across 3 phases. Phase 1 involved drafting items and gathering evidence that items measured intended dimensions. Phase 2 involved analyzing relevance of items for target groups, examining interitem correlations and factor loadings in a preliminary analysis, and obtaining reliability estimates. Phase 3 involved a confirmatory factor analysis and collecting convergent and discriminant validity evidence. RESULTS:In phase 1, 63 of 91 draft items were retained following a content validation exercise gauging how well items measured targeted dimensions (mean [SD] item ratings within dimensions, 4.16 [0.36] to 4.39 [0.34]) on a 5-point Likert scale (with 1 indicating not at all well and 5 indicating extremely well). In phase 2, 30 items were removed due to low factor loadings (i.e., < 0.40), and 4 items per dimension were selected (factor loadings, 0.42-0.89). In phase 3, a confirmatory factor analysis supported the 4-dimensional model ( χ2 = 610.14, P < .001; comparative fit index = 0.90, Tucker-Lewis Index = 0.87, root mean squared error of approximation = 0.11, standardized root mean squared residual = 0.06), with convergent and discriminant validity evidence showing hypothesized positive correlations with a justice measure ( r = 0.93, P < .001), trait positive affect ( r = 0.46, P < .001), and emotional stability ( r = 0.33, P < .001) and negative correlations with trait negative affect ( r = -0.39, P < .001). CONCLUSIONS:Results indicate the measure's potential utility in understanding justice perceptions and designing targeted interventions.
Organizational citizenship behavior (OCB) is often viewed as an unequivocal boon. However, differing motivations and external pressures can change OCB's relationship with counterproductive work behavior (CWB) and sexual harassment. We take a novel approach to understanding the relationship between OCB, CWB and sexual harassment by exploring the role of engaging in interpersonally directed OCB and CWB because of targeted colleagues' sex or gender. We use the terms "gendered OCB" and "gendered CWB" to refer to engaging in OCB or CWB because of the gender or sex of the target of the behavior (e.g. a colleague). We examined the relationships among OCB, CWB, and sexual harassment in a sample of 503 Prolific users (60.2% men) in the United States. Interpersonally directed OCB that was sex/gender agnostic had near-zero correlations with general CWB and sexual harassment. However, gendered OCB had significant and positive relationships with both general CWB (r = .15) and engaging in sexual harassment (r = .35). Gendered OCB's positive association with both desirable and undesirable behaviors is reflected in models fitting best when gendered OCB loaded on both the sexual harassment and OCB latent factors. Such findings challenge views of citizenship behaviors as universally "good."
General mental ability (GMA) tests have long been at the heart of the validity-diversity trade-off, with conventional wisdom being that reducing their weight in personnel selection can improve adverse impact, but that this results in steep costs to criterion-related validity. However, Sackett et al. (2022) revealed that the criterion-related validity of GMA tests has been considerably overestimated due to inappropriate range restriction corrections. Thus, we revisit the role of GMA tests in the validity-diversity trade-off using an updated meta-analytic correlation matrix of the relationships six selection methods (biodata, GMA tests, conscientiousness tests, structured interviews, integrity tests, and situational judgment tests) have with job performance, along with their Black-White mean differences. Our results lead to the conclusion that excluding GMA tests generally has little to no effect on validity, but substantially decreases adverse impact. Contrary to popular belief, GMA tests are not a driving factor in the validity-diversity trade-off. This does not fully resolve the validity-diversity trade-off, though: Our results show there is still some validity reduction required to get to an adverse impact ratio of .80, although the validity reduction is less than previously thought. Instead, it shows that the validity-diversity trade-off conversation should shift from the role of GMA tests to that of other selection methods. The present study also addresses which selection methods now emerge as most valid and whether composites of selection methods can result in validities similar to those expected prior to Sackett et al. (2022). (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Cognitive ability tests are widely used in employee selection contexts, but large race and ethnic subgroup mean differences in test scores represent a major drawback to their use. We examine the potential for an item-level procedure to reduce these test score mean differences. In three data sets, differing proportions of cognitive ability test items with higher levels of difficulty or subgroup mean differences were removed from the tests. The reliabilities of these trimmed tests were then corrected back to the lengths of the original tests, and the subgroup mean differences of the trimmed tests were compared to those of the original tests. Results indicate that it is not possible to come anywhere close to eliminating subgroup differences via item trimming. The procedure may modestly reduce subgroup mean differences in test scores, with effects becoming stronger as higher proportions of items are removed from the tests. Removing items based on difficulty or subgroup differences have roughly similar impacts on test score mean differences for Black-White test taker comparisons, but results are more mixed for Hispanic-White comparisons. Our results also provide preliminary evidence that removing items on the basis of subgroup mean differences may have relatively little effect on test criterion-related validity, but the impact of removing difficult items was more mixed.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
While the dominant finding indicates a monotonic relationship between cognitive ability and academic performance, some researchers have suggested the existence of cognitive thresholds for challenging coursework, such that a certain level of cognitive ability is required for reaching a satisfactory level of academic achievement. Given the significance of finding a threshold for understanding the relationship between cognitive ability and academic performance, and the limited studies on the topic, it is worth further investigating the possibility of cognitive thresholds. Using a multi-institutional dataset and the necessary condition analysis (NCA), we attempted to replicate previous findings of cognitive thresholds on the major GPA of mathematics and physics-majored students, as well as the course grade of organic chemistry, to examine whether high SAT math scores constitute a necessary condition for obtaining satisfactory grades in these courses. The results from the two studies do not indicate an absolute cognitive threshold point below which students are doomed to fail regardless of the amount of effort they devote into learning. However, we did find that the chance of students with a low level of quantitative ability to succeed in highly quantitative courses is very small, which qualifies for the virtually necessary condition.
Purpose To examine whether gender differences exist in medical trainees’ (residents’ and fellows’) evaluations of faculty at a number of clinical departments. Method The authors conducted a single-institution (University of Minnesota Medical School) retrospective cohort analysis of 5,071 trainee evaluations of 447 faculty (for which trainee and faculty gender information was available) completed between July 1, 2019, and June 30, 2022. The authors developed and employed a 17-item measure of clinical teaching effectiveness, with 4 dimensions: overall teaching effectiveness, role modeling, facilitating knowledge acquisition, and teaching procedures. Using both between- and within-subject samples, they conducted analyses to examine gender differences among the trainees making ratings (rater effects), the faculty receiving ratings (ratee effects), and whether faculty ratings differed by trainee gender (interaction effects). Results There was a statistically significant rater effect for the overall teaching effectiveness and facilitating knowledge acquisition dimensions (B = −0.28 and −0.14, 95% CI: [−0.35, −0.21] and [−0.20, −0.09], respectively, P < .001, medium corrected effect sizes between −0.34 and −0.54); female trainees rated male and female faculty lower than male trainees on both dimensions. There also was a statistically significant ratee effect for the overall teaching effectiveness and role modeling dimensions (B = −0.09 and −0.08, 95% CI: [−0.16, −0.02] and [−0.13, −0.04], P = .01 and < .001, respectively, small to medium corrected effect sizes between −0.16 and −0.44); female faculty were rated lower than male faculty on both dimensions. There was not a statistically significant interaction effect. Conclusions Female trainees rated faculty lower than male trainees and female faculty were rated lower than male faculty on 2 teaching dimensions each. The authors encourage researchers to continue to examine the reasons for the evaluation differences observed and how implicit bias interventions might help to address them.
I-O psychologists often face the need to reduce the length of a data collection effort due to logistical constraints or data quality concerns. Standard practice in the field has been either to drop some measures from the planned data collection or to use short forms of instruments rather than full measures. Dropping measures is unappealing given the loss of potential information, and short forms often do not exist and have to be developed, which can be a time-consuming and expensive process. We advocate for an alternative approach to reduce the length of a survey or a test, namely to implement a planned missingness (PM) design in which each participant completes a random subset of items. We begin with a short introduction of PM designs, then summarize recent empirical findings that directly compare PM and short form approaches and suggest that they perform equivalently across a large number of conditions. We surveyed a sample of researchers and practitioners to investigate why PM has not been commonly used in I-O work and found that the underusage stems primarily from a lack of knowledge and understanding. Therefore, we provide a simple walkthrough of the implementation of PM designs and analysis of data with PM, as well as point to various resources and statistical software that are equipped for its use. Last, we prescribe a set of four conditions that would characterize a good opportunity to implement a PM design.
Sackett et al. (2022) identified previously unnoticed flaws in the way range restriction corrections have been applied in prior meta-analyses of personnel selection tools. They offered revised estimates of operational validity, which are often quite different from the prior estimates. The present paper attempts to draw out the applied implications of that work. We aim to a) present a conceptual overview of the critique of prior approaches to correction, b) outline the implications of this new perspective for the relative validity of different predictors and for the tradeoff between validity and diversity in selection system design, c) highlight the need to attend to variability in meta-analytic validity estimates, rather than just the mean, d) summarize reactions encountered to date to Sackett et al., and e) offer a series of recommendations regarding how to go about correcting validity estimates for unreliability in the criterion and for range restriction in applied work.