In this reply to the commentaries by Mirowska (2025), Hickman (2025), and Holtrop and Bronzwaer (2026), we expand on our initial provocation article regarding the effects of candidate use of Generative AI (GenAI) in personnel selection (Lievens and Dunlop, 2025). First, we update the discussion by highlighting recent technological developments (agentic AI and AI-integrated wearables) that accelerate the threat of candidate GenAI use as a substitute for candidate effort. Second, we clarify and build upon our original arguments concerning the effects of candidate GenAI use on construct-related validity and subgroup differences, introducing the concept of "GenAI literacy" as a potential confounding construct. In doing so, we elaborate on the concept of AI-enabled assessment designs. Finally, we integrate the insights from the three commentaries with our own thinking to introduce the FAIR framework (Forbid, Advise, Insulate, Reimagine), which should help employers navigate the complex landscape of candidate GenAI use. The FAIR framework distinguishes between strategies aimed at preventing GenAI misuse and those designed to embrace and integrate GenAI into assessment processes. We conclude that the future of selection lies not in banning candidate GenAI use entirely. Instead, we argue for a strategic shift toward reimagining assessments for an AI-augmented world.
SJTs have traditionally been conceptualized as low-fidelity simulations of how respond to work-related situations. Yet, several studies demonstrated that situation descriptions were necessary to solve some SJT items, whereas they were not needed for other SJT items. So far, no solid support was found for various factors (e.g., item characteristics, presentation format, instructions, and content domain) that make some situation descriptions relevant and others irrelevant for responding to SJT items. Building on trait activation theory, we posit that trait-relevant situational cues serve as an ignored factor in SJT situation descriptions. Across two main studies (N-1 = 269, N-2 = 1,092), we manipulated the availability of trait-relevant situational cues in SJT items to examine their impact on SJT scores and convergent validity. We found a main effect for the availability of trait-relevant situational cues on the proportion of correctly solved SJT items, but only a marginal impact on convergent validity. The current study thus helps explaining why some SJT items can be situated on the context-dependent side, whereas other items are on the context-independent side. At a practical level, this suggests that more systematically embedding trait-relevant situational cues into SJT descriptions may be a helpful design consideration to make SJT items more context-dependent, even if its impact on convergent validity appears to be limited.
In assessment and selection, organizations often include interpersonal interactions because they provide insights into candidates' interpersonal skills. These skills are then typically assessed via one-shot, retrospective assessor ratings. Unfortunately, the assessment of interpersonal skills at such a trait-like level fails to capture the richness of how the interaction unfolds at the behavioral exchange level within a role-play assessment. This study uses the lens of interpersonal complementarity theory to advance our understanding of interpersonal dynamics in role-play assessment and their effects on assessor ratings. Ninety-six MBA students participated in four different flash role-plays as part of diagnosing their strengths and weaknesses. Apart from gathering assessor ratings and criterion measures, coders also conducted a fine-grained examination of how the behavior of the two interaction partners (i.e., MBA students and role-players) unfolded at the moment-to-moment level via the Continuous Assessment of Interpersonal Dynamics (CAID) measurement tool. In all role-plays, candidates consistently showed mutual adaptations in line with complementarity principles: Affiliative behavior led to affiliative behavior, whereas dominant behavior resulted in docile, following behavior and vice versa. For affiliation, mutual influence also occurred in that both interaction partners' temporal trends in affiliation were entrained over time. Complementarity patterns were significantly related to ratings of in situ (role-playing) assessors but not to ratings of ex situ (remote) assessors. The effect of complementarity on validity was mixed. Overall, this study highlights the importance of going beyond overall ratings to capture behavioral contingencies such as complementarity patterns in interpersonal role-play assessment. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
There is increasing attention to storification of assessments (i.e., embedding a storyline into a non-storified assessment) in research and practice and to gamified and game-based assessment in general. However, there is a surprising lack of agreement and of recommendations regarding what level of fantasy of the storyline one should choose for the storification from the perspective of applicant reactions. A distinction is typically made between fantasy (e.g., fighting aliens) and realistic (e.g., workday simulations) storylines, with both choices having their advantages and disadvantages. In this study, a sample of 195 participants was shown either a storified realistic test, a storified fantasy test, or a non-storified test. Afterwards, they rated various applicant reaction measures. Both storified assessments were rated equally positively on perceived modernity of the organization and enjoyment but the storified realistic test was superior to the storified fantasy test in terms of perceived job-relatedness, procedural fairness, organizational attractiveness, and clarity of work activity. Thus, the level of fantasy of a storyline in a storified assessment plays an important role for applicant reaction variables, whereby the overall pattern of results showed that the storified realistic test was rated most favorably, followed by the non-storified test, and the storified fantasy assessment.
To date, a limited set of studies have compared the criterion-related validity of low-fidelity (SJT) versus high-fidelity (AC) simulations for predicting job performance. Unfortunately, these studies validated these simulations through the overall assessment rating (OAR) instead of on the basis of specific dimensions. Given SJTs and ACs were compared that measured different dimensions, our understanding of the relative and comparative validity of these assessment approaches in measuring the same set of dimensions is still limited. Therefore, this study aims to conduct a head-to-head comparison of the criterion-related validity of the AC and the SJT (and their incremental validity) while keeping the performance dimensions under investigation constant. Data were collected from 406 applicants for supervisory and management positions in a large Iranian steel industry company. In this process, a general mental ability test, a personality inventory, an SJT, and an AC were used as predictors, and supervisory ratings of job performance dimensions (Thinking, Feeling, and Power) served as criteria. The AC had relatively high validity for all three dimensions, whereas the SJT had a similar validity only for the Thinking dimension. So, the SJT was significantly weaker in assessing the Feeling and Power dimensions. These results were confirmed by incremental validity analyses. Overall, this study shows that understanding the relationships between predictor and criterion dimensions plays a critical role in developing valid selection systems.
General mental ability (GMA) tests have long been at the heart of the validity-diversity trade-off, with conventional wisdom being that reducing their weight in personnel selection can improve adverse impact, but that this results in steep costs to criterion-related validity. However, Sackett et al. (2022) revealed that the criterion-related validity of GMA tests has been considerably overestimated due to inappropriate range restriction corrections. Thus, we revisit the role of GMA tests in the validity-diversity trade-off using an updated meta-analytic correlation matrix of the relationships six selection methods (biodata, GMA tests, conscientiousness tests, structured interviews, integrity tests, and situational judgment tests) have with job performance, along with their Black-White mean differences. Our results lead to the conclusion that excluding GMA tests generally has little to no effect on validity, but substantially decreases adverse impact. Contrary to popular belief, GMA tests are not a driving factor in the validity-diversity trade-off. This does not fully resolve the validity-diversity trade-off, though: Our results show there is still some validity reduction required to get to an adverse impact ratio of .80, although the validity reduction is less than previously thought. Instead, it shows that the validity-diversity trade-off conversation should shift from the role of GMA tests to that of other selection methods. The present study also addresses which selection methods now emerge as most valid and whether composites of selection methods can result in validities similar to those expected prior to Sackett et al. (2022). (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Sackett et al. (2022) identified previously unnoticed flaws in the way range restriction corrections have been applied in prior meta-analyses of personnel selection tools. They offered revised estimates of operational validity, which are often quite different from the prior estimates. The present paper attempts to draw out the applied implications of that work. We aim to a) present a conceptual overview of the critique of prior approaches to correction, b) outline the implications of this new perspective for the relative validity of different predictors and for the tradeoff between validity and diversity in selection system design, c) highlight the need to attend to variability in meta-analytic validity estimates, rather than just the mean, d) summarize reactions encountered to date to Sackett et al., and e) offer a series of recommendations regarding how to go about correcting validity estimates for unreliability in the criterion and for range restriction in applied work.
Currently used Pareto-optimal (PO) approaches for balancing diversity and validity goals in selection can deal only with one minority group and one criterion. These are key limitations because the workplace and society at large are getting increasingly diverse and because selection system designers often have interest in multiple criteria. Therefore, the article extends existing methods for designing PO selection systems to situations involving multiple criteria and multiple minority groups (i.e., multiobjective PO selection systems). We first present a hybrid multiobjective PO approach for computing selection systems that are PO with respect to (a) a set of quality objectives (i.e., criteria) and (b) a set of diversity objectives where each diversity objective relates to a different minority group. Next, we propose three two-dimensional subspace procedures that aid selection designers in choosing between the PO systems in case of a high number of quality and diversity objectives. We illustrate our novel multiobjective PO approaches via several example applications, thereby demonstrating that they are the first to reveal the complete gamut of eligible PO selection designs and to faithfully capture the Pareto trade-off front in case of more than two objectives. In addition, a small-scale cross-validation study confirms that the resulting PO selection designs retain an advantage over alternative designs when applied in new validation samples. Finally, the article provides a link to an executable code to perform the new multiobjective PO approaches. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Recently, shorter assessments have emerged as potential alternatives for more resourceful traditional selection approaches. Multiple, speeded assessments (MSAs) represent such an alternative. In MSAs, candidates participate in a large number of short (a maximum of 5 min), behavioral simulations in which they face a variety of job situations. Initial psychometric evidence on the validity of MSAs is promising. Yet, validity represents only one piece of evidence. It is not known whether MSAs disadvantage specific subgroups, which may inhibit diversity. There is also no information on candidates' experience of going through an MSA, which is pivotal for the attractiveness of the organization's selection process. Therefore, this study investigates an MSA in terms of subgroup differences (gender and nationality) and applicant perceptions. Master of Business Administration (MBA) students (N = 96) proceeded through 18 short role-plays sampling junior management situations. Score differences between men and women were negligible. Yet, there were large score differences between national citizens and foreigners. There was no evidence for predictive bias for nationality, though. Of the applicant reaction measures, interpersonal treatment perceptions contributed most to overall fairness perceptions. These findings add to the evidence in support of MSAs, while also stressing to remain vigilant for potential score differences among subgroups.
Work effort has been a key concept in management theories and research for more than a century. Maintaining and increasing employee effort also is a persistent concern to managers. The goal of the present conceptual and meta-analytic review was to increase clarity and consensus regarding what effort is and how to measure it. First, we reviewed conceptualizations of effort and provided an integrated definition that views effort as a direct outcome of motivation that captures (a) what employees work on, (b) how hard they work, and (c) how long they persist in that work. Second, we identified four main ways researchers have operationalized effort and meta-analytically studied the effects of each operationalization on effort–job performance relationships. For example, measures that assessed multiple dimensions of effort (ρ = .37) tended to relate more strongly to performance than measures that focused on only one dimension (e.g., effort intensity) or on effort more generally (ρ = .18 to .29). Third, we developed and meta-analytically tested a nomological network to gain a better understanding of effort's antecedents (e.g., intrinsic motivation, ρ = .46; performance orientation, ρ = .12) and outcomes (e.g., job performance, ρ = .34; exhaustion, ρ = .04) as well as constructs that appear to overlap with effort (e.g., work engagement, ρ = .48; grit, ρ = .51). Finally, on the basis of our conceptual and meta-analytic reviews, we delineated an agenda for future research on this central, yet often misunderstood, construct.
There has been a growing interest in third-party employment branding (TPEB) because prospective and current employees perceive it more credible than the company-controlled employer branding. The academic research on TPEB has also been rapidly growing. This chapter reviews the TPEB research using a bibliometric analysis of 734 articles published between 1996 and 2021. The analysis shows that 'employer branding', 'recruitment'‚ 'Glassdoor' and 'word-of-mouth' are the major keywords in this domain. TPEB research can be grouped into three themes - (i) 'best employer status and its outcomes' (ii) 'antecedents and consequences of third-party employment branding' and (iii) 'word-of-mouth and recruitment'. We propose directions for future research in these three areas. Additionally, we recommend further research in the areas such as influence of TPEB on financial metrics, effect of negative TPEB information on companies, counter-productive effects of best employer surveys, inclusiveness of best employer surveys and cross-fertilization between research on employer branding, internal branding, and TPEB.
Personnel Psychology has a long tradition of publishing important research on personnel selection. In this article, we review some of the key questions and findings from studies published in the journal and in the selection literature more broadly. In doing so, we focus on the various decisions organizations face regarding selection procedure development (e.g., use multiple selection procedures, contextualize procedure content), administration (e.g., provide pre-test explanations, reveal target knowledge, skills, abilities, and other characteristics [KSAOs]), and scoring (e.g., weight predictors and criteria, use artificial intelligence). Further, we focus on how these decisions affect the validity of inferences drawn from the procedures, how use of the procedures may affect organizational diversity, and how applicants experience the procedures. We also consider factors such as cost and time. Based on our review, we highlight practical implications and key directions for future research.
Assessment center (AC) exercises such as role-plays have established themselves as valuable approaches for obtaining insights into interpersonal behavior, but they are often considered the "Rolls Royce" of personnel assessment due to their high costs. The observation and rating process comprises a substantial part of these costs. In an exploratory case study, we capitalize on recent advances in natural language processing (NLP) by developing NLP-based machine learning (ML) models to investigate the possibility of automatically scoring AC exercises. First, we compared the convergent-related validity and contamination with word count of ML scores based on models that used different NLP methods to operationalize verbal behavior. Second, for the model that maximized convergence while minimizing contamination with word count (i.e., a model that used both n-grams and Universal Sentence Encoder embeddings as predictors), we investigated the criterion-related validity of its scores. Third, we examined how the interrater reliability of the AC role-play scores affects ML model convergence. To do so, we applied seven NLP methods to 96 assessees' transcriptions and trained 10 sets of ML models across 18 speeded AC role-plays to automatically score assessee performance. Results suggest that ML scores recovered most of the original variance in the overall assessment ratings, and replacing one or more human assessors with ML scores maintained criterion-related validity. Additionally, ML models seemed to exhibit higher convergence when assessors consistently detected and utilized observable behaviors to make ratings (i.e., when interrater reliability was higher). Finally, we provide a step-by-step guide for practitioners seeking to implement ML scoring in ACs.
Recently, multiple, speeded assessments (e.g., "speeded" or "flash" role-plays) have made rapid inroads into the selection domain. So far, however, the conceptual underpinning and empirical evidence related to these short, fast-paced assessment approaches has been lacking. This raises questions whether these speeded assessments can serve as reliable and valid indicators of future performance. This article uses the notions of stimulus and response domain sampling to conceptualize multiple, speeded behavioral job simulations as a hybrid of established simulation-based selection methods. Next, we draw upon the thin slices of behavior paradigm to theorize about the quality of ratings made in multiple, speeded behavioral simulations. In two studies, various assessor pools assessed a sample of 96 MBA students in 18 3-min role-plays designed to capture situations in the junior management domain. At the individual speeded role-play level, reliability and validity were not ensured. Yet, aggregated across all assessors' ratings of all speeded role-plays, the overall score for predicting future performance was high (.54). Validities remained high when assessors evaluated only the first minute (vs. full 3 min) or received only a control training (vs. traditional assessor training). Aggregating ratings of performance in multiple, heterogeneous situations that elicit a variety of domain-relevant behavior emerged as key requirement to obtain adequate domain coverage, capture both ability and personality (extraversion and agreeableness), and achieve substantial validities. Overall, these results show the importance of the stimulus and response domain sampling logic and send a strong warning to using "single" speeded behavioral simulations in practice. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
This study draws from brand positioning research to introduce the notions of points-of-relevance and points-of-difference to employer image research. Similar to prior research, this means that we start by investigating the relevant image attributes (points-of-relevance) that potential applicants use for judging organizations' attractiveness as an employer. However, we go beyond past research by examining whether the same points-of-relevance are used within and across industries. Next, we further extend current research by identifying which of the relevant image attributes also serve as points-of-difference for distinguishing between organizations and industries. The sample consisted of 24 organizations from 6 industries (total N = 7171). As a first key result, across industries and organizations, individuals attached similar importance to the same instrumental (job content, working conditions, and compensation) and symbolic (innovativeness, gentleness, and competence) image attributes in judging organizational attractiveness. Second, organizations and industries varied significantly on both instrumental and symbolic image attributes, with job content and innovativeness emerging as the strongest points-of-difference. Third, most image attributes showed greater variation between industries than between organizations, pointing at the importance of studying employer image at the industry level. Implications for recruitment research, employer branding, and best employer competitions are discussed.
This paper reveals the characteristics and effects of nonverbal behavior and human mimicry in the context of application interviews. It discloses a novel analyzation method for psychological research by utilizing machine learning. In comparison to traditional manual data analysis, machine learning proves to be able to analyze the data more deeply and to discover connections in the data invisible to the human eye. The paper describes an experiment to measure and analyze the reactions of evaluators to job applicants who adopt specific behaviors: mimicry, suppress, immediacy and natural behavior. First, evaluation of the applicant qualifications by the interviewer reveals how behavioral self-management can improve the interviewer's opinion of the candidate. Secondly, the underlying mechanics of mimicry behavior are exposed through analysis of seven nonverbal actions. Manual data analysis determines the frequency features of the actions and answers how often the actions are performed and how often they are mimicked during application interviews. Two of the seven actions are here deemed negligible due too low frequency features. Finally, machine learning is employed to analyze the data in great detail and distinguish the four behavior categories from each other. A Random Forest classifier is able to achieve 55.2% accuracy for predicting the behavior condition of the interviews while human observers reach an accuracy of 32.9%. The feature set for the classifier is reduced to 130 features with the most important features relating to the correlations between the leaning forward actions of the interview participants.
The article presents evidence for the cross-validity potential of fixed-weight (FW) versus Pareto-Optimal (PO) selection systems in biobjective selection situations where both the goals of diversity and quality are valued and the importance of the goals is undecided a priori. The article extends previous research by also studying the cross-validity potential of selection systems in the practically most important sample-to-sample cross-validity scenario. We address three research questions: (a) Do different PO systems show comparable levels of relative (i.e., proportional) achievement upon cross-validation? (b) Do PO systems achieve higher levels of relative achievement upon cross-validation than FW selection systems?, and (c) How does the achievement of PO and FW systems, in terms of adverse impact ratios and average performance of the selected applicants, evolve under cross-validation? As a key result, in case of sufficiently large applicant pools (typically 100 applicants or more), PO systems had on average a higher cross-validity potential than the corresponding FW systems. Yet, even for applicant pools as large as 500, FW systems may match the merits of PO systems and we present a straightforward procedure to decide which FW systems may offer a comparable cross-validation potential than the PO systems. (PsycInfo Database Record (c) 2022 APA, all rights reserved).