
ABSTRACT Organizations sometimes embed humor into recruitment materials such as job postings to try to attract talent. Our research investigates the efficacy of humorous job postings. On the one hand, humorous job postings may be beneficial because the use of humor may signal positive organizational attributes. On the other hand, humorous job postings could also inadvertently convey a lack of seriousness in the organization's HR practices. Integrating signaling theory with benign violation theory, we develop and test a theoretical model of the competing signals that humor in job postings might send. We conducted three experiments using single stimulus approaches (Studies 1 and 2) and stimulus sampling approaches (Study 3). Studies 1 and 2 showed a negative effect of humor on job attractiveness, whereas Study 3 showed no statistically significant effect of humor on job attractiveness. Consistently across studies, humorous job postings led potential job seekers to expect the organization to display lower levels of procedural justice. These findings suggest exercising caution when using humor in recruitment materials.
ABSTRACT The widespread adoption of large language model (LLM) chatbots has raised concerns about their potential misuse in high‐stakes employment contexts, including cheating on pre‐employment assessments. Although prior research demonstrates that LLMs can achieve desirable scores on cognitive and personality tests, little evidence exists regarding whether their availability has produced measurable score inflation in real‐world selection settings. The present study examines whether mean scores on commonly used pre‐employment assessments shifted following the public release of ChatGPT in November 2022. Using large archival samples of job applicants, we analyzed trends in scores on two cognitive assessments and one personality assessment. We also examined age as a moderator given evidence that younger individuals are more frequent users of LLM technologies. Regression and time‐trend analyses indicated statistically significant but practically negligible changes in assessment scores following ChatGPT's release. Across cognitive and personality measures, effect sizes were near zero and inconsistent in direction, providing no meaningful support for widespread LLM‐assisted cheating. Overall, results suggest that despite concerns about LLM misuse, there is currently little evidence of large‐scale score inflation in operational pre‐employment testing contexts.
ABSTRACT In their “Provocation Article,” König et al. (2026) argue for well‐being as an explicit criterion in personnel selection and propose selecting for applicants with characteristics predictive of well‐being. We argue that if the goal is safeguarding employee well‐being, selecting for characteristics predictive of others' well‐being is far more critical than those predictive of one's own well‐being. We believe that safeguarding others' well‐being is applicable to all employees, not just leaders, and that there are other types of interpersonal conduct beyond organizational citizenship behaviors (OCBs) and counterproductive work behaviors (CWBs) that can affect others' psychological safety, trust, and sense of belonging and should be examined. We conclude with recommendations for implementing selection for those who will protect and promote the well‐being of others and the long‐term benefits of this emphasis for worker growth and development.
ABSTRACT Research on big data technologies in employee selection has expanded rapidly but remains fragmented. This study presents a systematic literature review of empirical applications of big data technologies in employee selection. The study analyzed 50 publications obtained from the Scopus and Web of Science databases. Specifically, this review synthesizes prior research on the theoretical foundations of big data technologies in employee selection, the data sources used, the analytic approaches applied, and their contributions to employee selection research and practice at theoretical, methodological, and practical levels. Results show limited theoretical anchoring, with only minor links to established theories such as person‐environment fit theory, trait theory, needs theory, trait activation theory, and Brunswik's lens model. Most empirical applications are based on field data collected from real selection or human resource management processes, with fewer studies relying on lab data, such as mock interviews or simulated selection tasks, or on mixed data sources. Analytically, researchers have employed a broad range of big data technologies, including supervised machine learning, unsupervised machine learning, natural language processing embeddings, n‐gram text mining, large language models, and others, with increasingly multimodal and stage‐specific applications across the selection process. Overall, the empirical application of big data technologies in employee selection is promising, but remains contingent on theory and method alignment, transparent targets and labels, and ongoing auditing.
ABSTRACT A major concern with personality assessment in high‐stakes settings is test‐takers misrepresenting (“faking”) their personalities. Although test developers have adopted difficult‐to‐fake forced‐choice (FC) formats, generative artificial intelligence (AI) now offers test‐takers a highly sophisticated “test‐coach,” and yet it is unclear how resilient FC formats are to faking by AI‐generated response sets. This investigation examined the faking resistance of single‐statement (SS), forced‐choice IRT (FC‐IRT), and forced‐choice CTT (FC‐CTT) versions of the HEXACO inventory. 219 Accounting/Finance professionals completed the three assessments under honest and applicant conditions, the latter for a role where Honesty‐Humility and Conscientiousness were target traits. Three frontier AI models were repeatedly prompted with analogous instructions to simulate AI‐assisted faking. We found that two AI models consistently outscored human test‐takers under applicant conditions across all formats. Among the human‐generated responses, the FC‐CTT format showed less inflation than FC‐IRT for Honesty‐Humility and better rank‐order stability in applicant conditions, but this pattern did not extend to Conscientiousness. Neither FC format demonstrated clear improvements in convergent or criterion‐related validity. Findings challenge the previous assumption that FC measurement approaches can fully mitigate human‐ or AI‐assisted faking.
ABSTRACT Structured online reference checks (SORCs) are predictive of job performance and cost effective given that they can be completed quickly and require no “live” interview by HR personnel. Additionally, they appear to reduce gender bias—a problem identified with other reference checking methods. However, additional research is needed to examine possible threats to this method's effectiveness. One such threat is leniency bias, which has been leveled as a general criticism of reference checking by both scholars and practitioners. The present research, involving eight experimental studies, represents the first systematic investigation of leniency bias in SORCs. The results, involving data from over 1700 full‐time managers in the U.K. and U.S., indicate that: (a) quantitative ratings of direct reports' performance, organizational citizenship behavior, and counterproductive work behavior are more lenient in a SORC condition than when reporting private assessments; (b) vicarious self‐promotion (i.e., managers' presenting themselves positively through lenient ratings of their employees) helps to account for this bias; and (c) leniency is reduced by providing instructions that encourage the referee to take the perspective of the hiring manager rather than the job candidate when completing the reference check. While this research suggests that SORCs are likely to suffer from leniency bias, it also identifies a promising and relatively simple method to mitigate leniency. Implications for future research and practice are explored.
ABSTRACT Assessment centers (ACs) could provide behavioral foundations for trait assessment but demand substantial resources, especially the labor‐intensive human rating process. As organizations adopt large language models (LLMs) for automated assessment, we establish initial psychometric benchmarks for LLM‐based trait inference from AC behavior. We processed recordings from 363 participants in brief simulated online AC group discussions through an automated pipeline: Whisper (speech recognition), Pyannote (speaker diarization), and Gemini 2.0 Flash (trait inference across the Big Five, narcissism, and intelligence). As a human benchmark, eight trained raters (four aggregated per transcript) independently inferred the same traits from a subsample of 100 group discussion transcripts. Aggregating multiple LLM runs yielded high reliability across traits (ICC(3,4) = 0.75–0.92). LLM convergent validity with trait tests was modest but above chance for all traits except conscientiousness, and varied by trait observability: openness and intelligence ( r = 0.20), narcissism, extraversion, and agreeableness ( r = 0.12–0.16), emotional stability and conscientiousness ( r = 0.07–0.11). These findings lend support to AC behavior as a basis for measuring personality even under minimal information, with these validities likely a lower bound on what richer inputs would yield. Most notably, an LLM could infer personality from this behavior, with convergent validities descriptively higher than the human benchmark for most traits. Overall, constrained by discriminant validity, gender‐related adverse impact, and input dependencies (amount and clarity), the LLM trait inferences remained insufficient for standalone selection but can be produced at negligible marginal cost, positioning LLMs as a potential low‐cost contributor to multi‐source assessment pipelines.
Based on cognitive consistency theory, this study developed a model linking conscientiousness to employees' performance. A three-wave survey sample from 193 supervisor-subordinate dyads was used to test this model. After controlling the effects of individuals' demographics and proactive personality, HLM analyses revealed that conscientiousness is positively associated with organization-based self-esteem. Organization-based self-esteem partially mediated the relationship of conscientiousness with job performance and fully mediated the relationship of conscientiousness with organizational citizenship behavior-altruism. However, organization-based self-esteem did not mediate the relationship between conscientiousness and organizational citizenship behavior-voice. Additionally, job autonomy moderated the relationship between conscientiousness and organization-based self-esteem, and this relationship was stronger for employees with higher rather than lower job autonomy. Moreover, with controlling the influences of individuals' demographics and proactive personality, the analyses of moderated mediation revealed the indirect effects of conscientiousness on job performance, organizational citizenship behavior-altruism, and organizational citizenship behavior-voice via organization-based self-esteem were stronger for employees with higher rather than lower job autonomy.
The rapid rise of generative artificial intelligence (GenAI) poses new challenges for the validity and fairness of pre-hire employment assessments. Across two applied studies using large applicant samples completing a standardized pre-hire assessment (which includes both cognitive and non-cognitive components), we examined the self-reported incidence of GenAI assistance, and the impact of warning statements designed to deter such behavior. In Study 1 (N = 5675), conducted in Q3 2024, fewer than 3% of applicants reported using GenAI, though up to 19% reported using GenAI in combination with algorithmic resources (e.g., search engines). All three warning statements (consequences, educational, and reasoning) reduced reported use relative to the control condition, with limited evidence favoring the consequences-based warning. Study 2 (N = 3356), conducted in Q3 2025, focused exclusively on the consequences-based warning. Self-reported GenAI use increased from 2024 to 2025. Warnings significantly reduced the incidence of GenAI use but did not alter motivations, contexts, perceived effectiveness, or applicant reactions. Finally, analyses of potential job fit scores indicated that GenAI use per se was not systematically related to potential job fit, though stress-related motivations for GenAI use showed a small negative association with lower potential job fit. These findings highlight both the likely growing prevalence of GenAI in selection contexts and the utility of warnings as a potential deterrence strategy.
To clarify contradictions in job choice theory, we offer a three-tier compensability model of job choice that explains when money can remedy workplace deficits and when money fails. While wage differential frameworks assume pay universally substitutes for negative job conditions, mounting evidence shows financial incentives failing to remediate foundational psychological needs. Our model integrates behavioral decision theory and image theory to establish: (1) non-compensable Career Development attributes (recognition, leadership, organizational justice) that resist monetary substitution; (2) marginally compensable Work Sustainability attributes (work-life balance, psychological safety) permitting constrained trade-offs; and (3) compensable Task-Transactional attributes (job demands, role clarity) operating within standard utility frameworks. Through an adaptive conjoint experiment with 798 U.S. professionals making 7,182 choices, we demonstrate that non-compensable violations trigger categorical rejection (rejection > 90% at maximum pay) while the compensable tier exceeds 50% acceptance through premium pay. Moreover, we demonstrate that distinct mechanisms underlie willingness to pay and willingness to accept. These findings challenge the prevailing fungibility assumption in job choice. Compensation is not a universal solvent but a precision tool, effective only for specific transactional problems after foundational psychological needs are met. Our results suggest organizations should first invest in their Career Development infrastructure; only then can monetary premiums for transactional issues succeed.
This study investigates the effect of the stimulus format (text- vs. VR-based) of a situational judgment test (SJT) on test performance, while also investigating the mediating role of test motivation and the moderating role of test-takers' age. Results of a two-condition (SJT stimulus format: text- vs. VR-based) within-subjects design among 121 participants showed no differences in test motivation between a VR- and a text-based format. Further, results showed slightly higher SJT scores for the text-based stimulus format than the VR-based format, particularly for middle-aged and older test-takers. Finally, test motivation did not mediate the relationship between stimulus format and test performance. While existing knowledge is limited, this study provides little reason for concern regarding the differential effects of VR- versus text-based stimuli on motivation and performance across age groups, paving the way for further research on technological innovations in SJTs.
Anonymous application procedures (AAPs), which hide applicants' demographic and other identifying information, aim to reduce discrimination during preselection. However, research about applicants' perceptions of such procedures is scarce. Accordingly, we compared perceptions of traditional resum & eacute;s, anonymized resum & eacute;s, and standardized application forms. Potential applicants preferred anonymized resum & eacute;s followed by standardized forms, with traditional resum & eacute;s being least favored. Participants' gender and migration background did not emerge as moderators. In a second study, explaining that standardized application forms are used to enhance equal employment opportunities improved perceptions, whereas emphasizing job relatedness or combining both explanations was less effective. Organizations might benefit from adopting AAPs to signal their commitment to diversity and improve applicant reactions. Limitations and potential areas for future research are discussed.
Personnel selection research demonstrates that structured methods reliably predict job performance. Yet this focus has narrowed the criterion space by equating performance at entry with long-term effectiveness. Drawing on Job Demands-Resources theory, I argue that sustainable performance unfolds within dynamic demand-resource systems that simultaneously shape employee well-being. Building on Konig et al. call to integrate well-being into selection, I propose a Fit-for-Sustainability model that shifts attention from static performance prediction to sustainable effectiveness. The model distinguishes demands-abilities fit at entry from needs-supplies fit, which must be secured through work design, and incorporates regulatory resources-such as proactive personality, self-control, and learning goal orientation-that enable employees to recalibrate fit as demands evolve. Organizational resources provide the regulatory latitude necessary for proactive adjustment. Together, these elements reposition selection as the starting point of continuous demand-resource optimization, advancing a theory of selection for sustainable performance and well-being over time.
What motivates potential applicants to apply to an organization and speak favorably about it are key recruitment questions. To answer these, prior studies have primarily focused on the content of the employer brand, i.e., instrumental attributes (e.g., salary, promotion, job security) and symbolic attributes (e.g., employer warmth, competence) as drivers of applicant attraction. However, this perspective overlooks the fundamental process-based characteristics of the employer brand, which form the crux of employer branding and HRM system strength theory. The process attributes refer to branding-related aspects of the employer brand, i.e., ensuring the organization's distinctiveness against competitors, consistency across multiple images/touchpoints, and shared understanding of the employer image. Accordingly, this study investigates the role and added value of employer brand process attributes (distinctiveness, consistency, and consensus) versus content (instrumental and symbolic attributes) in understanding potential applicants' attraction (job pursuit and word-of-mouth intentions). Additionally, we examine whether the impact of the three process attributes is stronger for those applicants who engage in higher job search metacognitive activities, further exploring applicants' motivations from a job search angle. A two-time survey of 191 US job applicants showed that organizations that are distinctive and consistent as an employer elicit higher job pursuit and word-of-mouth intentions. Moreover, process attributes explain significant incremental value beyond content attributes. Limited moderation effects emerged, with consistency particularly relevant for applicants with high metacognitive activities.
In this reply to the commentaries by Mirowska (2025), Hickman (2025), and Holtrop and Bronzwaer (2026), we expand on our initial provocation article regarding the effects of candidate use of Generative AI (GenAI) in personnel selection (Lievens and Dunlop, 2025). First, we update the discussion by highlighting recent technological developments (agentic AI and AI-integrated wearables) that accelerate the threat of candidate GenAI use as a substitute for candidate effort. Second, we clarify and build upon our original arguments concerning the effects of candidate GenAI use on construct-related validity and subgroup differences, introducing the concept of "GenAI literacy" as a potential confounding construct. In doing so, we elaborate on the concept of AI-enabled assessment designs. Finally, we integrate the insights from the three commentaries with our own thinking to introduce the FAIR framework (Forbid, Advise, Insulate, Reimagine), which should help employers navigate the complex landscape of candidate GenAI use. The FAIR framework distinguishes between strategies aimed at preventing GenAI misuse and those designed to embrace and integrate GenAI into assessment processes. We conclude that the future of selection lies not in banning candidate GenAI use entirely. Instead, we argue for a strategic shift toward reimagining assessments for an AI-augmented world.
This short communication reports the development of a competency model for autistic employment, derived directly from autistic employees' accounts of effective workplace performance. Critical Incident Technique interviews were conducted with 15 autistic professionals employed in large corporate and public-sector organizations, yielding 67 critical incidents and approximately 130 behavioral statements. Interview data were thematically coded against an established competency dictionary and refined through consensus, resulting in 12 competencies organized into four clusters: social, motivational, cognitive, and personal, with 89 associated behavioral indicators. The draft model was reviewed and refined through follow-up focus groups with participants to assess clarity, accuracy, and comprehensiveness. Although several competency labels overlap with generic models, their behavioral expressions reflect autistic employees' distinctive strategies for managing communication demands, cognitive load, and workplace expectations. The model is presented as a set of hypothesized competencies intended for subsequent quantitative validation. Planned next steps include psychometric testing with autistic and non-autistic employees to inform the development of more transparent, function-focused selection and management practices.
General mental ability (GMA) is an important predictor of training performance. This makes GMA highly relevant when selecting applicants for vocational training. However, the value of specific cognitive abilities among other indicators of knowledge, skills, abilities, and other characteristics (KSAOs) alongside GMA in predicting training performance is currently being debated. This study aims to contribute to this debate. Results from two methods are compared, incremental validity analysis and relative importance analysis, using two types of training performance outcomes: academic and scenario-based. The analyses were performed using a dataset of German police officer candidates (N = 1218), which encompassed data from the applications process, GMA, specific cognitive abilities, and non-cognitive KSAOs (including a situational judgment test, SJT, and physical tests and a structured interview), as well as performance data from several examinations over the course of a 3-year training program. Findings from both analyses show that GMA and verbal specific cognitive abilities are good predictors of academic training performance. However, findings also show that other KSAO indicators, such as an SJT, physical tests, and a structured interview, outperformed the cognitive measures in predicting scenario-based training performance. Together, the findings add nuance to the current debate on the value of cognitive abilities and non-cognitive KSAOs for the prediction of training performance.
Humour in interviews can increase the affiliation between interviewees and interviewers. Technology-mediated interviews may limit the perceived opportunities and effectiveness of humour attempts. A total of 271 participants (from Prolific) were randomly assigned to one of three conditions describing either a face-to-face, videoconference, or asynchronous video interview scenario. The effects on perceived effectiveness and intentions to use humour for reasons that could be described as self-focused or other-focused were assessed. Perceptions of effectiveness were not significantly different across interview media. The likelihood of using self-focused humour did not differ, but the likelihood of using humour for other-focused reasons was lower in asynchronous video interviews. Social presence was identified as a mediator of the relationship between interview media and likelihood of humour use. This study offers insight into the antecedents of social dynamics in different interview formats.
Distinguishing between subclinical and clinical personality assessments is essential for their appropriate use across organizational, research, and clinical contexts. Although some workplace measures assess maladaptive tendencies, the extent to which they overlap with clinical instruments remains unclear. Misinterpreting this overlap can result in the inappropriate clinicalization of workplace measures and the misuse of clinical instruments in personnel selection. This study compared a workplace measure, the Hogan Development Survey (HDS), with a clinical instrument, the Personality Inventory for DSM-5 (PID-5). Across multiple psychometric criteria, the two measures showed meaningful but nonredundant overlap. The HDS provided limited information at extreme trait levels and more strongly predicted work outcomes, whereas the PID-5 was more informative at the pathological end of the continuum and more strongly associated with mental health indicators. These findings clarify the boundaries between subclinical and clinical personality assessment and inform evaluations of measures with overlapping content but different applications.
In two studies, we explored the impact of body-related nonverbal information (BNI) on performance ratings in technology-mediated interviews (TMIs). Using video frames that varied in size (head and shoulders vs. head to knee), participants (N = 203 and 238) evaluated recorded interview responses. Larger frames increased the perceived extent of BNI but not its perceived usefulness. Perceived usefulness, however, showed stronger associations with performance ratings, which were (partially) mediated by social presence. These effects were consistent for varying degrees of dynamic nonverbal cues and across both samples. The findings highlight the relevance of subjective perceptions and judgments of nonverbal information in TMIs. To reduce bias, organizations should use similar video frames across applicants and train interviewers to prioritize verbal responses.