Personality assessments help identify qualified job applicants when making hiring decisions and are used broadly in the organizational sciences. However, many existing personality measures are quite lengthy, and companies and researchers frequently seek ways to shorten personality scales. The current research investigated the effectiveness of a new scale-shortening method called supervised construct scoring (SCS), testing the efficacy of this method across two applied samples. Using a combination of machine learning with content validity considerations, we show that multidimensional personality scales can be significantly shortened while maintaining reliability and validity, and especially when compared to traditional shortening methods. In Study 1, we shortened a 100-item personality assessment of DeYoung et al.'s 10 facets, producing a scale 26% the original length. SCS scores exhibited strong evidence of reliability, convergence with full scale scores, and criterion-related validity. This measure, labeled the Short 10, is made freely available. In Study 2, we applied SCS to shorten an operational police personality assessment. By using SCS, we reduced test length to 25% of the original length while maintaining similar levels of reliability and criterion-related validity when predicting job performance ratings.
Cognitive ability tests are widely used in employee selection contexts, but large race and ethnic subgroup mean differences in test scores represent a major drawback to their use. We examine the potential for an item-level procedure to reduce these test score mean differences. In three data sets, differing proportions of cognitive ability test items with higher levels of difficulty or subgroup mean differences were removed from the tests. The reliabilities of these trimmed tests were then corrected back to the lengths of the original tests, and the subgroup mean differences of the trimmed tests were compared to those of the original tests. Results indicate that it is not possible to come anywhere close to eliminating subgroup differences via item trimming. The procedure may modestly reduce subgroup mean differences in test scores, with effects becoming stronger as higher proportions of items are removed from the tests. Removing items based on difficulty or subgroup differences have roughly similar impacts on test score mean differences for Black-White test taker comparisons, but results are more mixed for Hispanic-White comparisons. Our results also provide preliminary evidence that removing items on the basis of subgroup mean differences may have relatively little effect on test criterion-related validity, but the impact of removing difficult items was more mixed.
The use of personality measures to predict work-related outcomes has been of great interest over the past several decades. The present study used machine learning (ML) to examine the optimal level in the personality hierarchy to use in developing predictive algorithms. This issue was examined in a sample of incumbent police officers (N = 1,043) who completed a multifaceted personality measure and were rated on their job performance. Criterion-related validity was investigated as a function of level of operationalization in the personality hierarchy (dimensions, facets, items), scoring method (unit weighting, ordinary least-squares regression, elastic net regression), content relevance (all items vs. job-related items), and sample size (100, 200, 300, 500, 800). Results showed that empirically derived scores outperformed unit weighting across all levels of the personality hierarchy. The highest validity estimates were consistently obtained using elastic net scoring (with hyperparameter tuning resulting in solutions closer to ridge regression) at the item level, with minimal differences between ordinary least squares and elastic net for dimensions or facets with at least moderate sample sizes (N ≥ 200). An exploratory modeling approach where all item content was used did not outperform scoring when the item pool was relegated to only job-relevant personality traits. Taken together, findings suggest that personality scoring should occur at narrow operationalizations down to at least the facet level. In addition, this study demonstrated how ML can be used to not only maximize criterion-related validity but also to test long-standing theoretical problems in the organizational sciences. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
The present study is an updated survey examining individual assessor beliefs and practices related to faking in the individual assessment context. The responses from a mix of quantitative and qualitative survey questions were compared across individual assessors from the original 2005 sample (n = 77) and an updated 2020 sample (n = 78). Results suggest that single stimulus personality assessments are still the predominant form of personality assessment in use, but many individual assessors employ other types of personality assessments such as forced-choice. In 2020, individual assessors do not appear to be heavily concerned about the effects of faking on their recommendations, do not believe that a large number of candidates fake, and believe that even fewer candidates successfully fake.
Virtual reality (VR) is the three-dimensional digital representation of a real or imagined space with interactive capabilities. The application of VR for organizational training purposes has been surrounded by much fanfare; however, mixed results have been provided for the effectiveness of VR training programs, and the attributes of effective VR training programs are still unknown. To address these issues, we perform a meta-analysis of controlled experimental studies that tests the effectiveness of VR training programs. We obtain an estimate of the overall effectiveness of VR training programs, and we identify features of VR training programs that systematically produce improved results. Our meta-analytic findings support that VR training programs produce better outcomes than tested alternatives. The results also show that few moderating effects were significant. The applied display hardware, input hardware, and inclusion of game attributes had non-significant moderating effects; however, task-technology fit and aspects of the research design did influence results. We suggest that task-technology fit theory is an essential paradigm for understanding VR training programs, and no set of VR technologies is "best" across all contexts. Future research should continue studying all types of VR training programs, and authors should more strongly integrate research and theory on employee training and development.
This study examined differences in the psychometric scale properties of scores from a personality inventory as a function of cognitive ability when administered in low-stakes (incumbent employees; N = 1108) vs. high-stakes (job applicants; N = 79,339) testing contexts. Based on the idea that applicant response distortion increases trait score means and covariance and that cognitive ability is related to successful distortion, mean scores and loadings on a bifactor model were compared across groups that differed in motivation and ability to distort responses on the personality inventory. Results indicated that mean trait scores and loadings on a common factor were higher for the applicants than the incumbents. In addition, there was a positive linear trend relating cognitive ability to increases in scale scores and common factor loadings in the applicant sample; the observed changes in structure are a noted departure from what has often been found with regard to the differentiation of personality by intelligence hypothesis. Implications are discussed in terms of use of personality inventories in personnel selection contexts.
This study examined the intermediate role job satisfaction and organizational commitment play between leaders' perceived use of power and followers' performance. Based on a sample of 365 cadets at the U.S. Air Force Academy, this study found followers' job satisfaction and commitment mediated the positive relationships between their leaders' use of expert, referent, and reward power and the followers' organizational citizenship behavior. Further, while the use of legitimate or coercive power were both related negatively to followers' in-role job performance, these relationships were not mediated by the followers' job satisfaction or organizational commitment. This study then discusses the practical implications of these findings, highlights its theoretical contributions toward understanding power's direct and indirect relationships with performance in the leadership dynamic, and recommends future research avenues to leverage and build upon these findings.
Public sector testing has evolved over the past two centuries in terms of test content and test format. Public sector employment covers a very wide range of jobs and thus requires use of correspondingly wide range of assessment tools. Position classification invariably results in a class specification document that at the very least provides a starting point for construction of a job-relevant examination. Civil service examinations are seen as embodying merit principles in selection in that they are based on job-related criteria and provide a ranking of candidates in terms of their relative degree of qualification. One approach that helps meet the demand for a separate examination for each job, given the multiplicity of jobs within a given organization, is a systematic approach of analysis across jobs with a focus on the commonality among jobs. By combining more traditional tests and surveys with video based scenarios requiring preferred response a more complete picture of each candidates' overall job fitness emerges.
It is hard to argue with the central thesis of the focal article (Ruggs et al., 2016) that industrial–organizational (I-O) psychology has much to offer police departments in helping them meet their mission. As an example, the Los Angeles Police Department provides a clear statement that is representative of most police agencies: It is the mission of the Los Angeles Police Department to safeguard the lives and property of the people we serve, to reduce the incidence and fear of crime, and to enhance public safety while working with the diverse communities to improve their quality of life. Our mandate is to do so with honor and integrity, while at all times conducting ourselves with the highest ethical standards to maintain public confidence. (Los Angeles Police Foundation & LAPD, 2016)
SummaryCurrent methodologies in training evaluation studies largely employ a single method entitled random confirmatory trials, prompting several concerns. First, practitioners and researchers often analyze the effectiveness of their entire omnibus training, rather than the individual elements or identifiable components of the training program. This slows the testing of theory and development of optimal training programs. Second, a common training is typically administered to all employees within an organization or workgroup; however, certain factors may cause individualized training to be more effective. Given these concerns, the current paper presents two training evaluation methodologies to overcome these problems: the multiphase optimization strategy and sequential multiple assignment randomized trials. The multiphase optimization strategy is a method to evaluate a standard training, which emphasizes the importance of a multi‐stage training evaluation process to analyze individual training elements. In contrast, sequential multiple assignment randomized trial is used to evaluate an adaptive training that varies over time and/or trainees. These methodologies jointly overcome the problems noted earlier, and they can be integrated to address several of the key challenges facing training researchers and practitioners. Copyright © 2016 John Wiley & Sons, Ltd.
Followers’ perceptions of their leaders’ ethics have the potential to impact the way they react to the influence of these leaders. The present study of 365 U.S. Air Force Academy Cadets examined how followers’ perceptions of their leaders’ ethics moderated the relationships found between the leaders’ use of power, as conceptualized by French and Raven (Studies in social power, 1959), and the followers’ contextual performance. Our results indicated that leaders’ use of expert, referent, and reward power was associated with higher levels of organizational citizenship behaviors (OCBs) among their followers when the followers perceived these leaders to be more ethical. Moreover, when followers perceived their leaders to be less ethical, these followers reported lower levels of OCBs when their leaders’ utilized referent power. Practical implications, limitations, and future research are also discussed.
A popular means of predicting multitasking success is measuring individuals’ polychronicity, which is the preference for multitasking instead of focusing on single tasks in a linear fashion. An experimental investigation of the potential problems of using polychronicity measures for personnel selection purposes was conducted. The present study discovered that measures of polychronicity are seriously flawed for use in personnel selection. Due to the ease with which the intention behind the questions can be determined, individuals are able to distort their responses to match what the job requires. Practical issues and implications for personnel selection are discussed.