
The impact of calculator provision on the reliability and validity of a version of the Canadian Forces Aptitude Test—Problem Solving subtest was investigated in order to inform testing accommodation policy. Two hundred and fifty-four Canadian Armed Forces recruits undergoing basic training participated in the experimental research design, which consisted of a calculator and a no-calculator condition. Results supported that the convergent validity of the test was maintained in the calculator condition, as indicated by similar validity coefficients with other measures of cognitive ability in the two conditions; however, several items showed increased correct responding, and there was mixed support for criterion-related validity when calculators were provided. Additionally, the calculator condition moderated the relationship between calculator dependency and test performance, such that the relationship was positive in the calculator condition and negative in the no-calculator condition. Implications of these results with respect to accommodation policy are discussed, as well as additional research that would be beneficial in this area.
In a Behavior Description Interview (BDI), candidates are asked to describe past experiences that demonstrate skills and abilities important for the position (Janz, 1982). A recent study by Huffcutt et al. (2020) found that only around half of participants (48.1 percent) describe an experience reflecting maximal performance capability. Random mixing of maximal capability with day-to-day typical performance tendencies is problematic psychometrically because candidates are not all providing comparable information and top candidates could be overlooked. Given notable methodological concerns with Huffcutt et al.’s approach, our first purpose was to provide empirical confirmation that maximal responding in BDIs is, in fact, inconsistent. Our estimate of the proportion of maximal responding was even lower (41.3 percent), further amplifying concerns when assessment of maximal performance capability is desired (e.g., for many professional positions). The second purpose was to investigate two factors that could increase the consistency of maximal responding: rewording the main BDI question to focus directly on absolute top-end experiences (i.e., priming) and longer response length. Both were found to have significant effects. A number of directions for future research were identified, which, along with these findings, could help researchers move closer to the long-term goal of uniform description of experiences that reflect each candidate’s maximal capability (or typical tendencies if so desired).
Asynchronous video interviews (AVIs) have become increasingly popular as alternatives (or complements) to more traditional face-to-face interviews. Yet, AVI research has been largely focused on applicant reactions or behaviors, and we still know very little about what influences how applicants are rated. Importantly, because AVIs afford applicants the flexibility to record their responses from their homes, the background they choose could influence raters’ judgments. This study examines whether raters’ (N=276 Prolific respondents with prior hiring experience) initial impressions and final ratings differ if applicants record their AVIs from a home-office, a bedroom, or use background blurring settings, as well as the role played by response quality. Final interview scores were positively associated with both initial impressions and applicant response quality. Yet, background type (or the use of blurring) was not associated with initial impressions or final interview scores.
The COVID-19 pandemic highlighted the need for a massive workforce of contact tracers to help end the global pandemic. Rapidly accelerating the recruitment, selection, and training of contact tracers proved to be difficult, though, due in part to the lack of a valid, structured, and systematic approach to hiring and training contact tracers. This demonstration presents the results of the first steps in developing a systematic selection and training program: a combined (worker- and task-oriented) job analysis of the contact tracer job. Using archival records and structured interviews with 15 subject matter experts, we identified 25 unique characteristics related to successful performance as a contact tracer. We also identify which KSAOs are needed as part of selection versus those that can be trained and develop predictive hypotheses for each. Results jump-start the process of developing a systematic approach to selecting and training contact tracers to help navigate current and future public health emergencies.
nterviewers regularly make personality-related attributions in interviews, whether purposefully or not. In this study, we examined whether changing a contextual cue in a videoconference interview (the cleanliness of the room where the interviewee is located) influenced interviewers’ ratings of interviewee conscientiousness and interview performance ratings. We conducted a between-subjects experiment (N = 389) and manipulated three factors: background cleanliness (clean vs. messy) x location (office vs. home) x gender of job candidate (man vs. woman). The dependent variables were conscientiousness ratings and interview performance ratings. There was a main effect of cleanliness on conscientiousness and on interview performance ratings; these results were consistent in both the office and the home setting. The findings may inform best practices for participants in videoconference interviews.
Oddball interview questions have gained both popular and academic traction in recent years. Regardless of the intentions behind these questions, job seekers will form judgments about the employer based on its selection tactics. This paper examined the effect of oddball interview questions on organizational personality perceptions and subsequent attraction to the organization. In a time-lagged online experiment, we found organizations that asked oddball interview questions (vs. traditional interview questions) were perceived as more innovative and stylistic, which had a positive indirect effect on organizational attraction. Despite the positive effect of oddball interview questions on these organizational personality perceptions, oddball interview questions did not improve participants' overall attraction to the organization. The effect was not dependent on the job seekers’ personalities. Practitioners aimed to improve recruitment success by asking unorthodox interview questions should look elsewhere.
Intense competition for talent has led to increased organizational focus on improving how applicants perceive and respond to selection tools. Because of the recent increased use of technology in selection, we tested whether modifying aspects of videoconference interviews could improve applicant reactions. We tested two interventions—structured rapport building and question provision—with 205 applicants applying for a research assistant position. Applicants were randomly assigned to either an experimental condition (rapport or question provision) or the control condition and participated in a structured videoconference interview, followed by a survey. Structured rapport building had no significant effect on applicant reactions. However, question provision improved applicants’ perceptions of overall fairness and chance to perform—but not their reported anxiety, relative to the control condition. Question provision appears to be a simple and cost-effective intervention that could be used in a structured videoconference interview context to help to improve the applicant reactions.
This study examined current personnel selection practices in Australia including (a) the types of assessments used, (b) the factors considered when choosing assessments, and (c) the characteristics targeted in successful applicants. Participants from 68 organizations responded to a questionnaire that asked about current selection practices. Several areas where current practice deviated from research-supported best practice were identified. First, psychometric tests were used rarely: Cognitive ability tests were used by 26% of organizations and self-report questionnaires (e.g., personality inventories) by 18% of organizations. Second, when choosing assessments, the three most important considerations (in order) were the candidate experience, reducing bias, and that the assessment provides consistent scores; validity of the assessment was fourth. Finally, the most common characteristic organizations considered when selecting applicants was “culture fit.” Supplementary analyses to determine how culture fit was defined and assessed suggested there is little consistency in what it means and how it is measured.
We review the present state of research on police training in the United States, highlighting gaps in the literature, and limitations of trainings in use by local policing agencies. We focus on training content relevant to the volatile situations that are at the center of controversy, we evaluate content areas that focus on successfully navigating real-time, unpredictable, and potentially dangerous interactions, and discuss training needs in these areas. We suggest that one common response to the issue of bias—implicit bias training—lacks evidence of efficacy. Accordingly, we recommend alternative training content to address bias and discrimination. Finally, we call attention to potential barriers, including the highly charged political environment and officer resistance, that could limit the effectiveness of new training programs.
Recurrent police-public conflict suggests misalignment in desired police behavior between police and the public. We explored differences in desired police characteristics between police and members of the American public. Although racial minorities endorsed more negative attitudes of police overall, we found no meaningful differences in desired police characteristics between police and the public or between racial minority and majority participants. Second, we combined multiple criterion-related validation studies in similar jobs via meta-analyses and synthetic validity analyses to identify personality predictors of police performance dimensions. Third, we assessed base rates and adverse impact of these personality characteristics in police. Incumbent officers scored significantly lower on desired characteristics and higher on undesired characteristics than applicants. Overall, scales measuring Emotional Stability, Agreeableness, Conscientiousness, Excitability, and Skepticism seem job-relevant across samples, predictive of performance, and unlikely to cause adverse impact. Focusing on these characteristics in hiring could contribute to positive changes in police performance.
Pathology, personality, and integrity-related construct assessments have been widely used in the selection of police officers. However, the incidence of police brutality and misconduct is still concerning. The present study explored the feasibility of the assessment of cultural competence in police officers. We explored the extent to which the change to the agency’s first ever Black CEO would affect cultural competence of the officers as well as incidence of misconduct. Results showed that scores on a cultural competence factor of an in-basket simulation used for promotional assessments at a state highway patrol agency were not predictive of either supervisor-rated performance or incidence of misconduct. Whereas results showed that misconduct was not predicted by the agency’s first Black CEO, cultural competence of the officers did increase after the change in command. Practical implications for law enforcement agencies and suggestions for future research are discussed.
Calls for police reform have become frequent in the United States. Efforts to enact meaningful organizational change will require support from senior law enforcement leadership. Personnel selection for several of these positions (e.g., Sheriff) occurs via local election. Little is known, however, about the factors that influence voters’ assessment of candidates for these positions and the extent to which decision-making for personnel selection via election is influenced by the same beliefs (e.g., person-job and person-organization fit) as more traditional approaches to hiring. This study explores the extent to which voters’ perceptions of two candidates for the position of Sheriff are affected by their job-related qualifications and political affiliations – and the extent to which these perceptions influence voter behavior. Results suggest that the combination of participants’ and candidates’ political affiliations accounted for substantial incremental variance in evaluations of the candidates’ person-job fit, person-organization fit, and overall suitability for employment above-and-beyond differences in their job-related qualifications; and that participants were approximately 39% more likely to vote for the candidate with lower job-related qualifications when that candidate shared their political affiliation. Reports about the extent to which republicans and democrats value/ support police practices provide insights as to why these effects were observed.
Prepared in response to the weight and seriousness of social concerns with regard to the state and future of policing, this special issue was developed in order to feature research that examined a wide range of personnel and assessment decisions relating to policing. The focus was broad in scope, welcoming conceptual/theoretical papers, quantitative or qualitative reviews, empirical papers, and think pieces. To address the questions and areas identified in the initial call for papers, six articles are presented covering the themes of individual differences in personnel selection group composition and macro-level influences on policing, and practical recommendations and the future of policing. It is our expectation that these manuscripts will serve as a wellspring for further discussion and consideration of the role of psychology and assessment in improving police departments.
While organizations around the world recognize the importance of gender diversity and inclusion, many struggle to reach gender parity (Sneader & Yee, 2020). Particularly, women account for less than 15% of all sworn police officers (Donohue Jr, 2020). Considering signaling theory and novel research in organizational impression management, we examined the utility of various recruitment messaging techniques for attracting women job seekers to professions dominated by men, at both a consulting firm and law enforcement agency. Women evaluating consulting firm materials perceived greater behavioral integrity and were subsequently more attracted to the organization if recruitment messages included both high gender diversity signals and an explicit acknowledgement of the lack of gender diversity. With the law enforcement agency, a direct effect of the proposed interaction was identified, in that women were more attracted to police recruitment materials signaling gender diversity and explicitly acknowledging the lack of gender diversity within the agency. Materials had no adverse effect on men’s attraction. Last, research questions surrounding person-organization fit and risk propensity were analyzed to further explore the acknowledgement tactic.
Despite decades of attention paid to police reform, cases of office misconduct still continue to plague policing organizations. Assuming that organizations may still experience such officer malfeasance even when attempting to pursue best practices, we aim to explore how things can go wrong when everything else seems right. Specifically, we rely on trickle-down models of organizational justice, group engagement, and social identity to articulate how otherwise desirable organizational outcomes may produce detrimental outgroup biases. Based on our theoretical premise, we articulate specific changes that may be made to personnel systems that can avoid such officer misconduct in policing contexts.
Against best practice recommendations, interviewers prefer unstructured interviews where they are not beholden to regimentation. In cases where interviews are less structured, the interviewer typically generates his or her own set of interview questions. Even in structured interviews though, the initial interview content must be generated by someone. Thus, it is important to understand the different factors that influence what types of questions individuals generate in interview contexts. The current research aims to understand the types of interview questions individuals generate, factors that affect the quality of those questions, how skill in generating interview questions relates to skill in evaluating existing interview questions, and how individual traits relate to skill in generating interview questions. Results show that respondents who are skilled in evaluating existing interview questions are also skilled in writing interview questions from scratch, and these skills relate to general mental ability and social intelligence. Respondents generated questions that most commonly assessed applicant history and self-perceived applicant characteristics, whereas only 30% of questions generated were situational or behavioral.
Steele and Aronson (1995) showed that stereotype threat affects the test performance of stereotyped groups. A careful reading shows that threat affects test performance but does not eliminate Black–White mean score gaps. Sackett et al. (2004) reviewed characterization of this research in scholarly articles, textbooks, and popular press, and found that many mistakenly inferred that removing stereotype threats eliminated the Black– White performance gap. We examined whether the rate of mischaracterization of Steele and Aronson had decreased in the 15 years since Sackett et al. highlighted the common misinterpretation. We found that the misinterpretation rate dropped from 90.9% to 62.8% in journal articles and from 55.6% to 41.18% in textbooks, though this is only statistically significant in journal articles.
Applicant faking poses serious threats to achieving personality-based fit, negatively affecting both the worker and the organization. In articulating this “faking-is-bad” (FIB) position, Tett and Simonet (2021) identify Marcus’ (2009) self-presentation theory (SPT) as representative of the contrarian “faking-is-good” camp by its advancement of self-presentation as beneficial in hiring contexts. In this rejoinder, we address 20 of Marcus’ (2021) claims in highlighting his reliance on an outdated empiricist rendering of validity, loosely justified rejection of the negative and moralistic “faking” label, disregard for the many challenges posed by blatant forms of faking, inattention to faking research supporting the FIB position, indefensibly ambiguous constructs, and deep misunderstanding of person–workplace fit based on personality assessment. In demonstrating these and other limitations of Marcus’ critique, we firmly uphold the FIB position and clarify SPT as headed in the wrong direction.
This paper comments on Tett and Simonet’s (2021) outline of two contradictory positions on job applicants’ self-presentation on personality tests labelled “faking is bad” (FIB) versus “faking is good” (FIG). Based on self-presentation theory (Marcus, 2009) Tett and Simonet assigned to their FIG camp, I develop the ideas of (a) understanding self-presentation from the applicant’s rather than the employer’s perspective, (b) avoiding premature moral judgment on this behavior, and (c) examining consequences for the validity of applicant responses with a focus on the intended use for, and the competitive context of, selection. Conclusions include (a) that self-presentation is motivationally and morally more complex than assumed by proponents of the FIB view; (b) that its consequences for validity are ambivalent, which implies that simple credos like “FIB” or “FIG” are equally unjustified; and (c) that the label “faking” shall be abandoned from the scientific inquiry on the phenomena at hand, as it contributes to prejudiced and often erroneous conclusions.
Effective pre-hire assessments impact organizational outcomes. Recent developments in machine learning provide an opportunity for practitioners to improve upon existing scoring methods. This study compares the effectiveness of an empirically keyed scoring model with a machine learning, random forest model approach in a biodata assessment. Data was collected across two organizations. The data from the first sample (N=1,410), was used to train the model using sample sizes of 100, 300, 500, and 1,000 cases, whereas data from the second organization (N=524) was used as an external benchmark only. When using a random forest model, predictive validity rose from 0.382 to 0.412 in the first organization, while a smaller increase was seen in the second organization. It was concluded that predictive validity of biodata measures can be improved using a random forest modeling approach. Additional considerations and suggestions for future research are discussed.