Social media (SM) assessments are a new type of selection procedure. Their frequency of use in hireability decision-making is substantial, encompassing both utilitarian (e.g., LinkedIn) and hedonic (e.g., Facebook) platforms. Yet, there is limited evidence of SM assessment validity (prediction of job performance). This mismatch (use vs. lack of validity evidence) motivated our study of why decision-makers embrace such information. We used (i) the unified technology acceptance model and (ii) organizational justice theory to list positively framed facets of SM assessments (e.g., provide information about applicant skills, used consistently, fun to use). Decision-makers were assigned to one of three conditions—information retrieved from Facebook, LinkedIn, or structured interviews (as a baseline)—and were asked to rate how well these explanatory facets characterized information from those SM (and comparison) sources. Empirical composites (principal components) of these facets were developed. We found that all composites were positively related to use of SM across all conditions. These findings are disconcerting given the equivocal validity of SM assessments (and potential for much extraneous, non-job-related information). We urge continued research on reasons for frequent use of SM assessments, as well as implications of that use. Such research for other predictors is also encouraged.
Situational judgment tests (SJTs) are popular assessment approaches that present scenarios describing situations that one may experience in a job. Due to its long history and cross-disciplinary nature, today’s SJT literature is quite fragmented. In this integrative review, we start by systematically taking stock and synthesizing the SJT literature from the different scientific disciplines via bibliometric techniques on 524 unique documents. We identify six literature clusters (i.e., SJTs in the medical sciences, SJTs in personnel selection, methodological issues and SJTs for specific constructs, SJTs to assess emotional intelligence and related constructs, technological advances in SJTs, SJTs for teacher assessment and development) that correspond to academic disciplines and research streams within them. We also identify current trends in SJT research by examining the clusters formed by a recent subset of the SJT literature. We then build on the bibliometric analysis by categorizing the identified themes in an organizing framework with two fundamental dimensions: the main purpose of a study (i.e., conceptual understanding, prediction, other [e.g., understanding mean group differences, applicant reactions]) and its research focus (i.e., SJTs holistically, content, and design and methods). Finally, on the basis of this framework, we provide recommendations to encourage greater knowledge sharing between scientific disciplines. In addition, we outline an agenda for future research in terms of four broad directions: SJT theory, SJT constructs, SJT design and methods, and SJT application domains.
Cognitive ability tests are widely used in employee selection contexts, but large race and ethnic subgroup mean differences in test scores represent a major drawback to their use. We examine the potential for an item-level procedure to reduce these test score mean differences. In three data sets, differing proportions of cognitive ability test items with higher levels of difficulty or subgroup mean differences were removed from the tests. The reliabilities of these trimmed tests were then corrected back to the lengths of the original tests, and the subgroup mean differences of the trimmed tests were compared to those of the original tests. Results indicate that it is not possible to come anywhere close to eliminating subgroup differences via item trimming. The procedure may modestly reduce subgroup mean differences in test scores, with effects becoming stronger as higher proportions of items are removed from the tests. Removing items based on difficulty or subgroup differences have roughly similar impacts on test score mean differences for Black-White test taker comparisons, but results are more mixed for Hispanic-White comparisons. Our results also provide preliminary evidence that removing items on the basis of subgroup mean differences may have relatively little effect on test criterion-related validity, but the impact of removing difficult items was more mixed.
Summary Questionable research practices (QRPs) among researchers have been a source of concern in many fields of study. QRPs are often used to enhance the probability of achieving statistical significance which affects the likelihood of a paper being published. Using a sample of researchers from 10 top research‐productive management programs, we compared hypotheses tested in dissertations to those tested in journal articles derived from those dissertations to draw inferences concerning the extent of engagement in QRPs. Results indicated that QRPs related to changes in sample size and covariates were associated with unsupported dissertation hypotheses becoming supported in journal articles. Researchers also tended to exclude unsupported dissertation hypotheses from journal articles. Likewise, results suggested that many article hypotheses may have been created after the results were known (i.e., HARKed). Articles from prestigious journals contained a higher percentage of potentially HARKed hypotheses than those from less well‐regarded journals. Finally, articles published in prestigious journals were associated with more QRP usage than less prestigious journals. QRPs increase in the percentage of supported hypotheses and result in effect sizes that likely overestimate population parameters. As such, results reported in articles published in our most prestigious journals may be less credible than previously believed.
Experiments in male rodents demonstrate that sensitivity to the organizational effects of steroid hormones decreases across the pubertal window, with earlier androgen exposure leading to greater masculinization of the brain and behavior. Similarly, some research suggests the timing of peripubertal exposure to sex steroids influences aspects of human psychology, including visuospatial cognition. However, prior studies have been limited by small samples and/or imprecise measures of pubertal timing. We conducted 4 studies to clarify whether the timing of peripubertal hormone exposure predicts performance on male-typed tests of spatial cognition in adulthood. In Studies 1 (n = 1095) and 2 (n = 173), we investigated associations between recalled pubertal age and spatial cognition in typically developing men, controlling for current testosterone levels in Study 2. In Study 3 (n = 51), we examined the relationship between spatial performance and the age at which peripubertal hormone replacement therapy was initiated in a sample of men with Isolated GnRH Deficiency. Across Studies 1-3, effect size estimates for the relationship between spatial performance and pubertal timing ranged from. -0.04 and -0.27, and spatial performance was unrelated to salivary testosterone in Study 2. In Study 4, we conducted two meta-analyses of Studies 1-3 and four previously published studies. The first meta-analysis was conducted on correlations between spatial performance and measures of the absolute age of pubertal timing, and the second replaced those correlations with correlations between spatial performance and measures of relative pubertal timing where available. Point estimates for correlations between pubertal timing and spatial cognition were -0.15 and -0.12 (both p < 0.001) in the first and second meta-analyses, respectively. These associations were robust to the exclusion of any individual study. Our results suggest that, for some aspects of neural development, sensitivity to gonadal hormones declines across puberty, with earlier pubertal hormone exposure predicting greater sex-typicality in psychological phenotypes in adulthood. These results shed light on the processes of behavioral and brain organization and have implications for the treatment of IGD and other conditions wherein pubertal timing is pharmacologically manipulated.
Although cognitive ability is generally considered the best predictor of job performance (e.g., Schmidt & Hunter, 2004), some individuals who perform well on cognitive ability assessments may not ultimately have high job performance. For instance, there is evidence to suggest that expressions of certain personality disorders have a negative effect on job performance (e.g., Moscoso & Salagado, 2004). Yet, many psychologists now consider personality to exist on a continuum, such that individuals may have some symptoms associated with personality disorders (and decreased job performance), but not a clinically diagnosable disorder (e.g., De Fruyt & Salagado, 2003; Trull & Durrett, 2005). Therefore, as part of a selection process, it may be beneficial to identify individuals that are likely to exhibit behaviors associated with subclinical levels of personality disorders in the workplace. Situational judgement tests (SJTs) are particularly well suited to assessing subclinical levels of personality as part of an employment selection process. First, SJTs are low fidelity simulations (e.g., Motowidlo, Dunnette, & Carter, 1990) that use workplace-specific scenarios and response options. Second, SJTs with knowledge instructions (e.g., how effective is each behavior likely to be) are less susceptible to faking than typical personality inventories (Nguyen,
Social media assessments (SMAs) for hiring job applicants have received a great deal of coverage from the business press as well as somewhat modest coverage from human resource management researchers. At the same time, there is no systematic study of how often SMAs are actually used for hiring by recruiters in areas such as human resources, management information systems, and education. We meta- analytically reviewed the use of SMAs in hiring. Our meta-analysis showed that the overall mean percentage of SMA use by recruiters was 56%, including substantial use of Facebook and LinkedIn, as well as some use of Twitter and Google. Most recent estimates suggested that the mean percentage of use of SMAs by recruiters was 72%. Thus, SMAs appear to be used relatively frequently as forms of individual assessment. Yet, SMAs continue to be associated with relatively few studies in human resource management.
Meta-analytic reviews are considered the primary means for generating cumulative scientific knowledge and their results are often used by practitioners to inform evidence-based practice. However, the robustness of meta-analytic summary estimates is rarely examined. Consequently, the results of published meta-analyses may be misestimated and, thus, untrustworthy. Outliers can inflate the amount of residual heterogeneity in meta-analytic datasets, which can lead to biased meta-analytic and publication bias analysis results. We introduce a tool that will help researchers to conduct a meta-analysis that adheres to recommended reporting standards and best practices. Specifically, we describe and demonstrate a comprehensive sensitivity analysis tool that can assist in accounting for outlier-induced heterogeneity when performing a meta-analysis and the corresponding publication bias analyses. In addition, we use a dataset from a recently published meta-analysis to illustrate the functionality of the comprehensiv...
The focal article (Grand et al., 2018) addresses one of the most important issues across virtually all areas of science (Goldstein, 2010): the trustworthiness and credibility of a scientific discipline. Once these attributes are lost, it is difficult to regain them within a reasonable time frame, if ever. In contrast to previous articles on this topic (e.g., Kepes & McDaniel, 2013), the authors of the focal article provide a detailed review of the stakeholders surrounding industrial and organizational (I-O) psychology, including their potential effect on the robustness and trustworthiness of our scientific discipline. In essence, the focal article describes I-O psychology's ecosystem responsible for fostering robust and credible science. The authors should be commended for their comprehensive undertaking, and we have no substantive disagreements. However, implicitly, as with most articles on this vital topic, the focal article tends to take a bottom-up approach to decision making and change. The bottom-up approach is an emergent process where the individuals involved in the day-to-day activities are primarily responsible for the decision-making process and resulting change (Kindler, 1979). Thus, changes resulting from this process are incremental and typically involve making minor adjustments to existing processes (Bartunek & Moch, 1987).
In recent years, situational judgment tests (SJTs) have made strong inroads in assessment practices. Despite the importance of scoring for the validity of SJTs, little attention has been paid to different SJT scoring methods. This study investigated the influence of scoring methods on the criterion-related validity of SJTs. We examined five different consensus scoring methods (i.e., raw, standardized, dichotomous, mode, and proportion scoring) and several integrated scoring methods for scoring the same SJT. Results showed that one of the most popular scoring approaches (raw consensus scoring) is associated with an extreme response tendency and yields the lowest scale validity of all scoring approaches examined. Moreover, the mean item validity of midrange items was good only when they were scored by the mode consensus method. Thus, this study extends previous work (McDaniel et al., 2011) by deepening our understanding of how different scoring methods improve the validities of SJTs. Our findings suggest that using scoring methods that control the influence of extreme response tendency on the scores of SJTs yields higher validities. Finally, this study is the first to suggest that scoring SJTs with integrated methods yielded higher mean item validities than using any single method.
Tipping represents a form of compensation valued at over $50 billion a year in the United States alone. Tipping can be used as an incentive mechanism to reduce a principal-agent problem. An agency problem occurs when the interests of a principal and agent are misaligned, and it is challenging for the principal to monitor or control the activities of the agent. However, past research has been limited in the investigation of the extent to which tipping is effective at addressing this problem. Following an examination of 74 independent studies with 12,271 individuals, meta-analytic results indicate that there is a small, positive relation between service quality and percentage of a bill tipped ((rho) over bar = .15 without outliers). Yet, in support of the idea behind tipping, relative weights analyses illustrate that service qualitywas a stronger predictor of percentage of the bill tipped than food quality, frequency of patronage, and dining party size. Evidence also suggests that racial minority servers tend to be tipped less than White servers (Cohen's d = .17), and women tend to be tipped more than men (Cohen's d = .15). Still, given the magnitude of the effect, one might question if tipping is an effective compensation practice to reduce the principal-agent problem. We discuss theoretical and practical implications for future research.
Meta-analytic studies are the primary way for systematically synthesizing quantitative research findings to cumulate knowledge. As such, they have substantial influence on research and practice. Recently, however, the robustness of results from some meta-analytic studies in management has been questioned. Despite this, very few studies assess the presence and impact of publication bias (PB) and outliers, two factors influencing the non-robustness of meta-analytic results. In this study, we use a comprehensive sensitivity analysis approach to reexamine datasets from nine meta-analyses of correlation coefficients published in Psychological Bulletin that were categorized in the area of learning, behavior, and performance. We reexamined 123 distributions from these nine meta-analytic studies. Our results indicate that 88% of the meta-analytic results reported in the nine meta-analytic studies are unlikely to be robust. The degree of the non-robustness was classified as being ‘severe’ (i.e., > 40%) in 78% of the meta-analytic distributions. These results suggest that most of the meta-analytic results and associated conclusions and recommendations in this area may not be trustworthy. This adds to a growing body of evidence suggesting that our research practices may need to be revised to improve the trustworthiness of our cumulative knowledge in management and the social sciences.
In an effort to improve the predictive validity of a situational judgment test (SJT), this research explores SJT factor structure using multidimensional item response theory (MIRT). We found that using MIRT approaches to scoring SJTs increased the predictive validity of the SJT and that the SJT derived scores accounted for more variance explained than either personality or GMA.
Despite limited and conflicting support for Cognitive Evaluation Theory (CET), many researchers and practitioners still believe that extrinsic rewards can be harmful because they undermine intrinsic motivation. In this study, we utilize comprehensive sensitivity analysis techniques to determine the trustworthiness of our cumulative knowledge of the relation between extrinsic rewards and intrinsic motivation. These techniques help to determine the extent to which outliers and publication bias (PB) impacted prior meta- analytic results and conclusions. Results of our sensitivity analyses on 40 distributions from two meta-analyses indicated that many of the meta-analytic mean estimates provided in the originally published meta-analyses tend to be severely misestimated. Of the 40 distributions, all had meta-analytic mean estimates that were deemed non-robust. These results suggest that there seems to be no trustworthy meta-analytic support for CET or the notion that extrinsic rewards undermine intrinsic motiv...
Carotid atherosclerosis may be associated with cognitive decline in elderly stroke-free non-demented individuals. However the studies are few, especially amongst adults in midlife. Rarely was socioeconomic status adequately accounted for.In a cross-sectional examination of the population-based Jerusalem Lipid Research Clinic (LRC) cohort, carotid intima-media thickness (IMT), plaque and volumetric flow were determined by ultrasound in 507 cohort members, aged 48–52. At the same visit, global cognitive function and its five specific component domains were assessed with a NeuroTrax computerised test battery. Multivariable linear regression and logistic models were applied.In sex-adjusted, but not in multivariable-adjusted linear regression models, IMT and plaque were significantly inversely associated with measures of cognition. However, in binary logistic models adjusted for sex, age, education, childhood and adult SES, and multiple cardiovascular risk factors, IMT was associated with low-ranked (lowest fifth) global cognition (OR per 0.1 mm, 1.25, 95%CI 1.01–1.55, p = 0.044); this association was confirmed in multinomial logistic modelling. The presence of atheromatous plaque was associated with elevated risk for low-ranked global cognitive function in midlife (OR, 1.93, 95%CI, 1.04–3.59, p = 0.037). Common carotid volumetric flow was not associated with cognitive function, and adjustment for volumetric flow did not materially affect the associations of IMT or plaque with global cognition.In this population-based cohort in midlife, subclinical carotid atherosclerosis measured as higher IMT and the presence of atherosclerotic plaque was associated with low-ranked global cognitive function. In light of the absence of an independent association of carotid volumetric flow with cognition, and although reverse causation cannot be excluded in this cross-sectional study, we infer that carotid atherosclerosis in midlife may be a marker of intracranial atherosclerosis and possible resulting lower cognitive function.
The construct validity of situational judgment tests (SJTs) is a “hot mess.” The suggestions of Lievens and Motowidlo (2016) concerning a strategy to make the constructs assessed by an SJT more “clear and explicit” (p. 5) are worthy of serious consideration. In this commentary, we highlight two challenges that will likely need to be addressed before one can develop SJTs with clear and explicit constructs. We also offer critiques of four positions presented by Lievens and Motowidlo that are not well supported by evidence.
Improvements in stable, or dispositional, mindfulness are often assumed to accrue from mindfulness training and to account for many of its beneficial effects. However, research examining these assumptions has produced mixed findings, and the relation between dispositional mindfulness and mindfulness training is actively debated. A comprehensive meta-analysis was conducted on randomized controlled trials (RCTs) of mindfulness training published from 2003-2014 to investigate whether (a) different self-reported mindfulness scale dimensions change as a result of mindfulness training, (b) key aspects of study design (e.g., control condition type, population type, and intervention type) moderate training-related changes in dispositional mindfulness scale dimensions, and (c) changes in mindfulness scale dimensions are associated with beneficial changes in mental health outcomes. Scales from widely used dispositional mindfulness measures were combined into 5 categories for analysis: Attention, Description, Nonjudgment, Nonreactivity, and Observation. A total of 88 studies (n = 5,787) were included. Changes in scale dimensions of mindfulness from pre to post mindfulness training produced mean difference effect sizes ranging from small to moderate (g = 0.28-0.49). Consistent with the theorized role of improvements in mindfulness in training outcomes, changes in dispositional mindfulness scale dimensions were moderately correlated with beneficial intervention outcomes (r = .27-0.30), except for the Observation dimension (r = .16). Overall, moderation analyses revealed inconsistent results, and limitations of moderator analyses suggest important directions for future research. We discuss how the findings can inform the next generation of mindfulness assessment. (PsycINFO Database Record