
The present study investigates the effectiveness of generative large language models (LLMs) as raters across three common rating tasks: (a) social desirability ratings, (b) content validity ratings, and (c) trait importance ratings. Specifically, we examine reliability and validity of LLM-generated ratings across varying occupational contexts, rating methods, LLM families (i.e., GPT-4, GPT-5, and Sonnet 4.5), and inference configurations (i.e., prompt design and temperature settings). Results indicate that LLM ratings exhibit strong reliability and convergent validity in social desirability ratings across occupational contexts, as well as acceptable convergence with human ratings in Likert-type content validity evaluations. In contrast, reliability and convergent validity for trait importance ratings were inconsistent across occupational contexts. Variations in prompt design and temperature settings generally produced small to negligible effects on reliability and validity. Overall, the findings suggest that LLMs can function as effective supplementary raters in test development and validation processes, although greater caution is warranted for certain rating tasks. Practical implications and directions for future research are discussed.
A common belief in the organizational sciences is that estimated interaction effects in ordinary least squares regression cannot be artifacts of common method variance (CMV). This belief rests on the claim that CMV universally attenuates estimated interaction effects. As a result, researchers frequently dismiss CMV concerns when testing moderated relationships. We present an analytic closed-form regression model demonstrating that this universal attenuation claim is false in commonplace scenarios. In particular, when a quadratic term legitimately affects the dependent variable (DV) in the presence of CMV, an estimated interaction effect may be purely artifactual. The common belief holds only in two special cases: (1) when quadratic terms have zero effect on the DV, or (2) when one primary term that enters the interaction is unambiguously unaffected by CMV. We conclude that researchers should apply all standard CMV precautions when testing moderated hypotheses; the fact that one is assessing interactions should not be viewed as a “get out of CMV jail free” card. In addition, we recommend estimating and reporting model specifications both with and without quadratic terms to assess robustness.
We argue that reporting control variable results strengthens the transparency and credibility required for programmatic knowledge building in regression-based empirical research. Control variables are not just technical adjustments; their results provide diagnostic information that can reveal otherwise hidden biases and model misspecification. We present a practical multistep approach that employs control variable coefficients to uncover harmful combinations of multicollinearity, omitted variable biases and correlated measurement error. These harmful combinations, in turn, may generate type 1 errors (false positives) among variables of theoretical interest. We also show how to distinguish true suppressor effects from artifacts of poor model specification. Full reporting supports programmatic research by enabling scholars to compare results across studies, build on prior findings, and refine theory over time. Based on these benefits, we challenge a recent call to omit control variable results from manuscripts. Instead, we recommend that journals and reviewers require their inclusion in all published results tables.
Measurement equivalence/invariance (ME/I) is a prerequisite for cross-group comparisons when using survey data. Although popular structural equation modeling software programs, including Mplus and lavaan, enable tests of ME/I using simple commands, identifying noninvariant items when full ME/I is rejected is more challenging. This paper reviews current procedures for identifying noninvariant items, particularly when there are more than two groups. We recommend systematically rotating the reference items and conducting pairwise comparisons on the factor loadings estimated in the configural invariance model and the intercepts estimated in the metric invariance model. The results are then summarized with the list-and-delete method to identify sets of invariant items and clusters of invariant groups. A custom R package, MEI, is developed to implement our recommended procedures. With simple commands, MEI automatically conducts ME/I tests, identifies noninvariant items, and compares latent means with partial measurement invariance. This allows researchers to interpret cross-group comparison results more precisely. Finally, our procedures for testing ME/I from cross-group comparisons and the MEI package are extended to longitudinal studies with panel data, congruence studies with dyadic data, and multilevel studies with nested data.
Rhythms form an essential part of organizational life, involving embodied patterns of repetition and difference that structure work processes, against the ongoing background of wider organizational and environmental rhythms. Organizational literature increasingly recognizes the importance of rhythms; yet little methodological work exists, either at the level of theorization or practical guidance. The current study draws on Lefebvre's foundational work on rhythmanalysis to elaborate an organizational methodology for studying rhythms. We argue that rhythmanalysis provides a critically oriented approach to understanding social dynamics and advances theorizing about organizational environments by overcoming the dichotomy between entities and processes, stability and change. In this article, we propose methodological guidelines for developing the field of rhythmanalysis in organizational settings by illustrating how it can be conducted through the reanalysis of ethnographic material. We discuss the methodological contributions of rhythmanalysis for a critical exploration of organizational dynamics.
Discrete choice experiments (DCEs) are a promising yet underutilized research method that can provide rigorous empirical evidence about the microfoundations of choice. This method has potential applications in many areas of management research. While DCE methods are similar to conjoint analysis and policy capturing, they are conceptually and methodologically distinct and deliver unique and valuable results that can-in the right contexts-improve external validity. This paper disambiguates DCE from related methods and provides detailed guidance on best practices for conducting DCE research, with emphasis on the elements of experimental design and analysis that are unique to DCEs. Based on a systematic literature review, we identify several emergent and canonical research domains within management where DCE methods could be used to generate novel theoretical and empirical insights. We supplement this review and best-practice guidance with a demonstration experiment and provide all the code and documentation needed for researchers to conduct DCEs with skill and confidence.
Organizational research increasingly uses natural language processing (NLP) to measure textual similarity. Despite common usage, the meaning and consistency of similarity measures (e.g., cosine similarity and Euclidean distance) across common NLP methods (e.g., n-grams and document embeddings) is unclear. This risks misalignment between theoretical constructs and textual measures, undermining the comparability of findings across studies. To address this gap, we review studies using textual similarity in organizational and psychological research, finding a jingle-jangle fallacy: identical labels are used for similarity estimates from different NLP methods, and different labels are used for the same method. Additionally, we examine the consistency of similarity measures across and within NLP methods. Different transformer-based embeddings' similarity results are interchangeable. However, n-grams yield distinct, inconsistent results and are less appropriate for estimating similarity with distance measures. When applied to multi-word inputs, dictionaries and word embeddings return similar results reflecting linguistic style. We provide best practice recommendations and example code for operationalizing textual similarity, including clarifying which NLP methods correspond to content similarity, linguistic style similarity, and semantic similarity at the word, sentence, and document-levels of analysis.
Taxonomies provide a systematic way to organize phenomena and have various practical and theoretical benefits for organizational researchers and practitioners. While taxonomy development and maintenance is often a burdensome process (e.g., time-consuming, costly, and prone to judgmental error), advances in natural language processing (NLP) have the potential to streamline this process. In this study, we employed various evaluation metrics (e.g., cosine similarity) to investigate how machine learning (ML) methods and large language models (LLMs) can automate taxonomy development and maintenance. We examined two embedding models, six clustering algorithms, and three generative LLMs (for creating cluster labels) to construct taxonomies and compared their alignment with four established taxonomies (CABIN, IPIP-NEO-120, ATAF, and O*NET). The confirmatory taxonomic method we examined resulted in effective clustering (i.e., similar text statements were consistently grouped), frequently yielded structures similar to the original taxonomies for ATAF, IPIP-NEO-120, and CABIN (with O*NET being more variable), and resulted in extremely efficient taxonomy title generation. These findings can provide researchers with a foundation for how to approach NLP-based taxonomy development and maintenance activities for their own contexts.
Qualitative researchers face an enduring question: How many interviews do I need? While a variety of guidelines exist, there is limited consensus over which specific factors should determine the number of interviews required. We examined the determination of interview sample sizes in 562 qualitative studies across six high-impact management and organizational journals over a decade. Our findings reveal considerable variance in interview numbers, yet limited information is often provided on the criteria used to determine them. To promote clearer alignment between sample sizes and methodology, we examined studies with detailed descriptions of their interview sampling. We identified specific "sampling moves" used to determine the number of interviews, categorized into three types-opening, focusing, and closing sampling moves-that researchers use to establish confidence in the sample and support theoretical insights. By implication, our study refutes the notion of a "magic" interview number. Instead, sampling moves are heuristic tools that qualitative researchers can thoughtfully adapt to their analytical aims when determining appropriate sample sizes.
Dominance and unfolding response processes describe two ways in which individuals may respond to rating scale items. The dominance process assumes a monotonic relationship between a latent trait and the probability of endorsement and is typically modeled using a linear factor model within structural equation modeling (SEM). In contrast, the unfolding process assumes single-peaked response functions, with endorsement most likely when item and person locations are close on the latent continuum. Fitting unfolding models usually requires specialized software, which limits their integration with SEM. In this article, we proposed the ordered categorical response unfolding model (OCRUM), which can be estimated in Mplus. We illustrated its use with two empirical datasets and found that item and person locations were comparable to those obtained from the generalized graded unfolding model (GGUM). We also conducted Monte Carlo simulations to examine parameter recovery under varying sample sizes, test lengths, and response formats. Finally, we demonstrated that OCRUM can serve as the measurement component of a general structural equation model, enabling dominance and unfolding response processes to be represented within a single SEM framework.
Alternative approaches to personality measurement, such as open-ended narrative-based assessments, have potential advantages for organizational research and practice. In this research, we investigate factors that affect valid application of natural language processing (NLP) for scoring open-ended personality assessments and when, how, and why such assessments capture personality-related variance. Using a large sample of responses to open-ended assessments, convergence between NLP scores and self-report target scores increased as the degree of customization and the sophistication of the underlying model increased, with the worst psychometric performance occurring for zero-shot large language model (LLM) scores and the best for fine-tuned LLM scores. However, all scoring methods exhibited evidence of validity. Additionally, when trained to predict direct evaluations of the narrative responses, correlations with target scores were large (M = .83). NLP scores also exhibited discriminant and criterion-related validity evidence. However, validity was contingent upon the methodological rigor employed in developing writing prompts. Prompts designed to elicit trait-relevant information outperformed generic prompts, and this occurred because trait-specific prompts increased the amount of trait-relevant information (i.e., narrative units), which was associated with enhanced convergence with target scores.
Sentiment analysis (SA) has grown considerably in organizational science research over the past two decades, particularly in the last few years. While enthusiasm for integrating advanced natural language processing algorithms is encouraging, authors are not reaping the benefits of such tools fully. Our systematic review of SA application in the organizational sciences suggests that authors struggle to appreciate all of the decisions that are inherent to SA, the choices that are available at each decision point, and the consequences of each choice. To address this gap, we use a working example to illustrate four critical decision points authors confront when conducting SA, and the subsequent impact different choices can have on one's conclusion. Decision points include selecting the SA method, computing a sentiment score, preprocessing the data, and using an appropriate level of analysis. We conclude with a framework outlining five dimensions (e.g., accuracy, interpretability, computational cost) to guide the selection of an SA approach based on study goals and needs, along with seven recommendations to authors wishing to apply SA.
Researchers, engineers, and entrepreneurs are enthusiastically exploring and promoting ways to apply generative artificial intelligence (GenAI) tools to qualitative data analysis. From promises of automated coding and thematic analysis to functioning as a virtual research assistant that supports researchers in diverse interpretive and analytical tasks, the potential applications of GenAI in qualitative research appear vast. In this paper, we take a step back and ask what sort of technological artifact is GenAI and evaluate whether it is appropriate for qualitative data analysis. We provide an accessible, technologically informed analysis of GenAI, specifically large language models (LLMs), and put to the test the claimed transformative potential of using GenAI in qualitative data analysis. Our evaluation illustrates significant shortcomings that, if the technology is adopted uncritically by management researchers, will introduce unacceptable epistemic risks. We explore these epistemic risks and emphasize that the essence of qualitative data analysis lies in the interpretation of meaning, an inherently human capability.
Machine learning and artificial intelligence (AI) are increasingly used within organizational research and practice to generate scores representing constructs (e.g., social effectiveness) or behaviors/events (e.g., turnover probability). Ensuring the reliability of AI scores is critical in these contexts, and yet reliability estimates are reported in inconsistent ways, if at all. The current article critically examines reliability estimation for AI scores. We describe different uses of AI scores and how this informs the data and model needed for estimating reliability. Additionally, we distinguish between reliability and validity evidence within this context. We also highlight how the parallel test assumption is required when relying on correlations between AI scores and established measures as an index of reliability, and yet this assumption is frequently violated. We then provide methods that are appropriate for reliability estimation for AI scores that are sensitive to the generalizations one aims to make. In conclusion, we assert that AI reliability estimation is a challenging task that requires a thorough understanding of the issues presented, but a task that is essential to responsible AI work in organizational contexts.
Online data are constantly growing, providing a wide range of opportunities to explore social phenomena. Large Language Models (LLMs) capture the inherent structure, contextual meaning, and nuance of human language and are the base for state-of-the-art Natural Language Processing (NLP) algorithms. In this article, we describe a method to assist qualitative researchers in the theorization process by efficiently exploring and selecting the most relevant information from a large online dataset. Using LLM-based NLP algorithms, qualitative researchers can efficiently analyze large amounts of online data while still maintaining deep contact with the data and preserving the richness of qualitative analysis. We illustrate the usefulness of our method by examining 5,516 social media posts from 18 entrepreneurs pursuing an environmental mission (ecopreneurs) to analyze their impression management tactics. By helping researchers to explore and select online data efficiently, our method enhances their analytical capabilities, leads to new insights, and ensures precision in counting and classification, thus strengthening the theorization process. We argue that LLMs push researchers to rethink research methods as the distinction between qualitative and quantitative approaches becomes blurred.
Websites represent a crucial avenue for organizations to reach customers, attract talent, and disseminate information to stakeholders. Despite their importance, strikingly little work in the domain of organization and management research has tapped into this source of longitudinal big data. In this paper, we highlight the unique nature and profound potential of longitudinal website data and present novel open-source code- and databases that make these data accessible. Specifically, our codebase offers a general-purpose setup, building on four central steps to scrape historical websites using the Wayback Machine. Our open-access CompuCrawl database was built using this four-step approach. It contains websites of North American firms in the Compustat database between 1996 and 2020—covering 11,277 firms with 86,303 firm/year observations and 1,617,675 webpages. We describe the coverage of our database and illustrate its use by applying word-embedding models to reveal the evolving meaning of the concept of “sustainability” over time. Finally, we outline several avenues for future research enabled by our step-by-step longitudinal web scraping approach and our CompuCrawl database.
While machine learning (ML) can validly score psychological constructs from behavior, several conditions often change across studies, making it difficult to understand why the psychometric properties of ML models differ across studies. We address this gap in the context of automatically scored interviews. Across multiple datasets, for interview- or question-level scoring of self-reported, tested, and interviewer-rated constructs, we manipulate the training sample size and natural language processing (NLP) method while observing differences in ground truth reliability. We examine how these factors influence the ML model scores’ test–retest reliability and convergence, and we develop multilevel models for estimating the convergent-related validity of ML model scores in similar interviews. When the ground truth is interviewer ratings, hundreds of observations are adequate for research purposes, while larger samples are recommended for practitioners to support generalizability across populations and time. However, self-reports and tested constructs require larger training samples. Particularly when the ground truth is interviewer ratings, NLP embedding methods improve upon count-based methods. Given mixed findings regarding ground truth reliability, we discuss future research possibilities on factors that affect supervised ML models’ psychometric properties.
When assessing text, supervised natural language processing (NLP) models have traditionally been used to measure targeted constructs in the organizational sciences. However, these models require significant resources to develop. Emerging “off-the-shelf” large language models (LLM) offer a way to evaluate organizational constructs without building customized models. However, it is unclear whether off-the-shelf LLMs accurately score organizational constructs and what evidence is necessary to infer validity. In this study, we compared the validity of supervised NLP models to off-the-shelf LLM models (ChatGPT-3.5 and ChatGPT-4). Across six organizational datasets and thousands of comments, we found that supervised NLP produced scores were more reliable than human coders. However, and even though not specifically developed for this purpose, we found that off-the-shelf LLMs produce similar psychometric properties as supervised models, though with slightly less favorable psychometric properties. We connect these findings to broader validation considerations and present a decision chart to guide researchers and practitioners on how they can use off-the-shelf LLM models to score targeted constructs, including guidance on how psychometric evidence can be “transported” to new contexts.
Our paper provides a conceptualization of magnitude-based hypotheses (MBHs). We define an MBH as a specific type of hypothesis that tests for relative differences in the independent impact (i.e., effect size difference) of at least two explanatory variables on a given outcome. We reviewed 1,715 articles across eight leading management journals and found that nearly 10% (165) of articles feature an MBH, employing 41 distinct methodological approaches to test them. However, approximately 40% of these papers show missteps in the post-estimation process required to evaluate MBHs. To address this issue, we offer a conceptual framework, an empirical illustration using Bayesian analysis and frequentist statistics, and a decision-tree guideline that outlines key steps for evaluating MBHs. Overall, we contribute a framework for applying MBHs, demonstrating how they can shift theoretical inquiry from binary questions of whether an effect exists, to more comparative questions about how much a construct matters,compared to what, and under which conditions.
Over the past two decades, forced-choice (FC) measures have received considerable attention from researchers and practitioners in industrial and organizational psychology. Despite the growing body of research on FC measures, there has not yet been a comprehensive review synthesizing the diverse lines of research. This article bridges this gap by presenting a systematic review of post-2000 literature on FC measures, addressing ten critical questions, including: 1) validity evidence, 2) faking resistance, 3) FC IRT models, 4) FC test design, 5) FC measure development, 6) test-taker reactions and response processes, 7) measurement and predictive bias, 8) reliability, 9) computerized adaptive testing, and 10) random responding. The review adopts a historical perspective, tracing the development of FC measures and highlighting key empirical findings, methodological advances, current trends, and future directions. By synthesizing a substantial body of evidence across multiple research streams, this article serves as a valuable resource, providing insights into the psychometric properties, theoretical underpinnings, and practical applications of FC measures in organizational contexts such as personnel selection, development, and assessment.