
Objective: This study aimed to develop and validate the International Student Resilience Scale (ISRS) for use with international students in U.S. higher education. Method: International students completed an initial item pool for exploratory factor analysis (EFA) (n = 347) and an independent sample completed the retained items and a resilience measure for confirmatory factor analysis (CFA) (analytic n = 290 from 336 collected responses). Result: EFA supported a three-factor structure with 48 retained items. CFA supported the final model with acceptable fit indices. ISRS scores showed acceptable internal consistency. ISRS scores correlated in expected directions with resilience. Conclusions: ISRS scores may support counselors and international students service professionals in identifying resilience resources and tailoring supports. Future studies should examine score generalization across regions and student subgroups.
Objective This study compares the psychometric properties of scale structures developed by human experts and various generative artificial intelligence (GenAI) models to measure motivation for using GenAI in education. Method Grounded in a theoretical framework, six different item pools were generated by human experts and GenAI prompt bots based on five different language models. Based on the reviews conducted by human experts and GenAI bots, the three forms demonstrating the strongest psychometric performance were selected for the final analysis. Result The confirmatory factor analysis indicated that all three scale structures demonstrated acceptable and good model fit, and high internal consistency and convergent validity. The high correlations between the structures revealed that both human- and AI-generated factors measure similar structures. Conclusions GenAI models can generate psychometrically sound scales; however, human expertise remains essential for ensuring theoretical depth and cultural contextualization. Therefore, hybrid human-AI approaches appear to offer the most robust outcomes.
Objective The purpose of this article is to discuss the effects of missing data (under different conditions) on the accuracy of parameter and standard errors and to help researchers evaluate the appropriate treatment approaches for effectively addressing it in counseling and social science research. This article will review deletion procedures, single imputation procedures, maximum likelihood methods, and multiple imputation methods. Further, this article will provide a justification for maximum likelihood and multiple imputation procedures as superior treatment methods compared to deletion and single imputation approaches. Conclusion The author discusses the practical recommendation of the importance of attending to missing data throughout all aspects of the study, from preplanning to post data collection and results reporting. Further, this article also examines the importance of properly evaluating missing data in a dataset and effectively contextualizing study results given the potential for bias and error related to missing data.
Objective: This study examined the validity of Adjustment Disorder New Module-8 (ADNM-8) among a sample of Filipino college students from Central Visayas, Philippines, in the post-pandemic situation. Methods: The study utilized 5036 college students. Confirmatory factor analysis was employed to identify the best-fitted model. Measurement invariance (i.e. configural, metric, and scalar) was then conducted across year levels. A latent class analysis was performed to determine the number of latent classes. Finally, zero-order correlation was performed to explore adjustment disorder's association with related psychological constructs. Results: The findings revealed that adjustment disorder is best represented by two factors: preoccupation with the stressor and failure to adapt. Moreover, the two-factor model demonstrated measurement equivalence across year levels. LCA revealed a 3-class solution, identifying three distinct symptom classes. Lastly, criterion-related validity was supported by significant positive correlations with psychological distress, depression, and anxiety. Conclusion: This study established the psychometric soundness of ADNM-8 as a measure of post-pandemic adjustment disorder.
Developing self-regulation in early childhood is critical for later development, making the accurate and reliable measurement of this construct essential. This study conducted a reliability generalization (RG) meta-analysis of the Preschool Self-Regulation Assessment (PSRA) following REGEMA guidelines. A total of 35 studies reporting coefficient alpha for the Attention/Impulse Control (AIC) subscale and 12 studies for the Positive Emotion (PE) subscale were included. Bonett's transformation was applied to normalize alpha coefficients for AIC, and random-effects meta-analyses using restricted maximum likelihood estimation were conducted. Results indicated high mean reliability of the scores for AIC (alpha = 0.903, 95% CI [0.885-0.919]) and acceptable reliability of the scores for PE (alpha = 0.779, 95% CI [0.719-0.840]). Moderator analyses showed that larger sample sizes, kindergarten samples, and test administration methods were associated with higher reliability estimates. Overall, findings provide strong evidence for the generalizability of PSRA reliability across diverse samples and contexts.
ObjectiveThe aim of this study was to develop and initially validate scores on Ed Neukrug's Theoretical Orientation Survey (ENTOS). ENTOS is an inventory for identifying the extent to which respondents' worldviews are consistent with the following five major theoretical schools of counseling and psychotherapy: psychodynamic, existential-humanistic, first and second wave cognitive behavioral methods, post-modern orientations, and third wave cognitive behavioral approaches. Method A nationwide non-probability sampling technique was employed (N = 461 counselors and counseling students). Results Rasch analyses supported that ENTOS is comprised of the five major theoretical schools/subscales and person reliability estimates were in the acceptable range for all ENTOS scores. Conclusion The collective results suggest that ENTOS can be used as one way to help counseling students, educators, and practitioners identify their primary theoretical counseling orientation(s). Significance Statement Identifying one's primary theoretical orientation to counseling is a critical task that begins in counselor training programs. The process continues throughout the rest of the counselor's career. This study found that ENTOS was appropriately calibrated for assisting test takers identify which of the five major school(s) of counseling best corresponds with their worldviews. Counseling students, practitioners, and educators can use ENTOS as one way to identify and monitor their alignment with counseling theory.
The convention of identifying findings as statistically significant should no longer be the practice in presenting results of statistical analyses in counseling research and other aligned professions. The representation of statistically significant findings is often attributed to certainty, accuracy, application to a population beyond from which the data were collected, and replication. Four essential problems with the declaration of statistical significance will be highlighted including the (a) misunderstanding of statistical significance, (b) the probability of nonreplicable results, (c) the inefficiency of the declaration of statistical significance, and (d) the influence of power in null hypothesis statistical testing (NHST). Each of these problems may lead to the misunderstanding of research results by consumers of research and the general population from which research is meant to serve. Recommendations include removing the declaration of statistical significance and developing standards related to reporting counseling research.
Study 1 involved 279 university students who completed the EAC-B at baseline and measures of working alliance and psychological distress after the intervention. An exploratory factor analysis (EFA), together with reliability and predictive validity analyses, supported a two-factor structure - Counselor Expertise and Personal Commitment - which explained 40.9% of the variance. Both factors showed excellent internal consistency. Personal commitment significantly predicted all components of the working alliance (goals, tasks, bond) and greater reduction in psychological distress, indicating its central role in the effectiveness of brief counseling. Counselor expertise predicted only the bond dimension of the working alliance and was not related to decreased distress. Study 2, conducted on 315 students, examined the refined version of the 48-item scale. A confirmatory factor analysis (CFA) replicated the two-factor solution with satisfactory fit indices. Overall, the results provide strong evidence for the psychometric validity of the Italian EAC-B and highlight personal commitment as the most robust predictor of alliance quality and counseling outcomes. Assessing expectations at intake may therefore be a valuable component in optimizing engagement and effectiveness in short-term university counseling interventions.
Objective: There is a growing population of individuals with intellectual and developmental disabilities (IDD). Despite growing in prevalence, there remains a significant dearth of evidence-based assessment options to examine a person with IDD's health and wellness. Thus, this study sought to examine the feasibility of an adapted version of the Perceived Wellness Survey (PWS) for individuals with IDD. Method: Adults with IDD (N = 204) took the adapted PWS online. Results: The refined PWS-IDD demonstrated improved model fit over the original version (e.g. CFI increased from .73 to .91), along with acceptable internal consistency reliability (alpha = .61-.82; omega = .69-.84) and evidence of convergent validity of score interpretation with the HRQoL-IDD-16. Invariance testing indicated the revised PWS-IDD performed equivalently across those who did and did not self-report an intellectual disability. Conclusions: Preliminary results support the reliability of PWS-IDD scores and provide evidence for the validity of their interpretation in adults with IDD.
Objective : This study examined the psychometric properties of the Korean version of the Reasons for Living Inventory for Young Adults-II (RFL-YA-II).Method : Exploratory factor analysis was conducted in a sample of Korean young adults (N = 230), followed by Confirmatory factor analysis (CFA) and measurement invariance testing across individuals with recurrent nonsuicidal self-injury (n = 511) and controls (n = 840). Bifactor analyses were performed to further examine the dimensionality. Convergent validity evidence was examined based on correlations with suicide risk.Results : The original four-factor structure was replicated. CFA and bifactor models demonstrated a good fit, and configural, metric, and scalar invariance of RFL-YA-II scores were supported. RFL-YA-II scores showed negative correlations with suicide risks.Conclusions : Findings support the reliability of the scores and the validity of their intended interpretations and uses in research and clinical contexts related to suicide prevention among young adults in Korea.
Fairness is a fundamental aspect of validity and a core ethical principle in counseling, yet remains understudied as a measurement property in counseling research. Differential item functioning (DIF) analysis within the Rasch model offers a robust approach for assessing issues of fairness by evaluating test items for potential bias. This article provides a step-by-step guide for counseling researchers to assess DIF using the Rating Scale Model (RSM) in Winsteps. We outline procedures for evaluating model requirements and fit and identifying DIF using both statistical and graphical approaches. We present an applied example using CES-D data to demonstrate implementation and interpretation.
Objective: In this instructional article, we aim to provide an accessible guide for counseling researchers and practitioners on interpreting item fit statistics within two common polytomous IRT models: the Generalized Partial Credit Model (GPCM) and the Graded Response Model (GRM). Method: Using a simulated dataset (N = 839) from the 20-item Center for Epidemiologic Studies Depression Scale (CES-D), we demonstrate how to interpret item parameters, category response curves, item information, key item fit statistics (e.g. infit, outfit, chi(2), G(2), S-chi(2), RMSEA), and graphical analysis methods. Results: Results demonstrate how different item fit statistics may yield complementary or contrasting information, depending on the model structure and sample characteristics, emphasizing the need for contextual interpretation. Conclusion: Evaluating item fit provides essential validity evidence based on internal structure and offers practical strategies for improving measurement precision and conceptual clarity in counseling research and assessment.
We introduce and demonstrate person fit analysis in the context of counseling research. Person fit is an important but often overlooked component of item response theory (IRT) and Rasch model analyses that has implications for the validity, reliability, and fairness of score interpretations for individual participants. It involves evaluating the degree to which individual person estimates from the model can be appropriately interpreted and used. When person misfit occurs, it may not be appropriate to interpret and use a person's estimate as an indicator of their location on the construct. We use data from the CESD scale to illustrate person fit analysis techniques that combine numeric and graphical evidence. Counseling researchers can incorporate these techniques into their IRT analyses to inform the interpretation and use of scores for individual participants. We emphasize practical interpretations and uses of person fit techniques that can be applied in counseling research contexts.
Objective The purpose of this article is to provide counseling researchers with accessible guidance for interpreting graphical outputs generated in item response theory (IRT) analyses. Although the use of IRT in counseling research has increased substantially, many scholars, students, and practitioners find it challenging to interpret results, particularly the graphical representations that convey important information about item functioning and measurement precision. Method We focus on five common IRT visualizations: item-person maps, item characteristic curves, category response functions, item information curves, and test information functions. Using publicly available data from the 20-item Center for Epidemiological Studies Depression scale (CES-D), analyzed with a graded response model, we provide a step-by-step tutorial in R to demonstrate how to generate, interpret, and apply these graphs in counseling research. Results We highlight how each visualization offers unique insights into item functioning, person fit, and ability distribution, thereby enhancing both counseling research quality and practical application. Conclusion This tutorial demonstrates how IRT graphical outputs can inform item functioning, measurement precision, and scale coverage, and these interpretive strategies offer broad applicability across IRT models for counseling research and practice.
Questionnaires and surveys used to measure psychological constructs in counseling research often rely on rating scales, such as Likert responses. Rasch measurement models are suitable for developing scales to evaluate constructs, and to generate evidence to support valid interpretations and inferences. The philosophical foundations of Rasch measurement theory are straightforward, and its statistical modeling is accessible to researchers and practitioners without extensive quantitative training. This paper provides practical guidelines for constructing sound measures in counseling research and addresses two key questions for rating scale analysis: (a) How should a suitable Rasch model be selected for rating data? and (b) What should be reported from a Rasch-based rating scale analysis? A flowchart is provided to guide model selection, and key indicators for effective reporting are summarized. An analysis of a depression scale is used for illustrating empirical rating scale structures. Implications of using Rasch measurement models in counseling research are discussed.
Objective: The purpose of this study was to develop and validate the LGBTQ+ Microaggressions in Counseling Scale (LMCS), which measures sexual orientation and gender identity-related microaggressions in counseling. Method: Participants were adult U.S. citizens who identified as LGBTQ+ and had received counseling in the last five years. In Study 1 (n = 170), we adapted the Sexual Orientation Microaggressions in Psychotherapy Scale and added new items. We conducted face validity, content validity, and an Exploratory Factor Analysis (EFA). In Study 2 (n = 170), we examined the factor structure through Confirmatory Factor Analysis (CFA). In Study 3 (n = 60), we evaluated the construct validity of the LMCS. Results: We found that a 9-item scale was adequate. The LMCS was significantly associated with working alliance and cultural humility of counselors, supporting the construct validity of the LMCS. Conclusions: The LMCS contributes to improving the counseling environment for LGBTQ+ clients.
Multiple linear regression analysis is one of the most frequently used strategies among counseling researchers and evaluators. While providing a robust strategy to test theories, identify associations between client and intervention characteristics with treatment outcomes, and explore the relationships between program features and intended impacts, the analyses depend on points of scientific rigor and a series of statistical model assumptions. This article reviews two points of scientific rigor and six model assumptions: (a) measurement precision; (b) sufficiency of statistical power; (c) absence of influential cases, outliers, and leverage points; (d) linearity between predictor and criterion variables; (e) normality of residuals; (f) independence of observations; (g) absence of multicollinearity; and (h) homoskedasticity. In each instance, we provide background information and related strategies for assessing and addressing violations. Implications for scientist-practitioners and researchers are discussed.
Objective: We examined the longitudinal psychometric properties of the Counselor Burnout Inventory (CBI) using item response theory with a sample of 590 counselors. Method: We analyzed the response scores of the CBI of three timepoints at six-month intervals in a one-year period. We specifically examined model fit, item-level, and person-level data. Results: We found that counselor burnout was a stable construct across the three timepoints and generally identified continuum of counselor burnout. We also identified item-level issues for specific items at each timepoint, including disordered thresholds. While the CBI was generally well-targeted for participants, we also identified some misfit, suggesting the items may not adequately capture counselor burnout for some participants in this study. Conclusion: We provide additional context for the use and interpretation of test scores associated with the CBI as a measure of counselor burnout, including for whom the scores may or may not be well-targeted.
Objective Effective Single-Session Therapy (SST) hinges on the support provider's specific perspectives that align with SST thinking. This study provides validity evidence for the scores of the Belief and Attitude Toward Therapy Questionnaire (BAT-Q) and develops the Single-Session Therapy Mindset Scale (SSTMS). Method A diverse global sample of 415 practicing and trainee mental health support providers involved with individual psychotherapy provided data online. Results The BAT-Q demonstrated strong psychometric properties in our sample (Cronbach's alpha = 0.833), confirming its continued relevance and reliability. Exploratory Factor Analysis helped with item reduction of the newly developed SSTMS. It demonstrated internal consistency (Cronbach's alpha = 0.826) and significant correlations with the BAT-Q. Receiver operating characteristic curve analysis could identify a cutoff value of 46/60 on SSTMS to identify support providers with the mindset for successful SST practice. Conclusions These scales empower researchers to explore SST implementation, training, and cultural impacts and support providers for self-assessment, ultimately advancing SST.