
Abstract: In this study, we evaluate four German-language versions of the PSC (parent long: 35 items, short: 17 items, children and adolescents long: 35 items, short: 17 items). In two online surveys (COMPARE, ASK), we examined the psychometric quality criteria. The COMPARE data set included n = 388 parent–child dyads (age: parents M (SD) = 42.20 (7.37); age: children M (SD) = 10.90 (2.91)), and the ASK data set included n = 278 dyads (parents M (SD) = 43.00 (6.14); children M (SD) = 12.50 (2.11)). Data analysis revealed that all 35 items in the two versions fulfilled the quality criteria of classical test theory. The internal consistency for all four versions was acceptable to good. The three-factor structure was replicated and the discriminant and convergent validity compared to the CBCL and KIDSCREEN were satisfactory. Overall, the PSC is an economical, valid, and comprehensive screening instrument for mental health problems in children and adolescents.
Abstract: Objective: This study examined the preliminary psychometric quality and construct validity of the meinBeruf.ch assessment system and evaluated gender-related fairness at the level of input variables and algorithmic recommendations. Methods: The sample comprised 914 users of the meinBeruf.ch platform. Participants completed standardized personality (BFI-2), interest (RIASEC), cognitive (HMT), and domain-specific skill measures, followed by algorithmic career recommendations. Reliability and construct validity were assessed using internal consistency coefficients and correlations between psychological scales and industry-fit scores. Gender fairness was examined by testing whether self-reported gender added incremental predictive value beyond psychological profiles, and by reporting supplementary input-variable analyses of gender differences, DIF, measurement invariance, and a technical counterfactual audit of the ranking layer. Results: The psychological scales demonstrated adequate to high internal consistencies, and correlations with algorithmic recommendations showed theoretically consistent patterns. Adding gender to random-forest prediction models changed predictive accuracy only minimally (M |ΔR2| = .011), indicating that gender provided little systematic predictive value beyond the psychological profile. Supplementary analyses documented expected gender differences in some input variables and no counterfactual rank changes when only the gender label was varied. Limitations: The study used a self-selected online sample, relied primarily on cross-sectional platform data, and lacked long-term criterion outcomes such as later career choice, satisfaction, or persistence. Conclusion: The meinBeruf.ch assessment system provides preliminary evidence for psychometric quality and construct validity. The algorithmic ranking layer behaved technically gender-fair, while gender-related differences in personality, interests, and self-estimated skills remain important for interpretation and future research, especially in STEM-related vocational guidance.
Abstract: Large language models (LLMs) have gained increased attention as tools for automatic item and scale generation in psychological assessment. Yet, it is not known how prompt instructions can affect the psychometric quality of LLM-generated items and scales. To test whether different prompt instructions influence the quality of generated items and scales, a general-purpose LLM (GPT4o) was prompted with seven different prompt instructions (one adhering to prompt engineering principles solely and six combining these principles with five different psychometric characteristics). Items were generated for two constructs, one established (emotional stability) and one rather novel (climate change anxiety). The LLM-generated items and scales were evaluated in terms of content validity, item difficulty, model fit, factor structure, reliability, and validity. Overall, there were no systematic differences regarding content validity, item difficulty, model fit, reliability and convergent validity, but unsystematic differences in terms of validity.
Abstract: Psychological tests are commonly evaluated using internal consistency, indexed by Cronbach's α. This emphasis fosters the assumption that maximizing α is sufficient for building useful scales – a position we call the reliability-first myth. The present study illustrates a known but practically important implication of classical test theory: reliability does not uniquely determine how well observed scores support individual-level interpretation. Using a Monte Carlo simulation within a tau-congeneric framework, we varied item-difficulty composition, test length, number of response categories, and mean factor loading. Differentiation was operationalized using the Critical Difference (CD). Response-category count, item-difficulty composition, and test length explained substantial variation in CD, whereas achieved α contributed little unique value. Extreme-difficulty sets yielded the smallest CDs but only by compressing observed-score variance. Mixed-difficulty sets cover the trait more broadly without losing precision. Reliability is therefore necessary but not sufficient for individual-level interpretation.
Abstract: Over decades, the literature on educational and psychological assessment of latent attributes has proposed a number of approaches for obtaining and interpreting scores from tests and questionnaires, with measurement and prediction standing out. These approaches have traditionally been portrayed as fundamentally different, and it has been claimed that they would pursue different goals and have distinct interpretations. However, such views might reflect a myth – one perpetuated through repetition rather than theoretical necessity. With this article, we aim to challenge this long-standing belief by arguing that the distinction between measurement and prediction is not necessary. Specifically, we propose that both the measured score and the predicted score can be viewed as estimators of the same underlying quantity, the true score, thereby unifying them under the broader framework of estimation theory. This unification helps clarify conceptual relationships between assessment approaches and promotes greater coherence in how scores are interpreted across research and practice.
Abstract: Vocational interests represent relatively stable preferences for specific activities and contexts and constitute key predictors of educational choices and career development. The Basic Interest Markers (BIM) is a public-domain inventory designed to assess 31 theoretically derived basic-interest dimensions. A total of 1,977 participants completed a 337-item dichotomous BIM adaptation. Exploratory factor analyses were conducted to examine both the item-level structure and the higher-order organization of the 31 interest dimensions. Most BIM dimensions replicated the proposed structure, with majority of items loading on their hypothesized factor and showing moderate-to-high loadings. Mathematics, Law, Teaching, Sales, and Human Relations Management showed the strongest internal coherence. Two pairs of dimensions – Finance and Business, and Medical Services and Life Science – merged into unified factors, whereas Family Activity split into two components. The Management factor showed weak cohesion. At a higher-order level, dimensions were organized into eight broader domains, broadly consistent with previous clustering structures. Findings provide strong evidence for the replicability of the BIM’s 31-dimension structure in Spanish, supporting its use in vocational research. Results also support a hierarchical organization of vocational interests and point areas for possible refinement.
Abstract: Full Achievement Emotions Questionnaire (AEQ) forms are lengthy, and no validated brief version exists for Rioplatense adolescents or young adults. We therefore evaluated a shortened AEQ class-domain form in a secondary-school student sample (n = 374). We contrasted two CFA specifications: an independent cluster model (ICM) and a third-order hierarchical model in which eight emotions loaded on two valence dimensions that in turn loaded on a general emotional engagement factor (GEE), reflecting overall affective involvement in class. The hierarchical (CFI = .952, RMSEA = .062) showed slightly better fit than the ICM (CFI = .949, RMSEA = .066) and was retained as the primary measurement solution. Schmid–Leiman results supported a strong GEE factor (ω = .98, ωh = .76), and this factor predicted procrastination as expected (β = −.71, p < .001). Results support two alternative scoring approaches – eight subscales or a model-estimated GEE – to index either discrete emotions or a broad engagement continuum within a measurement/prediction framework.
Abstract: The Multidimensional Emotional Disorder Inventory (MEDI) is a brief, transdiagnostic measure assessing nine dimensions of emotional disorders. This study evaluated the psychometric and clinical properties of the German version of the MEDI in a large sample (N = 1,129) including healthy individuals and patients from two outpatient psychotherapy clinics. The results showed high internal consistency (Cronbach's α = .73–.92) and acceptable test–retest reliability (rtt = .58–.78) over a 7-month interval. Exploratory structural equation modeling confirmed the original nine-factor model, although with reduced consistency for the Avoidance scale. Correlations with established symptom and personality measures, as well as clinical diagnoses, indicated good convergent and discriminant validity. Overall, the MEDI – German version demonstrated good psychometric properties, making it suitable for evaluating therapeutic interventions for emotional disorders in clinical practice and research.
Abstract: While artificial intelligence (AI) gained attention for eliciting diagnostic evidence from text answers using NLP or for generating visual stimuli, few studies investigate its use for analyzing visual data such as free-hand sketches from graphical response formats. The present case study is based on a formative assessment including instructional considerations and illustrates the application of three AI approaches to graphical responses from 96 students. Students answered two tasks assessing the conceptual understanding of fractions. Comparisons of AI approaches to expert ratings reveal promising results of two approaches (rule-based approach and ResNet). The third approach using a pretrained clip model showed lower performance, especially in tasks requiring counting. Additional comparisons to diagnostic evidence from other items highlight the relevance of graphical response items as a distinct item format. We discuss strengths and weaknesses of the approaches, as well as the case study, and hint to topics for further research.
Abstract: The impostor phenomenon (IP) refers to individual differences in difficulties in internalizing positive feedback and success, and fear of being exposed as an intellectual fraud. The 2015 wave of the SOEP-IS study included five of the 20 items of the German-language Clance Impostor Phenomenon Scale. This study analyzed the psychometric properties and validity of the IP measure used in the SOEP-IS data (N = 2,643). Confirmatory factor analysis supported a unidimensional model and invariance across sex. The internal consistency was good for a very brief measure (α and ω = .78). Correlations with self-esteem, the Big Five traits, sadness and worrying met expectations and were stable across a 2-year interval (rchange ≤ .05). The findings support the reliability and validity of the abbreviated IP measure. Its use for analyzing panel data such as the SOEP-IS is recommended considering its limitations (e.g., limited reliability and coverage of the IP).
Abstract: Despite the profound impact of artificial intelligence (AI) in diverse contexts, large-scale socio-economic panel studies have rarely addressed the use and evaluation of AI for individual respondents. Therefore, the Artificial Intelligence Experience and Attitude Survey (AIEAS) is introduced to measure awareness, experience, attitude valence, and usage intention regarding AI in the work, healthcare, and education domains. The vignette-based items describe different AI applications that are administered in a planned missingness design to meet the brevity requirements of large-scale studies. A qualitative interview study (N = 30) confirmed the comprehensibility of the items, while a quantitative web-based study (N = 1,084) demonstrated satisfactory measurement structure and precision. These findings attest to good psychometric properties of the AIEAS that allow for a multidimensional measurement of AI experiences and evaluations. Integrating the instrument into panel studies will support examining societal trends over time to inform the scientific community and policies in different domains.
Abstract: Background: The Motivations to Eat Meat Inventory (MEMI) assesses four motives for eating meat: Natural, Necessary, Normal, and Nice. This study aims to psychometrically evaluate the German MEMI. Methods: We reanalyzed data from two German-speaking samples (N = 389; N = 1,498) who completed the MEMI online, one with an importance-based (Sample 1), one with an agreement-based (Sample 2) response format. We ran confirmatory factor analyses, tested measurement invariance, and examined validity. Results: Across both samples, bifactor models showed the best fit, with slightly better fit in Sample 1 but acceptable fit in both formats. Measurement invariance across gender, age, and education largely supported factorial validity. Results suggested acceptable convergent validity. Regarding discriminant validity, a strong general factor underpinned responses, with Normal and Nice explaining meaningful additional variance. Nice meaningfully predicted meat consumption. Limitations: Results are limited by sample differences. Conclusions: The German MEMI is a valid tool to assess motivations to eat meat and best modeled with a bifactor structure.
Abstract: Adults’ attitudes toward governmental public health measures are critical to assess because they relate to people’s compliance with public health measures and policies. A psychometrically sound measure for assessing trust in governmental public health measures and satisfaction with information politics in German is still missing. Based on theoretical approaches, we developed a measure to assess this trust (four items) and satisfaction (four items). We tested it longitudinally (Sample 1, n = 1,038 adults at T1 and T2, MAge = 43.56, representative for the German population) and cross-sectionally (Sample 2, n = 1,346 adults, MAge = 23.03). Results from both samples suggest initial evidence for high degrees of reliability and validity in Germany. Longitudinal results suggested initial evidence for high degrees of retest reliability, structural validity, and strict measurement invariance over time. Thus, this brief measure can be utilized for future opinion polls, political surveys, or public health-related policy decisions.
Abstract: Lexical studies seek universal personality dimensions by analyzing trait-descriptive words. This study examined whether large language model (LLM) agents, endowed with Big Five profiles, could reproduce human lexical structures. GPT-4o agents rated Japanese trait words, and principal component analyses with Bass-Ackward comparisons were conducted. The broadest dimensions, corresponding to the Big Two, were robustly recovered, supporting their cross-linguistic universality. Four stable components also emerged, including prosociality and conscientiousness, whereas activity/introversion and Neuroticism appeared less consistently. Compared with human data, LLM responses showed higher interitem correlations, greater internal consistency, and an exceptionally large first component, suggesting blurred distinctions between traits. Congruence with human data was lower than human–human benchmarks, reflecting multilingual training influences and biases. These findings indicate that LLMs can recover broad dimensions but impose structural compression. Future research should test prompt variations, alternative models, and baseline conditions to clarify sources of LLM-generated structures.
Abstract: Workplace compassion has been recognized as a key resource for employee well-being and organizational functioning, yet valid instruments for its assessment remain scarce in Portuguese contexts. This study aimed to validate the Workplace Compassion Scale (WCS) in a sample of 455 higher education professionals (67.7% female; M = 44.23 years; SD = 9.44). Confirmatory factor analysis supported a second-order hierarchical model with four interrelated dimensions (Noticing, Empathizing, Sensemaking, and Acting), reflecting a sequential process of compassion that involves recognizing suffering, emotional connection with it, making sense of its causes, and taking action to alleviate it, showing good model fit. Internal consistency was excellent for the total scale (α = .91) and acceptable to good for subscales. Test–retest analyses confirmed temporal stability for the total score and three of the subscales. Evidence for nomological validity emerged from positive associations with self-compassion. Nomological validity was partially supported, with trivial associations for climate soothing-safeness/drive, a near-zero association for the WCS total scores with emotional exhaustion, and an inverse association between Acting and emotional exhaustion. At the same time, small positive correlations with work–family conflict, burnout total, and climate threat indicate nomological complexity that warrants further investigation. Unexpected positive associations with work–family conflict suggest that compassion may also involve empathic strain or role spillovers. Overall, the Portuguese WCS demonstrates strong psychometric properties and offers a theoretically grounded, practical tool for investigating how compassion operates in organizational settings. Its application can inform interventions aimed at fostering well-being, reducing strain, and cultivating compassionate work environments. Future research should further examine these associations to clarify whether they reflect contextual pressures, empathic strain, or cultural specificities in the Portuguese higher education sector.
Abstract: Introduction: A validated German version of the Everyday Discrimination Scale (EDS) to assess the frequency of perceived ethnic discrimination is lacking. Despite most validation studies being conducted in the United States, suggesting measurement invariance across ethnic groups, systematic cross-cultural validation is essential. Consequently, we validated a German version of the EDS across three studies in Austria. Results: Factor analyses in a diverse student sample (N = 1,067; Study 1) and a sample of immigrants from Türkiye (N = 583; Study 2) revealed a robust single-factor structure and high internal consistency. Study 1 further supported criterion and convergent validity. Multigroup CFA in Study 2 indicated partial scalar invariance across gender. Cognitive interviews (N = 20 immigrant men from Türkiye; Study 3) indicated high item comprehensibility but variable representativeness. Discussion: Findings support the psychometric adequacy of the German version of the EDS, despite its limited sensitivity to subtle discrimination.
Abstract: The BFI-2 (Soto & John, 2017) has been adapted to various languages and cultures, but not yet to a Swedish context. Its predictive power has been investigated mainly on domain level, but to a lesser extent on facet level. In three large samples, N = 824, N = 1,104, and N = 823, using an online questionnaire, the intended five-factor structure for the domains Extraversion, Agreeableness, Conscientiousness, Negative Emotionality, and Open-Mindedness was reproduced in all three samples, both based on items and the facet scores. Reliability was good for the domains, more diverse for the facets, but still acceptable in most cases. Invariance analyses revealed high correspondence with the American original. The predictive validity was evaluated for individual (psychological well-being), interindividual (attachment styles), and intergroup (social dominance orientation and ambivalent sexism) variables. Regression analyses revealed significant and theoretically meaningful patterns, both for the domains and their facets.
Abstract: Developing valid measurement models for latent variables, such as personality traits, is essential for accurate psychological assessment. A critical aspect of this process is evaluating the fit of psychometric models. However, commonly used model fit indices are often affected by nuisance parameters – such as sample and model size – making the use of conventional cutoff values problematic, as these thresholds are typically based on narrow simulation scenarios. Recently, a machine learning-based approach to model fit evaluation has been introduced by Partsch and Goretzko (2025), offering a more flexible and data-informed alternative. This approach considers not only various indicators of model (mis)fit but also multiple characteristics of the data and model. In this paper, we discuss how nuisance parameters can distort model fit evaluation and present the core principles of this new evaluation strategy. We further demonstrate how interpretable machine learning can reveal the decision-making process of a pretrained predictive model in identifying model misspecification.
Abstract: This study validates the French version of the Digital Jealousy Scale (DJS), a self-report instrument assessing social media-induced jealousy (SoMJ) in romantic contexts. A sample of 551 French-speaking participants completed the DJS along with measures of jealousy, attachment, envy, relationship satisfaction, and self-esteem. Confirmatory factor analyses supported a unidimensional structure with good fit indices and high internal consistency (α = .86). Measurement invariance was established across gender, relationship status, and divorce status, although results for engaged participants should be interpreted cautiously due to limited power. DJS scores correlated positively with multidimensional jealousy, attachment anxiety, and malicious envy, and negatively with self-esteem, relationship satisfaction, and age. These findings support the scale’s validity and reliability within attachment and social comparison frameworks. The DJS thus provides a brief, robust tool for assessing SoMJ in French-speaking populations. Future work should examine its predictive value and cross-cultural applicability.
Abstract: Despite the critical role of formative assessment (FA) in teaching and learning, there are few instruments that enable its evaluation from students’ perspective in mathematics education. We aimed to develop and validate a scale to collect secondary education students' perspective of the FA practices developed by their teachers in mathematics class. The instrument was designed starting from a rational criterion, followed by content and face validity evidence analysis. Then, the internal structure was studied through an exploratory factor analysis (N = 534), and a second-order solution was reached that preserved 24 items distributed in four first-order dimensions (collaborative learning, formative feedback, responsive teaching, and students’ reflection) and a general factor. This structure was verified by confirmatory factor analysis (N = 1,943). The final version of the instrument was examined regarding its factorial invariance and temporal stability. The scale promotes teachers’ and students’ involvement in collaborative dynamics and evaluation of mathematics learning environments.