
Abstract: The social climate of prisons is increasingly recognized as a key factor in rehabilitation, institutional safety, and staff well-being. The Essen Climate Evaluation Schema (EssenCES) is the most extensively studied tool in this context, assessing safety and support. This study aimed to confirm its psychometric properties and examine its comparability and fairness between inmates and prison officers. Using pooled data from five German and Swiss research projects ( n = 1,512 inmates, n = 1,807 prison officers), previously observed group-specific differences were replicated: inmates reported higher perceptions of safety, while officers rated therapeutic hold higher. Confirmatory factor analyses supported the proposed three-dimensional oblique structure for both groups (Inmates: CFI = .946; RMSEA = .054; Prison officers: CFI = .960; RMSEA = .052), and subscale reliability ranged from acceptable to good (ω = .76–.85). Partial measurement invariance indicated that the EssenCES is largely comparable across groups. However, differential item functioning analyses showed that fairness and difficulty varied for most items. Despite these minor group differences, the EssenCES remains a suitable tool for assessing the prison social climate from different perspectives.
Abstract: Research on socially/ethically aversive personality traits has increasingly focused on their common core, the dark factor of personality (D). Based on large samples (total N > 167,000), we herein evaluate the French-language version (in comparison to the English original) of the 70-item D measure (D70), including psychometric properties, measurement invariance, sources of noninvariance, and criterion-related validity. Results confirm highly satisfactory psychometric properties and validity. Configural and metric invariance, as well as partial scalar and strict invariance are demonstrated across languages. The predominant sources for noninvariance were translation and comprehension bias. All findings also held for the shorter D35 and D16 measures, which are subsets of the D70. In sum, the D70 is a reliable measure of aversive personality in French, nomologically consistent with previously validated language versions suitable for cross-cultural research.
Abstract: The 9-item Shared Decision Making Questionnaire (SDM-Q-9) is widely used to measure shared decision-making (SDM) across clinical settings, but no study so far has synthesized its psychometric properties. In the present study, we systematically reviewed and synthesized psychometric evidence on the SDM-Q-9. The MEDLINE, Web of Science, PsycInfo, and CINAHL databases were searched for original studies in English or German, providing any information on psychometric properties of the SDM-Q-9 considering all possible interpretations and constructs measured by the instrument. Details on study design, sample characteristics, and information on psychometric evidence were extracted. We included 101 studies, with 49 articles contributing evidence on reliability, 91 on validity (87 on interrelations with other variables and 68 on the internal structure), and three on fairness. We found a high amount of psychometric evidence on the validity and reliability of the SDM-Q-9 as a measure of SDM perceived by the patient. Evidence is lacking for other interpretations and fairness. Notable is the weak association of the SDM-Q-9 with physician-rated SDM. In summary, substantial evidence supports reliability and validity of the SDM-Q-9 for assessing the subjectively experienced level of SDM by the patient, but further research is needed on alternative interpretations and on fairness of the measure.
Abstract: The level of alexithymia affects a person’s mental and physical health. Nevertheless, this trait is still measured using unreliable tools. The Perth Alexithymia Questionnaire (PAQ) is a promising measuring tool, but no study has examined its psychometric characteristics using Item Response Theory (IRT). Thus, we aimed to explore the psychometric features of the PAQ using IRT and Classical Test Theory. A sample of Czech adults ( n = 848, age: M = 34.95, SD = 11.90, females: 81.13%) participated in an online survey. We measured alexithymia, empathy, sensory processing sensitivity (SPS), neuroticism, anxiety, and depression. Confirmatory factor analysis provided evidence of the good fit of a five-factor solution: Reliability of the PAQ was good (omega = 0.70–0.96). Measurement invariance testing revealed that the PAQ measures alexithymia invariantly between people who are single and in partnership. Partial invariance was found in sex. The PAQ items had high discrimination, and their measurement precision was highest in individuals with above-average alexithymia. Higher alexithymia was present in males. Finally, alexithymia was positively associated with depression, anxiety, and neuroticism, and with SPS after neuroticism was taken into account. In conclusion, the PAQ represents a reliable and valid instrument for assessing alexithymia.
Abstract: Evidence on the structure of the Depression Anxiety Stress Scales for Youth (DASS-Y) is both limited and equivocal, hindering the use and interpretation of the scale scores. The present study aimed to re-examine the factor structure of the DASS-Y using exploratory structural equation modeling (ESEM) and bifactor-ESEM frameworks and to investigate whether the structure of DASS-Y responses varies by country and educational stage. Convergent validity evidence was also evaluated through associations with externalizing problems, bullying victimization, and subjective well-being scores. We recruited three samples: Serbian early adolescents ( n = 412; age range = 11–14 years), Australian early adolescents ( n = 616; age range = 11–14 years), and Serbian middle-to-late adolescents ( n = 595; age range = 15–19 years). Both the ESEM model, allowing cross-loadings, and the bifactor-ESEM model, incorporating cross-loadings and a general distress factor, yielded meaningful solutions for the DASS-Y responses. Evidence supported the measurement invariance of both models across countries and educational stages. We conclude that more complex solutions beyond the standard three-factor model should be considered when testing the structure of DASS-Y responses.
Research on socially/ethically aversive personality traits has increasingly focused on their common core, the dark factor of personality (D). Based on large samples (total N > 167,000), we herein evaluate the French-language version (in comparison to the English original) of the 70-item D measure (D70), including psychometric properties, measurement invariance, sources of noninvariance, and criterion-related validity. Results confirm highly satisfactory psychometric properties and validity. Configural and metric invariance, as well as partial scalar and strict invariance are demonstrated across languages. The predominant sources for noninvariance were translation and comprehension bias. All findings also held for the shorter D35 and D16 measures, which are subsets of the D70. In sum, the D70 is a reliable measure of aversive personality in French, nomologically consistent with previously validated language versions suitable for cross-cultural research.
: This study aims to develop and validate a Chinese version of the Short-Form Employability Five-Factor instrument. In Study 1 (N = 351), the original scale was translated into Chinese, and confirmatory factor analysis supported the correlational five-factor structure. In Study 2, confirmatory factor analysis was conducted to cross-validate the factor structure in an independent Chinese sample (N = 345). Moreover, multigroup confirmatory factor analysis indicated configural and metric invariance across Chinese and UK samples (N = 281). In addition, the Chinese scale showed significant correlations with affective commitment, performance, and subjective career success, supporting its criterion-related validity. Finally, the Chinese employability scale demonstrated satisfactory reliability, convergent validity, and discriminant validity. Overall, the findings supported the Chinese version of the competence-based employability scale as a reliable and valid measure.
: We examined whether self-report and behavioral measures of impulsivity fail to converge because they assess impulsivity at the trait and state levels, respectively. To this end, we hypothesized that the stable (trait) level of performance across administrations of behavioral impulsivity measures would be predictive of self-report impulsivity scores, such that the correlation between self-report and behavioral measures would increase with the number of assessments across which behavioral impulsivity measures are aggregated. A sample of 383 participants (Mage = 30.44, SDage = 4.33, 69% female) completed self-report and behavioral impulsivity measures three times over a week. We found that for two of the three behavioral measures, the association between self-report and behavioral measures increased when both assessed impulsivity at the trait level. To our knowledge, the current study is the first to test trait-level versus state-level assessment as an explanation for the divergence between self-report and behavioral impulsivity measures. Our findings offer practical suggestions on how behavioral measures can be used to assess trait-level impulsivity and theoretical insights for other fields with a similar problem of diverging self-report and behavioral measures of the same construct.
: This study showcases how genetic algorithms can be used to derive a short screening scale out of longer assessment instruments, in cultural adaptations. Leveraging a sample of 1,081 participants aged 6-18 years collected with the Conners-3 Parent version in Romania, we employed a genetic algorithm to identify an optimal item subset that maximizes screening accuracy for ADHD. The adaptation process balanced psychometric rigor with cultural sensitivity, comparing the newly derived short scales to the original C3AI. Validation analyses, including confirmatory factor analysis (CFA), reliability estimates (Cronbach's alpha and McDonald's omega), and receiver operating characteristic (ROC) analysis, demonstrated strong diagnostic power, acceptable reliability, and structural validity. Notably, the Romanian-adapted scale exhibited improvements in model fit and diagnostic utility compared to the original short screening form. This research highlights the potential of automated methodologies, such as genetic algorithms, to enhance cultural adaptations of psychological measures while retaining psychometric integrity. The findings underscore the utility of context-specific scale redevelopment to optimize clinical screening tools.
Evidence on the structure of the Depression Anxiety Stress Scales for Youth (DASS-Y) is both limited and equivocal, hindering the use and interpretation of the scale scores. The present study aimed to re-examine the factor structure of the DASS-Y using exploratory structural equation modeling (ESEM) and bifactor-ESEM frameworks and to investigate whether the structure of DASS-Y responses varies by country and educational stage. Convergent validity evidence was also evaluated through associations with externalizing problems, bullying victimization, and subjective well-being scores. We recruited three samples: Serbian early adolescents (n = 412; age range = 11-14 years), Australian early adolescents (n = 616; age range = 11-14 years), and Serbian middle-to-late adolescents (n = 595; age range = 15-19 years). Both the ESEM model, allowing cross-loadings, and the bifactor-ESEM model, incorporating cross-loadings and a general distress factor, yielded meaningful solutions for the DASS-Y responses. Evidence supported the measurement invariance of both models across countries and educational stages. We conclude that more complex solutions beyond the standard three-factor model should be considered when testing the structure of DASS-Y responses.
This research assessed the psychometric properties of the Employee Resilience Scale (ERS) and examined its relationship with work-family conflict (WFC) through two studies across four samples. In Study 1, confirmatory factor analysis across Chinese (Sample 1, N = 157; Sample 2, N = 173) and United Kingdom (Sample 3, N = 270) samples revealed a clear unidimensional structure with satisfactory reliability. Construct validity was established through evidence of convergent validity (positive correlations with work resilience), nomological validity (positive correlations with work engagement), and discriminant validity (the ERS remained empirically distinct from these related constructs). Measurement invariance analyses in the MIMIC framework indicated the ERS is equivalent across gender and age within the Chinese sample, but some items exhibited differential item functioning in the UK sample and in cross-cultural comparisons, highlighting sociocultural influences. In Study 2 (Sample 4, N = 385), employee resilience showed a negative association with WFC, with this relationship showing an indirect association through meaningful work. The validated ERS is a valuable tool for advancing research and practical applications in understanding resilience in Chinese and British contexts.
: Skin shame is a clinically relevant construct reflecting a psychological burden associated with chronic skin diseases. The 24-item Skin Shame Scale (SSS-24) aims to assess this construct, but it exhibits psychometric weaknesses and its applicability in time-constrained clinical settings is limited. In this study, we therefore developed and validated a psychometrically optimized short form of the SSS-24 using Ant Colony Optimization (ACO). In Study 1 (N = 464), an optimized 8-item version of the measure (SSS-OF) was developed using ACO in individuals diagnosed with atopic dermatitis (AD) and psoriasis. In Study 2 (N = 621), the psychometric properties and construct validity of the SSS-OF were tested in an independent sample of individuals diagnosed with AD and psoriasis. The SSS-OF demonstrated high internal consistency, good fit to a unidimensional model, and measurement invariance across the two different skin diseases in both samples. The scale converged with other measures of shame and was associated with psychological distress, reduced dermatological quality of life, and measures of (self-)disgust. The SSS-OF is a valid, reliable, and time-efficient instrument for assessing skin-related shame that outperforms its long form and can be recommended for research and clinical screening.
Valid instruments for assessing students' causal attributions of achievement are essential for capturing students' explanations of academic success and failure, which may influence their performance at school. Based on a sample of 467 German students, this study aimed to test the validity of the Scales for the Assessment of Causal Attributions for Achievement (SACA). These scales measure five causal attributions for both academic success and failure: ability, effort, task difficulty, luck, and perceived teacher popularity. Confirmatory factor analyses provided strong support for a ten-factor structure. Measurement invariance across gender and grade levels was fully established, indicating that the ten attributions can be equivalently measured across gender and grade levels by the SACA. Construct validity was supported through associations with various external criteria (i.e., academic self-concept, academic achievement, and achievement goals). Success attributions to ability, effort, and perceived teacher popularity showed the strongest positive correlations with academic self-concept, academic achievement, mastery goals, and performance-approach goals, while all failure attributions were negatively associated with these outcomes. Multiple regression analyses confirmed the SACA's predictive utility by assessing the unique contribution of each attribution dimension while controlling for gender and grade level.
: Modern information and communication technologies provide employees with greater flexibility, but they also blur the boundaries between work and personal life. Work-related extended availability (WREA) refers to the availability of workers for work-related matters and the availability of work tasks for workers beyond the boundary of the work domain. In the present study, we developed and validated a new instrument for assessing WREA. In Sample 1 (N = 310), we tested the initial item pool, and in Sample 2 (N = 591), we examined the nomological network of the finalized measure using data collected at two time points. Exploratory and confirmatory factor analyses supported a unidimensional structure. The scale demonstrated high correlations with conceptually closely related constructs, indicating convergent validity, and lower correlations with more distantly related constructs supported discriminant validity. Furthermore, associations with self-reported real-life behaviors, such as accepting work-related calls during personal time, and incremental validity in predicting work-to-family conflict demonstrated the scale's criterion validity. Overall, the scale offers a brief and psychometrically robust measure of WREA.
In psychological assessment, vocabulary tests are commonly used as reliable and efficient indicators of crystallized intelligence, as retrospective proxies for premorbid intelligence, and as measures of language proficiency. However, many of the widely used German vocabulary tests are outdated, proprietary, and lack a clear rationale for item selection. To address these limitations, we developed a new, openly available vocabulary test: the Next-Generation Open Vocabulary Assessment (NOVA). Therefore, we first constructed 110 multiple-choice vocabulary items with support from ChatGPT and administered them to 1,052 German-speaking adults using a multiple-matrix design, along with a declarative knowledge test for validation purposes. In a second step, we used Ant Colony Optimization to compile two parallel 30-item short forms, optimized for reliability and item difficulty and discrimination. The resulting tests assessed vocabulary unidimensionally and reliably, covered a large ability range, and correlated strongly with declarative knowledge. We provide a Shiny app for the calculation of standard values based on individual test results. Additional analyses revealed that 57% of the variance in item difficulties could be explained by word frequency and word length, which may be particularly useful for streamlining the future development of vocabulary tests.
Abstract: This study aimed to test the construct and criterion validity of the Serbian adaptation of the Short Dark Tetrad (SD4). In addition to testing measurement invariance between the Serbian ( N = 488) and Canadian samples ( N = 739), the construct validity of the SD4 was also assessed through correlations with more extensive measures of the Dark Tetrad, and criterion validity was evaluated through correlations with variables related to mental health. The results indicated good model fit indices for the SD4 in both samples and partial scalar invariance across samples. Validity correlations with extensive measures confirmed the construct validity of all SD4 scales, with caution noted for the psychopathy scale, which shares similar content with sadism. Profile similarity, based on construct and criterion correlations, revealed substantial dissimilarity between narcissism and the other scales but high similarity among the Machiavellianism, psychopathy, and sadism scales. Regardless of the similarity between the scales, they showed distinctive correlations with emotional distress and positive mental health aspects.
Abstract: Building on the long common history of board games and intelligence research, we developed a new deductive reasoning test based on the popular game Mastermind. The research questions of this registered report were: (a) Is a psychometrically sound measurement of the ability to solve Mastermind items possible (i.e., a reliable, uni-dimensional measurement with a good coverage of difficulty)? (b) Is the ability to solve Mastermind items substantially related to other measures of cognitive ability (i.e., matrix test, knowledge test) and need for cognition? (c) Can item difficulty be predicted by the number of colors, positions, premises, and a newly proposed entropy-based index? Based on the results of a pilot study, we developed 30 items and administered them to 351 participants in the preregistered main study using a multiple matrix sampling design. The deductive Mastermind test proved to be (a) a reliable and efficient measure of reasoning across a wide ability range, and (b) showed expectation-consistent patterns for the convergent and divergent measures. (c) The entropy-based index allowed for the prediction of item difficulty to a considerable degree ( R 2 = .47). We discuss the ideas of information theory, including entropy, as constructing principles for the rational test development of reasoning tests.
: The academic grit scale (AGS) is a new measure developed to assess the level of adolescents' grit in an academic-specific context. The main purpose of this study was to examine the psychometric properties, factor structure, and measurement invariance of the AGS among Chinese adolescents. A cross-sectional design and convenient sampling were conducted in a sample of Chinese adolescents (N = 619, 50.6% female, Mage = 14.56 years, SD = 1.47 years) using the AGS. Confirmatory factor analysis (CFA) supported the original unidimensional model of the AGS, and multiple-group CFA further verified the AGS scores had strong measurement invariance across gender (i.e., boys and girls) and grade level (i.e., middle school students and high school students). Moreover, the AGS also showed satisfactory internal consistency (using Cronbach's alpha, McDonald's omega, and mean inter-item correlation) and criterion validity, supported by expected relationships with scores on external variables (i.e., general grit, anxiety, and depression). In conclusion, these findings suggest that the AGS has satisfactory psychometric properties and can be a reliable tool to measure the academic grit level of Chinese adolescents.
Abstract: Speech Emotion Recognition (SER) is an umbrella term that encompasses all Machine Learning & Deep Learning (ML/DL) algorithms used for the very specific task of extracting emotional state from human speech. In literature, various techniques have been utilized to extract emotions from signals, including well-established speech analysis and classification techniques. Using the scoping review method, the paper maps techniques for speech-based emotion recognition. In doing so, it presents elements defining the use of algorithms to assess emotions, for example, databases used for emotion recognition, notions and types of emotions considered, and empirical investigations made toward SER and related limitations. The contribution places particular emphasis on the existing perspectives and practices in order to offer a series of recommendations for future developments.