
Comparative judgment has recently been explored as an alternative to traditional marking in education, offering advantages in certain contexts. Yet, the sheer number of possible script pairs makes full pairwise comparison impractical. Existing adaptive algorithms often lack a formal foundation, raising concerns about fairness in assessment. To address this, we examine comparative judgment through the Bradley-Terry model and propose an adaptive optimal design for paired comparisons. Drawing on experimental design theory, we develop efficient methods for selecting sets of pairs, adapting the multiple incomplete cyclic design and general equivalence theorem to ensure D-optimality. This framework provides a principled approach for constructing and validating adaptive comparative judgment. We evaluate the method using simulated data, comparing it against full factorial paired comparison. Results reveal similar patterns in perceived script quality across both approaches, but the adaptive method requires far fewer comparisons, offering a more efficient and equitable solution for educational assessment.
Composite indicators are mathematical tools that aid in understanding complex realities, such as poverty, sustainability, and the health system. This study explores the limitations of Principal Component Analysis in reconciling three fundamental elements for representing multidimensional phenomena that present weakly intercorrelated data. First, informational power, estimated by the average variance extracted from the sub-indicators, which corresponds to the proportion of information from the input data retained in the composite indicator. Secondly, interpretability, measured by the number of Principal Components needed to represent the multidimensional phenomenon, indicates how easy it is to understand and interpret. Third, explanatory power, measured by the correlation between the composite indicator and a conceptually significant sub-indicator, reflects its ability to capture the conceptual framework of the multidimensional phenomenon. In addition, the study develops an adaptive approach that allows reconciling interpretability, informational power, and explanatory power in the representation of social exclusion in the Brazilian city of Maring & aacute;.
This paper proposes a new two-parameter lifetime distribution, called the Power Xgamma exponential distribution (PXGED), obtained by applying a power transformation to the Xgamma exponential distribution introduced by Yadav et al. (2022). The PXGED provides enhanced flexibility for modeling diverse lifetime and reliability data and is particularly suitable for accommodating monotone/non-nonotone hazard rate patterns, including both increasing and decreasing shapes commonly encountered in survival and engineering applications. The characterization of the proposed distribution is carried out by studying several important statistical properties, including moments, the moment-generating function, order statistics, stochastic ordering, and entropy measures. Parameter estimation is carried out under both classical and Bayesian frameworks using doubly type-II censored data. In addition, optimal censoring schemes are determined using different optimality criteria. The performance of the proposed estimators is examined through an extensive Monte Carlo simulation study. Finally, the practical applicability and robustness of the PXGED model are demonstrated using three real-world datasets.
Geostatistical data in epidemiology, hydrology, and environmental studies frequently encounter detection limits, resulting in censored measurements (left-, right-, or interval-censored). Addressing these challenges requires specialized statistical methods for accurate inference and prediction. Although classical approaches, such as the Expectation-Maximization algorithm, Monte Carlo EM, and Stochastic Approximation EM have been employed, this study presents a Bayesian geostatistical framework that leverages Stan's No-U-Turn Sampler to obtain posterior simulations for parameter estimation. The performance of the methodology was evaluated through both uncensored and censored simulation studies, and further validated with real contamination datasets: dioxin in Missouri and arsenic in Michigan. Our results yielded credible intervals for model parameters, enabling robust inference and the identification of optimal covariance structures.
Recent work has explored dyad-network communication as a contributing factor in relationship development and outcomes. The present study is the first to operationalize dyad-network communication. Specifically, we move toward a measure of dyad-network communication. In Study 1, open-ended responses were collected from 80 couples (N = 160) regarding the methods through which they converge and/or diverge when discussing or in the presence of their respective social networks - as well as how those networks accommodate to them. These responses were used to develop an 18-item Likert-style scale, comprised of three subscales. Study 2 involved three waves of data collection. An exploratory factor analysis was performed on the first wave, narrowing the total items down to 14. The second wave was subjected to confirmatory factor analysis and bivariate correlation as an initial test of convergent/divergent validity. The third wave involved advanced regression analyses as further tests of construct validity. Results are discussed in terms of measurement use as well as the theoretical salience of dyad-network communication as a measured variable.
Analytic rubrics are commonly employed in translation assessment, but it is often unclear whether their domains represent distinct measurable aspects or mainly reflect a general judgment of translation quality. This empirical validation study investigated the internal structure of analytic Persian - English translation ratings. A total of 300 Iranian translators completed translation tasks from Persian to English and English to Persian, and trained expert raters evaluated their work using a 16-indicator rubric that addressed terminology, fluency, culture, and style. Since the indicators were scored on a five-point ordinal scale, confirmatory factor analyses were performed using weighted least squares mean and variance (WLSMV) estimation. Four models were compared: unidimensional, correlated four-factor, higher-order, and bifactor models. The bifactor model best captured the data's structure. Findings indicated that the general translation-performance factor explained most of the shared variance, supporting the use of a total score. Domain-specific factors were less robust: terminology and style offered limited diagnostic insight, the culture domain offered modest and task-sensitive insight, and fluency contributed little beyond the overall score. These results suggest that the rubric is most appropriate for overall score reporting, while domain scores should be used cautiously for formative feedback rather than for high-stakes, independent evaluations.
This study aimed to develop a VO(2)max prediction framework that integrates missing data imputation, utilization of missingness patterns, residual correction - based machine learning modeling, and SHAP interpretation using large-scale national physical fitness data. Various imputation methods were compared through a masked-value recovery evaluation, and MICE (BayesianRidge) demonstrated the best restoration performance and was therefore selected as the optimal imputer. For predictive modeling, a two-stage framework, OUR-MARC (ExtraTrees + Histogram-based Gradient Boosting Regressor), was proposed, combining baseline prediction with residual correction. On the test dataset, OUR-MARC showed the best numerical performance on the held-out test dataset. (R-2 = 0.8772, RMSE = 2.1001, MAE = 1.5339). These findings suggest that explicitly incorporating missingness pattern information and residual structure contributes to improved predictive stability and mitigation of overfitting. SHAP analysis identified HR Recovery, body mass index (BMI), muscular endurance, body fat percentage, and age group as key predictors of VO(2)max. Moreover, structural differences in variable importance were observed across sex-stratified analyses, highlighting the necessity of sex-specific considerations in VO(2)max prediction. Overall, the proposed framework demonstrates that high predictive accuracy and interpretability can be simultaneously achieved in high-missingness physical fitness data environments.
Truncated distributions are useful for modeling constrained data, yet matrix-variate truncated models remain relatively underexplored compared to their univariate and vector settings. This paper proposes a finite mixture model based on the truncated matrix-variate normal (TMVN) distribution, referred to as the FM-TMVN model, which generalizes the classical truncated multivariate normal distribution to the matrix setting and enables direct modeling of matrix-structured observations under truncation. We derive the corresponding likelihood function and develop an expectation conditional maximization (ECM) algorithm for maximum likelihood estimation. Our methodology is illustrated using multispectral satellite imagery, where the observations are naturally matrix-valued and subject to bounded constraints. The empirical results demonstrate that the proposed FM-TMVN model provides improved model fit and enhanced interpretability compared with conventional finite mixture of matrix-variate normal (FM-MVN) models.
Exercise plays a fundamental role in the prevention and management of obesity by reducing cardiometabolic risks, supporting physical function, and improving quality of life. The aim of this study is to examine the validity and reliability of the Turkish version of the Volition in Exercise Questionnaire (VEQ-TR) in obese adults. This cross-sectional validation study was conducted with obese adults attending the obesity center of a training and research hospital. Data were collected from 271 participants for exploratory factor analysis and 200 participants for confirmatory factor analysis. The questionnaire included the VEQ-TR and the Short Form-36 Health Survey (SF-36), which was used to assess convergent validity. Internal consistency and test-retest reliability analyses were also performed. As a result of exploratory and confirmatory factor analyses, the VEQ-TR was found to exhibit a three-factor structure in a sample of obese Turkish adults, consisting of "Self-Confidence and Coping with Failure" (SC&CF), "Reasons and Postponing Training" (R&PT), and "Unrelated Thoughts" (UT). Confirmatory factor analysis supported the three-factor structure with acceptable fit (CFI=.962, TLI=.953), and convergent/discriminant validity indices were satisfactory (AVE=.61-.85, CR=.89-.94, HTMT=.48-.52). The internal consistency coefficients for the scale's subscales were found to be alpha=.88 for SC&CF, alpha=.87 for R&PT, and alpha=.91 for UT, indicating high reliability. VEQ-TR can be used as a valid and reliable measurement tool for assessing volitional processes related to exercise in obese adults.
In self-report studies, participants' inaccurate responses due to insufficient effort can compromise data quality. Researchers have explored diverse methods to identify inattentive responses, each presenting unique strengths and weaknesses. One of these approaches is the factor mixture model (FMM), a model-based method used to identify inattentive responses within unidimensional measures, including items with opposite semantic polarity. The objective of this simulation study was to analyze the effectiveness of the FMM in identifying attentive responders from inattentive responders under various conditions, including inattentive response behavior type, sample size, the number of negatively worded items in a scale, and prevalence of inattentiveness in a dataset. Findings from 60 unique simulation conditions indicate that the model excels in identifying straight and almost straight-lining inattentive response behavior on a 5-point rating scale, and its performance is fairly good when inattentive response behaviors are random and mixed. In particular, the prevalence of inattentive participants emerged as a significant determinant of the model's detection accuracy. However, the sample size and the number of reversed-worded items in the scale exhibited negligible effect. These insights can be leveraged by researchers and practitioners to enhance data quality in survey methods.
Chronic stress contributes to cardiovascular disease, diabetes, and mental health disorders through its cumulative physiological toll, or allostatic load (AL). Building on our previously proposed Chronic Stress Indicator (CSI), this study aims to refine, validate, and compare the CSI using a data-driven framework that integrates physiological, socioeconomic, and behavioral factors. Using data from the MIDUS II biomarker project, a nationally representative sample of U.S. adults aged 34-84, we assessed the performance of the refined CSI and traditional AL indices in predicting multiple stress-related outcomes. Advanced statistical techniques, including the Boruta feature selection algorithm and factor analysis, were applied to optimize biomarker selection and weighting. Sensitivity analyses evaluated the robustness and reliability of each construction. Compared with the original CSI, the refined model incorporating socio-behavioral variables and data-driven weighting demonstrated improved predictive performance for short-term stress outcomes, while both traditional and extended models performed well for long-term outcomes. This work advances the measurement of chronic stress by validating and enhancing the CSI, providing a robust tool for identifying at-risk populations and guiding targeted interventions.
Model fit indices can be interpreted differently depending on the fitting functions used, which vary across estimation methods. The unweighted least squares (ULS) estimator has gained increasing attention due to its relatively simple discrepancy function and its applicability to models with ordinal variables. The growing use of ULS highlights the need to better understand how model fit indices should be interpreted and how they behave under this estimation. The current study aims to clarify understanding of ULS fit indices by examining the behavior of ULS-based RMSEA, CFI, and TLI under different types and levels of model misspecification in the context of confirmatory factor analysis (CFA). Results indicate that ULS-based fit indices are more sensitive to misspecified dimensionality than other types of misspecifications. Moreover, the effects of factor loading and model size interact with the type of misspecification, such that the resulting patterns differ from those typically observed under ML estimation. Practical implications for applied researchers are provided in light of study findings.
With the increasing integration of generative artificial intelligence (GenAI) tools such as ChatGPT into higher education, the need for valid and reliable instruments to assess students' literacy in using these systems effectively, critically, and ethically is growing rapidly. In the present study, the ChatGPT Literacy Scale (ChatGPT-LS) was adapted into Turkish, and its psychometric properties were evaluated in a sample of 572 university students in T & uuml;rkiye (Mage = 21.29, SD = 3.34), of whom 75.7% were female. Confirmatory Factor Analysis confirmed the five-factor structure of the 25-item scale, and measurement invariance was supported at the configural, metric, scalar, and strict levels. The scale demonstrated high internal consistency, and Classical Test Theory (CTT) analyses provided strong evidence for construct validity. Item Response Theory (IRT) analyses further showed that the items functioned appropriately in terms of difficulty and demonstrated strong discrimination. Associations with measures of AI literacy and attitudes toward AI supported convergent validity, whereas discriminant validity findings indicated that ChatGPT literacy was related to, yet distinct from, broader AI-related constructs. Beyond validating the Turkish version of the scale, the present study demonstrates that ChatGPT literacy can be conceptualized and measured as a stable, multidimensional competency construct across contexts, while also advancing GenAI literacy measurement through the combined use of CTT, IRT, and multi-group invariance testing. These findings highlight the value of the ChatGPT-LS for comparative research, curriculum development, and the evaluation of GenAI literacy initiatives in higher education.
This study addresses the challenge of measuring multidimensional social exclusion in cities, since it cannot be adequately captured by isolated indicators such as income or education. Although the operational framework of composite indicators provides the means to synthesize multiple sub-indicators of social exclusion into a single measure, the scores generated for dozens or hundreds of urban areas hinder their interpretability. In this regard, clustering techniques, such as the k-means algorithm, use similarities in characteristics across urban areas to form smaller groups of social exclusion, thereby simplifying their interpretation and facilitating the planning of public policies. Despite these advantages, k-means has limitations. Methods for determining the ideal number of clusters yield inconsistent results and do not guarantee reliable clustering. Furthermore, the data-driven definition of the number of groups completely ignores the concept of social exclusion, which can make it difficult to interpret. This proposed Smart k-means balances reliability and interpretability by introducing two main innovations: (i) flexibility in defining the number of clusters, allowing for a range specified by the decision-maker, aligning empirical results with the conceptual construct of the multidimensional phenomenon; and (ii) an iterative procedure that identifies and excludes sub-indicators that do not contribute to the classification of social exclusion until the average silhouette width reaches a threshold of 0.50, ensuring cohesive and well-separated clusters. The method's applicability is demonstrated through an analysis of social exclusion across eight cities in the state of Paran & aacute;, Brazil, highlighting its potential to support the development of more targeted and effective public policies.