Bayesian conditional transformation models (BCTMs) address the direct estimation of the conditional distribution function of a random variable Y $Y$ conditional on a set of explanatory variables X ${\rm variables}\ \bm{X}$ . The BCTMs infer the conditional distribution by applying a transformation function of Y $Y$ given X = x $\bm{X} = \bm{x}$ towards a baseline distribution free of parameters to be estimated. The benefit of these models is that the explanatory variables X = x $\bm{X} = \bm{x}$ impact the whole conditional distribution of Y $Y$ given X = x $\bm{X} = \bm{x}$ instead of only the mean, variance, kurtosis, or skewness. The transformation functions are an essential part of the model, and they range from loss-complex and low-parameterized functions to complex relationships between explanatory variables and response variables represented by nonlinear functions. The general construction of the BCTM class explores monotonic B-splines for parameterizing the transformation function. Smoothness and regularization are accomplished through an adequate prior distribution for the parameters. We proposed a new estimation procedure for the BCTM based on the integrated nested Laplace approximation, which is tested through a simulation study. Also, two longitudinal studies using real data are considered. The first application is a cardiovascular study and compares our proposed algorithm, named integrated Laplace with Bayesian conditional transformation models (ILBCTM), with the original Markov chain Monte Carlo-based algorithm for BCTM. We obtained similar results with a shorter computational time. The second application considers the ILBCTM in a study of the mortality rate of bronchial and lung cancer in Brazil.
Cognitive Diagnosis Models (CDMs) are widely used in latent-variable modeling for classification tasks that diagnose abilities or skills. Originally developed for dichotomous indicators, CDMs have been extended to polytomous and continuous responses, including bounded continuous variables (e.g. proportions or index scores on a 0-1 or 0-100 scale). We introduce a Bounded DINA (B-DINA) model, an extension of DINA for handling bounded continuous responses, using a Beta distribution with an appropriate mean-precision parameterization. We present a Bayesian estimation framework, define interpretable item parameters and compute posterior probabilities of membership in each latent-attribute profile. We explicitly address label-switching nonidentifiability and assess absolute model fit via posterior predictive p $$ p $$ -values (PPP). Also, we have conducted a simulation study to evaluate parameter recovery for our proposed method and its performance. Further, we illustrate the model mainly with municipal data from Southeastern Brazil, where bounded indices summarize economy, education and health. Our proposed B-DINA effectively classifies municipalities and reveals relationships between observed indicators and latent attributes. As bounded continuous variables are common across the social sciences and policy analysis, our proposed B-DINA could offer a broadly applicable classification tool in the practice.
Imbalanced binary data may be more common than expected in medical trials. In this paper, we propose a new class of link function for binary response based on the cumulative distribution function of the scale mixture of skew-normal distributions, which can be useful for fitting imbalanced binary data. The proposed link class has as special cases several link functions proposed in the literature, such as the probit and Student’s-t link, and we present a Bayesian approach for model fitting. Further, we develop Bayesian case-deletion influence diagnostics based on the Kullback-Leibler divergence. The newly developed procedures are illustrated with one example as well as a simulation, which illustrates the potential of the proposed class of links as an alternative for binary regression models when imbalanced binary data is presented.
This work proposes a new quantile regression model with a bounded response distribution that generalizes the L-logistic distribution. Following a Bayesian approach, estimation model comparison criteria and residual analysis are performed as well as a simulation study for prior sensitivity and parameter recovery, considering a computationally intensive approach. An application of the new distribution to model poverty vulnerability in Brazil and a regression analysis with poverty data from Peru is included. Comparison with the Beta and L-Logistic distributions are also performed showing the great flexibility of the new model.
In this article, we consider the discrete Bell distribution to introduce a new mixed-effects regression model that may be an interesting alternative to traditional mixed-effects models for count response variables. The new regression model can be applied in several areas including health data. We consider the frequentist and Bayesian approaches to perform inferences in this class of mixed regression models. We provide Monte Carlo simulation experiments to verify the performance of these approaches in estimating the mixed Bell regression model parameters. The simulation results are quite promising and indicate that these approaches are effective in doing that. We also consider model comparison criteria based on the frequentist and Bayesian approaches and simulations are considered to verify the performance of these criteria. Two empirical applications to real data of the proposed mixed-effects model are provided, and comparisons with the Poisson mixed-effects model, as well as the Poisson inverse Gaussian mixed-effects model, are made. The real data applications confirm that the proposed mixed-effects Bell regression model can be an interesting alternative in the modeling of count response variables.
Multidimensional item response theory (MIRT) models estimate multiple latent traits from individual item response data and have been applied to problems in several knowledge areas. This article explores some statistical aspects about the compensatory multidimensional graded response model with two hierarchical structures and Q-matrix. Specifically, the efficiency of the estimation method using No-U-Turn Sampler (NUTS) algorithm and the performance of the three model comparison criteria regarding the adequacy of the models with the same dataset were verified. An application was conducted with recent data from responses to a questionnaire on sustainability perception applied to residents in Brazil verified the model that best fit the data, thus justifying the choice of the appropriate model for the research. The results were used for the selection of a model for the application. The explored models provided detailed information about individuals, which may be useful for the design of public policies aimed at the development of specific educational actions for the economic, environmental and social fields. Codes are available for the proposed methodology to be explored in other applications.
The current research on evaluating the response modification factor R, related to the lateral strength demand of structures, has been generally based on inelastic single-degree-of-freedom systems. Nevertheless, most structures have more than one main analysis component and will be subjected to bidirectional ground motions. In this study, the results of the parametric evaluation of the response modification factor considering the bidirectional interaction (Rb) of inelastic two-degree-of-freedom systems (2DOF) are presented. The effects on the factor Rb of the vibration period, the ductility capacity, the hysteretic model, the seismic incidence angle, and the period ratio were evaluated. Analysis results show that the bidirectional interaction could increase the lateral strength demand of the 2DOF systems because of the coupling effect of the two components’ responses. To make this research useful for improving engineering practice and code provisions, the main contribution is the proposal of a simplified expression for estimating the factor Rb.
The choice of a prior distribution is a key aspect of the Bayesian method. However, in many cases, such as the family of power links, this is not trivial. In this article, we introduce a penalized complexity prior (PC prior) of the skewness parameter for this family, which is useful for dealing with imbalanced data. We derive a general expression for this density and show its usefulness for some particular cases such as the power logit and the power probit links. A simulation study and a real data application are used to assess the efficiency of the introduced densities in comparison with the Gaussian and uniform priors. Results show improvement in point and credible interval estimation for the considered models when using the PC prior in comparison to other well-known standard priors.
A Q-matrix is a binary matrix that defines the relationship between items and latent variables and is widely used in diagnostic classification models (DCMs), and can also be adopted in multidimensional item response theory (MIRT) models. The construction process of the Q-matrix is typically carried out by experts in the subject area of the items and statistical procedures can be used to verify its suitability. In DCMs, different approaches have been proposed for validating the Q-matrix through iterative algorithms. This article presents a method of empirical Q-matrix validation for MIRT models. A simulation study and an application to real data of morphology skills of elementary students are conducted to examine the viability of the method. Relevant issues regarding the implementation of the method and the results obtained are discussed.
In addition to the usual slope and location parameters included in a regular two-parameter logistic model (2PL), the logistic positive exponent (LPE) model incorporates an item parameter that leads to asymmetric item characteristic curves, which have recently been shown to be useful in some contexts. Although this model has been used in some empirical studies, an identifiability analysis (i.e., checking the (un)identified status of a model and searching for identifiablity restrictions to make an unidentified model identified) has not yet been established. In this paper, we formalize the unidentified status of a large class of fixed-effects item response theory models that includes the LPE model and related versions of it. In addition, we conduct an identifiability analysis of a particular version of the LPE model that is based on the fixed-effects one-parameter logistic model (1PL), which we call the 1PL-LPE model. The main result indicates that the 1PL-LPE model is not identifiable. Ways to make the 1PL-LPE useful in practice and how different strategies for identifiability analyses may affect other versions of the model are also discussed.
Bounded count response data arise naturally in health applications. In general, the well-known beta-binomial regression model form the basis for analyzing this data, specially when we have overdispersed data. Little attention, however, has been given to the literature on the possibility of having extreme observations and overdispersed data. We propose in this work an extension of the beta-binomial regression model, named the beta-2-binomial regression model, which provides a rather flexible approach for fitting a regression model with a wide spectrum of bounded count response data sets under the presence of overdispersion, outliers, or excess of extreme observations. This distribution possesses more skewness and kurtosis than the beta-binomial model but preserves the same mean and variance form of the beta-binomial model. Additional properties of the beta-2-binomial distribution are derived including its behavior on the limits of its parametric space. A penalized maximum likelihood approach is considered to estimate parameters of this model and a residual analysis is included to assess departures from model assumptions as well as to detect outlier observations. Simulation studies, considering the robustness to outliers, are presented confirming that the beta-2-binomial regression model is a better robust alternative, in comparison with the binomial and beta-binomial regression models. We also found that the beta-2-binomial regression model outperformed the binomial and beta-binomial regression models in our applications of predicting liver cancer development in mice and the number of inappropriate days a patient spent in a hospital.
Recommendation Systems have become prevalent in recent years, attracting the attention of researchers to investigate different methods to filter relevant information for users. This information is not always explicit and different proposals have emerged to obtain the latent values of individuals through their behavior. In educational areas, latent attributes of test-takers can be acquired by psychometric models such as the Cognitive Diagnostic Model. These models attempt to create a user's profile in order to explore the connections between students and subjects, just like a recommendation system does with its users and the products to be recommended. The objective of this work is to develop a new recommendation approach that incorporates Cognitive Diagnostic Models applied to data from media defined by discrete content (such as genres in movies and series) in order to generate its polytomous response in the form of the rating prediction that a user would give to each item. The proposed approach was applied to two datasets (MovieLens20M Dataset and Anime Recommendation Database). The new proposal was also considered with additional information regarding the popularity of the items, in an enhanced version of our model, and compared to classic recommendation systems found in the literature. Finally, this work also explored the performance of the models in ranking items to be recommended for the users. In general, the new method obtained better results than the classic recommendation ones for both the predicted rating and the item ranking.
This paper investigates the effectiveness of various metrics for selecting the adequate model for binary classification when data is imbalanced. Through an extensive simulation study involving 12 commonly used metrics of classification, our findings indicate that the Matthews Correlation Coefficient, G-Mean, and Cohen’s kappa consistently yield favorable performance. Conversely, the area under the curve and Accuracy metrics demonstrate poor performance across all studied scenarios, while other seven metrics exhibit varying degrees of effectiveness in specific scenarios. Furthermore, we discuss a practical application in the financial area, which confirms the robust performance of these metrics in facilitating model selection among alternative link functions.
The G-DINA model is a versatile cognitive diagnostic model used for individuals' classification. In this paper, we propose a Bayesian formulation of the G-DINA model and an estimation method using an MCMC Gibbs Sampling algorithm, implemented using the software JAGS. A simulation study was designed to evaluate the parameters recovery and the estimation accuracy of the Bayesian implementation of G-DINA and the results were compared with the standard frequentist approach in four scenarios. The results show that the proposed Bayesian implementation recovers all parameters and has good accuracy in the estimation, with performance similar to or better than the frequentist approach in all simulated scenarios. As an application, we propose a new methodology of classification of depression for respondents to the Beck Depression Inventory (BDI) based on an existing one, by replacing the original DINA model for the G-DINA. A comparative analysis of the application of these two methods in a data set from 1111 respondents of the BDI test was conducted. The results indicate that the methodology with the G-DINA model is more suitable for this application.
This article provides a review of the One Parameter Logistic Ability-based Guessing (1PL-AG) model (San Martín et al., 2006), which belongs to the family of ability-based guessing models within the Item Response Theory (IRT) framework. The model considers both the characteristics of the test items and the abilities of the individuals when estimating the probability of a correct guess and incorporates a general discrimination parameter to account for item difficulty. A comprehensive model that encompasses the 1PL-AG model as a specific instance, while employing a general Cumulative Distribution Function (CDF) as item characteristic curve (ICC) is introduced. Additionally, we explore another scenario, referred to as the One Parameter Normal Ogive Ability-based Guessing (1PNO-AG) model. 1PNO-AG employs the standard normal distribution as its link function and then a Bayesian approach is developed. Results considering simulations indicated that the R code developed with the use of JAGS successfully recovered the true parameter values. From an applied perspective, we compare the results obtained from applying various alternative models to a real dataset. We observed that the 1PNO-AG model exhibited superior performance in terms of the Deviance Information Criterion (DIC).
In the current literature on latent variable models, much effort has been put on the development of dichotomous and polytomous cognitive diagnostic models (CDMs) for assessments. Recently, the possibility of using continuous responses in CDMs has been brought to discussion. But no Bayesian approach has been developed yet for the analysis of CDMs when responses are continuous. Our work is the first Bayesian framework for the continuous deterministic inputs, noisy, and gate (DINA) model. We also propose new interpretations for item parameters in this DINA model, which makes the analysis more interpretable than before. In addition, we have conducted several simulations to evaluate the performance of the continuous DINA model through our Bayesian approach. Then, we have applied the proposed DINA model to a real data example of risk perceptions for individuals over a range of health-related activities. The application results exemplify the high potential of the use of the proposed continuous DINA model to classify individuals in the study.
Continuous clustered proportion data often arise in various areas of the social and political sciences where the response variable of interest is a proportion (or percentage). An example is the behavior of the proportion of voters favorable to a political party in municipalities (or cities) of a country over time. This behavior can be different depending on the region of the country, giving rise to groups (or clusters) with similar profiles. For this kind of data, we propose a finite mixture of a random effects regression model based on the L-Logistic distribution. A Markov chain Monte Carlo algorithm is tailored to obtain posterior distributions of the unknown quantities of interest through a Bayesian approach. To illustrate the proposed method, with emphasis on analysis of clusters, we analyze the proportion of votes for a political party in presidential elections in different municipalities observed over time, and then identify groups according to electoral behavior at different levels of favorable votes.
In 2010, the Samejima-Bolfarine-Bazan (SBB) Item Response Theory (IRT) models were introduced by (Journal of Educational and Be-havioral Statistics 35 (2010) 693-713) under a Bayesian approach. These models extend the regular Bayesian One and Two Parameter Logistic IRT models by incorporating a parameter accounting for asymmetry of the Item Characteristic Curve (ICC) which is named the complexity of the item. It includes the Logistic Positive Exponent (LPE) IRT model formulated ini-tially by (Psychometrika 65 (2000) 319-335) and the Reflection of the LPE (RLPE). In the present work, new properties of the SBB models are devel-oped including a random effect for testlet structures with a Bayesian inference through a Markov chain Monte Carlo (MCMC) algorithm which includes the parameter estimation and model comparison. The asymmetric behavior of the Item Characteristic Curve (ICC) is detected using a marginal item informa-tion function. Two simulation studies are developed to analyze the sensitive-ness of the penalized parameter in the asymmetric behavior of the ICC and to evaluate the parameter recovery of the proposed model. A real data set, with a testlet structure and empirical evidence of asymmetric behavior of the ICCs, is used to apply the models.
Some asymmetric Item Characteristic Curves (ICC)s have already been introduced in the IRT literature. These proposals include a new item parameter associated with the item complexity which explains the asymmetry in the ICC. Although the importance of proposing new models that have asymmetric ICC in IRT is already known, the relationship between these models and unbalanced binary responses in testing data in real applications has not been explored. In this work we propose new asymmetric IRT models that have an asymmetric ICC as their main feature. A special case of these models is the cloglog IRT model. Bayesian estimation of the proposed models is discussed and one application in educational data illustrates the benefits of the new ICC when we compare our IRT models with other IRT models proposed in the literature.
The Q-matrix is commonly used in diagnostic classification models and has recently been incorporated into the multidimensional item response theory (MIRT) models to add information about the relationship between items and dimensions of the latent trait. The reformulation of the MIRT models with Q-matrix (MIRT-Q) has presented to improve the precision of the parameters of these models and to provide a simple and intuitive method for users to define the item-trait relationship. This paper aims to explore the incorporation of the Q-matrix in the formulation of MIRT models for polytomous item responses. Specifically, we introduce the incorporation of the Q-matrix into two of the polytomous MIRT models most known and used: the multidimensional graded response (MGR) model, hereinafter called MGR-Q, and the multidimensional generalized partial credit (MGPC) model, hereinafter called MGPC-Q. We provide readers the code of the MGR-Q and MGPC-Q models in Stan, a Bayesian estimation software, and we conduct a simulation study in order to evaluate the parameter recovery of the estimation method. To illustrate the use of both models in practice, we fit them to an operational dataset from 2400 individuals on 13 items and demonstrate the estimation of MGR-Q and MGPC-Q using the Stan program.