Inductive item tree analysis is an established method of Boolean analysis of questionnaires. By exploratory data analysis, from a binary data matrix, the method extracts logical implications between dichotomous test items based on their positive item scores. For example, assume that we have the problems i and j of a test that can be solved or failed by subjects. With inductive item tree analysis, an implication between the items i and j can be uncovered, which has the interpretation "If a subject is able to solve item i, then this subject is also able to solve item j". Hence, in the current form of the method, (a) solely dichotomous items are considered, and (b) conclusions are drawn from only positive item scores. In this paper, we provide extensions to these restrictions. First, as remedy for (b), we focus on the dichotomous formulation of the inductive item tree analysis algorithm and describe a procedure of how to extend the dichotomous variant to also include negative item scores. Second, to address (a), we further extend our approach to the general case of polytomous items, when more than two answer categories are possible. Thus, we introduce extensions of inductive item tree analysis that can deal with nominal polytomous and ordinal polytomous answer scales. To show their usefulness, the dichotomous and polytomous extensions proposed in this paper are illustrated with empirical data and in a simulation study.
Mosaic plots are state-of-the-art graphics for multivariate categor ical data in statistical visualization. Knowledge structures are mathematical models that belong to the theory of knowledge spaces in psychometrics. This paper presents an application of mosaic plots to psychometric data arising from underlying knowledge structure models. In simulation trials and with empirical data, the scope of this graphing method in knowledge space theory is investigated.
This paper describes the technique of exploratory latent class cluster analysis. The classical analysis is a model-based statistical approach for identifying unobserved subgroups from observed categorical data and for classifying cases into the identified subgroups based on membership probabilities estimated directly from the statistical model. In the first part on mathematical modeling of the paper, we introduce the data and the sampling distribution for the data as required in the analysis of latent classes, the fundamental model assumptions are reviewed, and the general unrestricted latent class model is presented. Classification of cases into the clusters using modal assignment is discussed. In the second part on inferential statistics of the paper, we briefly review the classical maximum likelihood methodology related to parameter estimation and model testing, and the information criteria AIC and SIC for model selection. In the third part on case study of the paper, the General Social Survey data are analyzed using the software Latent GOLD. We present the Latent GOLD profile plot and tri plot options for the graphical representation of the results. The Latent GOLD classification output illustrating the assignment of respondents to the latent survey respondent types is also shown.
A modified particle is disclosed wherein a particle has an attached group having the formula: wherein Ar represents an aromatic group; R1 represents a bond, an arylene group, an alkylene group wherein R4 is an alkyl or alkylene group or an aryl or arylene group; R2 and R3, which can be the same or different, represent hydrogen, an alkyl group, an aryl group, -OR5, -NHR5, -NR5R5, or -SR5, wherein R5, which is the same or different, represents an alkyl group or an aryl group; and Q represents a labile halide containing species. Also disclosed is a modified particle or aggregate having attached a group having the formula: wherein CoupA represents a Si-containing group, a Ti-containing group, or a Zr-containing group; R8 and R9, which can be the same or different, represent hydrogen, an alkyl group, an aryl group, -OR10, -NHR10, -NR10R10, or -SR10, wherein R10 represents an alkyl group or an aryl group; Q represents a labile halide containing species; and n is an integer of from 1 to 3. Modified particles with attached polymers are also disclosed as well as methods of making the modified particles.
The two-sample problem for Cronbach's coefficient $\alpha_C$, as an estimate of test or composite score reliability, has attracted little attention, compared to the extensive treatment of the one-sample case. It is necessary to compare the reliability of a test for different subgroups, for different tests or the short and long forms of a test. In this paper, we study statistically how to compare two coefficients $\alpha_{C,1}$ and $\alpha_{C,2}$. The null hypothesis of interest is $H_0 : \alpha_{C,1} = \alpha_{C,2}$, which we test against one-or two-sided alternatives. For this purpose, resampling-based permutation and bootstrap tests are proposed. These statistical tests ensure a better control of the type I error, in finite or very small sample sizes, when the state-of-affairs \textit{asymptotically distribution-free} (ADF) large-sample test may fail to properly attain the nominal significance level. We introduce the permutation and bootstrap tests for the two-group multivariate non-normal models under the general ADF setting, thereby improving on the small sample properties of the well-known ADF asymptotic test. By proper choice of a studentized test statistic, the resampling tests are modified such that they are still asymptotically valid, if the data may not be exchangeable. The usefulness of the proposed resampling-based testing strategies is demonstrated in an extensive simulation study and illustrated by real data applications.
This paper introduces the R package CDM for cognitive diagnosis models (CDMs). The package implements parameter estimation procedures for two general CDM frameworks, the generalized-deterministic input noisy-and-gate (G-DINA) and the general diagnostic model (GDM). It contains additional functions for analyzing data under these frameworks, like tools for simulating and plotting data, or for evaluating global model and item fit. The paper describes the theoretical aspects of implemented CDM frameworks and it illustrates the usage of the package with empirical data of the common fraction subtraction test by Tatsuoka (1984).
Inductive Item Tree Analysis (IITA) comprises three data analytic algorithms for deriving reflexive and transitive precedence relations (surmise relations or quasi-orders) among binary items. With the help of simulation studies, the IITA algorithms were already compared concerning their ability to detect the correct precedence relations in observed data. These studies generate a set of surmise relations on an item set, simulate a data set from each of the surmise relations by applying some random response errors, and then try to recover the initial surmise relations from those noisy data. We show that, in the currently published studies however, the representativeness of sampled quasi-orders was not considered or implemented unsatisfactorily. This led to non-representative samples of quasiorders, and hence to biased or wrong conclusions about the quality of the IITA algorithms to reconstruct the underlying surmise relations. In our paper, results of a new, truly representative simulation study are reported, which correct for the problems. On the basis of this study, the three IITA algorithms can now be compared reliably.
Recently, performance profiles in reading, mathematics and science were created using the data collectively available in the Trends in International Mathematics and Science Study (TIMSS) and the Progress in International Reading Literacy Study (PIRLS) 2011. In addition, a classification of children to the end of their primary school years was conducted in accordance with these performance types. To create performance profiles and classifications, multidimensional item response theory and latent profile analysis were used. The focus in this study is on the comparison and usability of clustering methods in their application in large-scale assessments. In a first step, the cluster solutions of classic approaches such as kmeans, fuzzy c-means and hierarchical procedures are compared to one another and assessed in terms of their proximity to the reference typology of latent profile analysis. In the second step, the results of the model-supported latent profile analysis are compared directly with the findings of the classic model-free cluster analyses by means of appropriate measured values. The result is a high consistency in the classification of invariant ranked profiles. In the last step, the calculated “quantitative” cluster solutions are compared to the “qualitative” typology of the pupils derived from the content-based benchmarking of the competency levels in mathematics using the TIMSS guidelines. It is evident that, as a cluster solution, the benchmarking breakdown of the sample in the five competency levels does not show a high goodness of fit with the available data.
In psychometric latent variable modeling approaches such as item response theory one of the most central assumptions is local independence (LI), i.e. stochastic independence of test items given a latent ability variable (e.g., Hambleton et al., Fundamentals of item response theory, 1991). This strong assumption, however, is often violated in practice resulting, for instance, in biased parameter estimation. To visualize the local item dependencies, we derive a measure quantifying the degree of such dependence for pairs of items. This measure can be viewed as a dissimilarity function in the sense of psychophysical scaling (Dzhafarov and Colonius, Journal of Mathematical Psychology 51:290–304, 2007), which allows us to represent the local dependencies graphically in the Euclidean 2D space. To avoid problems caused by violation of the local independence assumption, in this paper, we apply a more general concept of “local independence” to psychometric items. Latent class models with random effects (LCMRE; Qu et al., Biometrics 52:797–810, 1996) are used to formulate a generalized local independence (GLI) assumption held more frequently in reality. It includes LI as a special case. We illustrate our approach by investigating the local dependence structures in item types and instances of large scale assessment data from the Programme for International Student Assessment (PISA; OECD, PISA 2009 Technical Report, 2012).
Multicollinearity is one of the main problems when using regression analytic approaches to predict outcome variables. The application of traditional regression analytic approaches often provides unstable and unreliable estimates of the parameters when multicollinearity occurs. In this paper we apply a regression analytic method called correlated component regression (CCR), developed by Magidson (Correlated component regression: re-thinking regression in the presence of near collinearity. In Abdi et al. (eds) New perspectives in partial least squares and related methods, Springer, Heidelberg, pp 65–78, 2013), for characterizing student performances in PIRLS/TIMSS 2011 (Martin and Mullis, Methods and procedures in TIMSS and PIRLS 2011, TIMSS & PIRLS International Study Center, Chestnut Hill, 2013) through selected background characteristics, such as cultural and socio-economic characteristics. On the basis of various criteria, we compare the findings of CCR with the results of OLS regression regarding the prediction of student performance values. An implemented cross-validation procedure and step-down algorithm are utilized to perform a special type of variable reduction. Thus, the results of our study will provide more reliable sets of background variables for characterizing large scale educational data in the domains of reading, mathematics, and science.
Fechnerian scaling as developed by Dzhafarov and Colonius (e.g., Dzhafarov and Colonius, J Math Psychol 51:290–304, 2007) aims at imposing a metric on a set of objects based on their pairwise dissimilarities. A necessary condition for this theory is the law of Regular Minimality (e.g., Dzhafarov EN, Colonius H (2006) Regular minimality: a fundamental law of discrimination. In: Colonius H, Dzhafarov EN (eds) Measurement and representation of sensations. Erlbaum, Mahwah, pp. 1–46 ). In this paper, we solve the problem of correcting a dissimilarity matrix for Regular Minimality by phrasing it as a convex optimization problem in Euclidean metric space. In simulations, we demonstrate the usefulness of this correction procedure.
The Programme for International Student Assessment (PISA; e.g., OECD, Sample tasks from the PISA 2000 assessment, 2002a; OECD, Learning for tomorrow’s world: first results from PISA 2003, 2004; OECD, PISA 2006: Science competencies for tomorrow’s world, 2007; OECD, PISA 2009 Technical Report, 2012) is an international large scale assessment study that aims to assess the skills and knowledge of 15-year-old students, and based on the results, to compare education systems across the participating (about 70) countries (with a minimum number of approx. 4,500 tested students per country). Initiator of this Programme is the Organisation for Economic Co-operation and Development (OECD; www.pisa.oecd.org ). We review the main methodological techniques of the PISA study. Primarily, we focus on the psychometric procedure applied for scaling items and persons. PISA proficiency scale construction and proficiency levels derived based on discretization of the continua are discussed. For a balanced reflection of the PISA methodology, questions and suggestions on the reproduction of international item parameters, as well as on scoring, classifying and reporting, are raised. We hope that along these lines the PISA analyses can be better understood and evaluated, and if necessary, possibly be improved.
For scaling items and persons in large scale assessment studies such as Programme for International Student Assessment (PISA; OECD, PISA 2009 Technical Report. OECD Publishing, Paris, 2012) or Progress in International Reading Literacy Study (PIRLS; Martin et al., PIRLS 2006 Technical Report. TIMSS & PIRLS International Study Center, Chestnut Hill, 2007) variants of the Rasch model (Fischer and Molenaar (Eds.), Rasch models: Foundations, recent developments, and applications. Springer, New York, 1995) are used. However, goodness-of-fit statistics for the overall fit of the models under varying conditions as well as specific statistics for the various testable consequences of the models (Steyer and Eid, Messen und Testen [Measuring and Testing]. Springer, Berlin, 2001) are rarely, if at all, presented in the published reports. In this paper, we apply the mixed coefficients multinomial logit model (Adams et al., The multidimensional random coefficients multinomial logit model. Applied Psychological Measurement, 21, 1-23, 1997) to PISA data under varying conditions for dealing with missing data. On the basis of various overall and specific fit statistics, we compare how sensitive this model is, across changing conditions. The results of our study will help in quantifying how meaningful the findings from large scale assessment studies can be. In particular, we report that the proportion of missing values and the mechanism behind missingness are relevant factors for estimation accuracy, and that imputing missing values in large scale assessment settings may not lead to more precise results.
Schrepp (2005) points out and builds upon the connection between knowledge space theory (KST) and latent class analysis (LCA) to propose a method for constructing knowledge structures from data. Candidate knowledge structures are generated, they are considered as restricted latent class models and fitted to the data, and the BIC is used to choose among them. This article adds additional information about the relationship between KST and LCA. It gives a more comprehensive overview of the literature and the probabilistic models that are at the interface of KST and LCA. KST and LCA are also compared with regard to parameter estimation and model testing methodologies applied in their fields. This article concludes with an overview of KST-related publications addressing the outlined connection and presents further remarks about possible future research arising from a connection of KST to other latent variable modeling approaches.