
An item bank of cognitive items targeting the spectrum of cognitive function, with a focus on advanced dementia was developed with data from residents in long term services and supports (LTSS) settings. Latent variable models, including item response theory and factor analyses were applied to fifty items from multiple cognitive screening measures assessing memory, orientation, naming, attention, calculation, and following directions. Because the intent was to use the item bank to measure overall cognition, subdomains were not modeled. Factor analyses tested for essential unidimensionality, followed by application of an item response theory-based graded response model, including differential item functioning to examine measurement equivalence. The total sample size was 6,921, with 1,874 (27.1 %) males and 5,040 (72.9 %) females. The average age was 82.5 (SD = 11.0). There were 1,089 (16.2 %) Black, 498 (7.4 %) Hispanic, and 5,136 (76.4 %) White respondents. The average number of years of education was 9.1 (SD = 5.5). Estimates of reliability across age, education, race, and sex subgroups were high with most > 0.96. Cognitive function of the participants in the study varied from no to minimal impairment (552 or 8.0 %); mild (2,000 or 28.9 %); moderate (2,686 or 38.8 %); severe (1,328 or 19.2 %) to very severe impairment (355 or 5.1 %). Items with low information and substantial DIF were recommended for removal. The final recommended item bank includes 38 items. Future work with this item bank may include development of computerized adaptive tests and short-forms.
Psychological test and assessment modeling 62 (2020) 1, S. 85-105 Padagogische Teildisziplin: Empirische Bildungsforschung;
Anadolu-Sak Intelligence Scale (ASIS) is the first intelligence scale developed and normed in Turkey. In this study, the concurrent validity of the ASIS was investigated. A total of 98 gifted and 232 average students were administered the ASIS. Gifted students’ mean IQ (135) on the ASIS was found to be above the traditional cutoff (130 IQ). ASIS general intelligence index was the best predictor of giftedness, followed by the Fluid Reasoning Index, the Crystalized Knowledge Index and the Memory Capacity Index. Correlations between the ASIS indexes and academic achievement in math, science, language and social studies for average students were high, ranging from .50 to .83. The General Intelligence Index had the highest correlations with academic achievement in all the subjects with a mean correlation of .81 at the fourth grade and .75 at the fifth grade. Results provided strong support for the validity of the ASIS in relation to academic achievement and giftedness.
Few studies have examined factors that can predict students’ mathematic ability, particularly subgroups of students who share the common characteristics that are associated with different levels of math ability. Based on PISA 2012 data from the United States and China, we used regression tree analysis to select the most salient predictors of math ability and identify the subgroups of 15-yearold students who were likely to be proficient in math ability. Based on the results from regression tree analysis, it was found that students whose math self-efficacy score was over 3.33 and their perceived positive peer math norm score was below 3.33 on a rating scale of 1 to 4 were most likely to be associated with proficient math ability. In contrast, students whose math self-efficacy score was below 2.81 were most likely to be associated with low or below-average math ability. However, for students whose math self-efficacy score was between 2.81 and 3.33, their math ability level was likely to be associated with their perceived positive peer math norm. The significance of the study is that it uniquely identified distinct subgroups of students who were more likely to be associated with different levels of math ability. Methodologically, the study demonstrated the application of regression tree analysis in studying students’ math ability.
The current study utilized a random forest regression analysis to predict post-experiment fatigue in a sample of 212 healthy participants (mean age = 20.5, SD = 2.21; 52% women) between the ages of 18 and 30 following a mildly stressful experiment. We used a total of 30 features of demographic variables, lifestyle variables, alcohol and other drug use behaviors and problems, state anxiety and depressive symptoms, and physiological indicators that were lab assessed or self-reported. A random forest regression analysis with 10-fold cross-validation resulted in accurate prediction of post-experiment fatigue (R2 equivalent = 0.93) with the average "out-of-bag" (OOB) R2 = 0.52. Not surprisingly, self-reported pre-experiment fatigue was the most important variable (54%) in the prediction of post-experiment fatigue. Feeling anxious (state anxiety) pre- and post-experiment (3%, 7%), feeling less vigorous post experiment (3%), systolic and diastolic blood pressure (3%, 2%) and LF HRV (2%) assessed at baseline, and self-reported alcohol-related problems (3%) and sleep (2%) additionally contributed to the prediction of post-experiment fatigue. Other remaining input variables had relatively minimal importance. Substantively, this study suggests that complex interactions across multiple systems domains that support regulation may be linked to fatigue. A random forest regression analysis can relatively easily be implemented with a built-in cross-validation function and reveal a web of connections undergirding health behavior and risks.
The investigation of developmental trajectories is a central goal of educational science. However, modeling and predicting complex trajectories in the context of large-scale panel studies poses multiple challenges. Statistical models oftentimes need to take into account a) potentially nonlinear shapes of trajectories, b) multiple levels of analysis (e.g., individual level, university level) and c) measurement models for the typically unobservable latent constructs. In this paper, we develop a new approach, termed the multilevel latent growth components model (ML-LGCoM) that can adequately address all three challenges simultaneously. A key feature of this new approach is that it allows researchers to test contrasts of interest among latent variables in a multilevel study. In our illustrative example, we used data from the National Educational Panel Study to model the (non-linear) development of students’ satisfaction with their academic success over four years while taking into account clusterand individual-level trajectories and measurement error.
The testlet model is a popular statistical approach widely used by researchers and practitioners to address local item dependence (LID), a violation of the local independence assumption in item response theory (IRT) which can cause various deleterious psychometric consequences. Same as other psychometric models, the utility of the testlet model relies heavily on accurate estimation of its model parameters. The two-parameter logistic (2PL) testlet model has only been systematically investigated in the psychometric literature regarding its model parameter recovery with one full information estimation methods, namely Markov chain Monte Carlo (MCMC) method, although there are other estimation methods available such as marginal maximum likelihood estimation (MMLE) and limited information estimation methods. In the current study, a comprehensive simulation study was conducted to investigate how MCMC, MMLE, and one limited information estimation method (WLSMV), all implemented in Mplus, recovered the item parameters and the testlet variance parameter of the 2PL testlet model. The manipulated factors were sample size and testlet effect magnitude, and parameter recovery were evaluated with bias, standard error, and root mean square error. We found that there were no statistically significant differences regarding parameter recovery between the three methods. When both sample size and magnitude of testlet variance were small, both WLSMV and MCMC had convergence issues, which did not occur to MCMC regardless of sample size and testlet variance. A real dataset from a high-stakes test was used to demonstrate the estimation of the 2PL testlet model with the three estimation methods. Keywords: IRT, testlet model, estimation, full-information, limited-information.
This study addresses the sample size question for multilevel latent contextual models (MLCM), which are commonly used in educational science to assess the effects of instructional quality. In terms of MLCM, only few studies have investigated whether the Bayesian toolbox helps to overcome small-sample issues. The main goal was to investigate the performance of maximum likelihood versus non, weakly, and highly informative Bayesian estimation techniques under small-sample conditions. We assumed that incorporation of prior information derived from TIMSS data would help to produce reasonable results with small samples. As expected, our results showed that the Bayesian approaches outperformed ML estimation under all conditions when informative priors were used, as these yield almost unbiased and highly accurate estimates even under unfavourable conditions (small number of level-2 groups and small group size). The study results are discussed in the light of published findings. Implications for applied educational research are derived.
The term Educational Measurement refers to the process of representing differences between persons or other entities in educational contexts in terms of numbers. This includes theory, research, and application concerning study designs, instruments, data collection, statistical analysis, and the usage of the results obtained. Educational measurement has a substantial overlap with psychometrics. The main distinction between educational measurement and psychometrics lies in the content typically focused on; educational measurement is concerned with educational aspects and psychometrics with internal psychological processes (see Jones & Thissen, 2007, for a historical overview of psychometrics).
(ProQuest: ... denotes formulae omitted.)IntroductionAt the beginning of the twentieth century, the first intelligence test was proposed by Binet. The test is known as Binet's intelligence scale. In addition to intelligence tests, however, eminent psychologists, such as Binet, Hylan and Spearman started using socalled prolonged work or continuous performance tasks (CPTs), in which subjects are required to engage in simple, repetitive activities, such as letter cancellation, detecting differences in simple shapes, adding three digits and so on. The obtained series of response (or reaction) times made it possible to examine the patterns of reaction times (Binet, 1900; Hylan, 1989). In this way, it was possible to study "the fluctuations which always occur in any persons continuous output of mental work, even when this is so devised as to remain of approximately constant difficulty" (Spearman, 1927, p. 320).Many years later, in the Netherlands, the Bourdon-Wiersma test (see Huiskamp & de Mare, 1947) was introduced, and subsequently a test version was designed for children. This test is known as the Bourdon-Vos Test (Vos, 1992). Such tests provide several measures that can be used as indicators of the mental ability of the testee. The simplest measure is the amount of time used (speed) to accomplish the test. Furthermore, because of the nature of such tests, multiple RTs are obtained for each subject. These intraindividual RTs are used to calculate the mean or median and the standard deviation in RT over series. For example, in earlier studies on Bourdon-like cancellation tests, the preferred measure of performance fluctuations was the percentile range, defined as the difference between fastest and slowest row(s) of the test. However, the ability to use different measures has caused some debate among researchers. For example, over a century ago Hylan (1898) and Binet (1900) stressed the importance of the fluctuation in the reaction times and suggested the mean deviation as a measure of performance.Several decades later, Spearman considered that oscillation was a separate universal factor in addition to what he called the general factor and perseveration (Spearman, 1927, p. 327). Indeed, Larson and Alderton (1990) reported the following: "Jensen (1982), discussing his reaction time (RT) experiments, noted that trial-to-trial variability (the standard deviation of each subject's reaction times) frequently surpassed response speed as a predictor of intelligence. That is, low aptitude individuals were excessively variable from one RT trial to the next, relative to brighter subjects. Currently, numerous studies suggest that his observation was correct and that variability has a robust statistical relationship to intelligence." This study is in line with the abovementioned suggestion and introduces theoretically based models explaining how RTs fluctuate during simple CPTs. More recently, a validation study with a new CPT test known as the Attention Concentration Test (ACT) was reported by Hotulainen et al. (2014). This study examined how attention (RT variation) correlated with and contributed to scientific reasoning (a modified version of Science Reasoning Tasks) and school achievement (GPA).Although measures of variation, such as the mean deviation and the standard deviation may be intuitively appealing for RT measures, these measures lack theoretical foundations to explain observed RT fluctuations. For example, one might rightly ask what exactly is captured in the different measures which are used. This question can be only answered by an explanatory theory of the fluctuations in the RTs. Hence, the aim of this study first is to introduce a theoretically sound conceptualization for the random variability of the RTs in the stationary part of the time series and second, to determine by testing whether the models chosen to explain the random variability of the RTs holds with the empirical data.Attention and attention concentration testsIn this study, attention is understood as a fundamental attentional capacity (Smit & van der Ven, 1995). …
The current study tests whether memory deterioration due to pro-active interference (PI) in verbal recall could be halted via block repetition potentially leading to an increased memory consolidation. We also tested whether bilinguals would be better shielded against memory deterioration than monolinguals because they constantly need to enrich their vocabulary to compensate for their smaller lexica in either language. We tested monolinguals and balanced bilinguals with an N-Back and a free verbal recall task. Repetition showed a significant main effect with a large effect size. In Study 1 (N=45), monolingual men showed less improvement in the repetition blocks, while bilingual men showed a significant doubling of their word recall on each repetition. In Study 2 (N=78), monolingual women were less likely to use the repetition opportunity to improve the word score. Thus, in both studies, a significant monolingual disadvantage showed. When the two data sets were merged (N=123), statistical effects showed that the single word list repetition had successfully and significantly increased resistance to PI, but all individual differences due to bilingualism and sex had disappeared. This supported a previous meta-analysis showing that a monolingual disadvantage does not hold in large samples with N > 100 (Paap effect).
In this primer we present a hands-on introduction to relative importance analysis as a way of exploring the relative importance of predictors in regression analysis. This method is particularly useful when predictors are correlated since it deals with issues of multicollinearity. We outline the benefits of two major approaches to relative importance, relative weights and dominance analyses, by contrasting these two relative importance analyses with correlations and multiple regressions. Based on two already published examples, we illustrate how relative importance analysis can be used to augment the interpretation of results and when relative weights importance is most appropriate. Finally, we discuss the advantages as well as the limitations of relative importance analysis on a more theoretical level. Our aim throughout is to present these analytical methods in a simple way that makes them accessible to a broad audience.
(ProQuest: ... denotes formulae omitted.)1 IntroductionThe main tools of experimental research in sociology and psychology is the theory of surveys and experiments as parts of Mathematical Statistics. Mathematical Statistics developed on the fundament of Probability Theory from the end of 19th century on. At the beginning of the 20th century, Karl Pearson and Sir Ronald Aylmer Fisher were notable pioneers of this new discipline. Fisher's book (1925) was a milestone providing experimenters such basic concepts as his well-known maximum likelihood method and analysis of variance as well as notions of sufficiency and efficiency.When we, in the sequel, speak about experiments, we understand this in the broader sense including also surveys - but see for the fundamental differences of experiments and surveys from the theory of science' point of view for instance Rasch, Kubinger, and Yanagida (2011). In concrete applications, the experiment first has to be planned, and after the experiment is finished, the analysis has to be carried out. We deal in this paper with the pre-experimental phase, i.e. the optimal planning of an experiment.Experimental designs originated in the early years of the 20-th century mainly in agricultural field experimentation. A centre was Rothamsted Experimental Station near London, where Sir Ronald Aylmer Fisher was head of the statistical department (since 1919). There he wrote one of the first books about statistical design of experiments (Fisher, 1935); a book which was fundamental, and promoted statistical technique and application.Everything presented in the following is, however, also very important and applicable in psychological research. The mathematical justification of the methods is not stressed, here, and proofs will be often barely sketched, rather omitted. Readers interested in this are referred to Rasch and Schott (2018).Fisher (1935) also outlined the problem of "Lady tasting tea", now a famous design of a statistical randomized experiment which uses Fisher's exact test and is the original exposition of Fisher's notion of a null hypothesis.We refer in the following first to Fisher's problem, that deals with soil fertility. Because soil fertility in fields varies enormously, a field is partitioned into so-called blocks (or strata in surveys) and each block subdivided into plots. It is expected that the soil within the blocks is relatively homogeneous so that the differences in the yield of the varieties planted at the plots of one block are suggested to be only due to the varieties but not due to soil differences. To ensure homogeneity of soil within blocks, the blocks must not be too large. On the other hand, the plots must be large enough so that harvesting (mainly with machines) is possible. Consequently, only a limited number of plots within the blocks is possible and only a limited number of varieties within the blocks can be tested. If all varieties can be tested in each block, we speak of a complete block design. The number of varieties is often larger than the number of plots in a block. Therefore incomplete block designs were developed, chiefly among them completely balanced incomplete block designs, ensuring that all yield differences of varieties can be estimated with equal variance using models of the analysis of variance. How all this is applicable in psychological research is shown in Rasch, Kubinger, and Yanagida (2011 ).The Experimental Designs originally developed in agriculture soon were used in medicine, in psychology and in engineering or more general in all empirical sciences. Varieties were generalized to treatments, and plots to experimental units. But even today the number v of treatments or the letter y (from yield) in the models of the analysis of variance recall us to the agricultural origin.Experimental designs are an important part in the planning (designing) of experiments. The main principles are (the three R-s):1. …
Educational Large-Scale Assessments (LSAs), such as the Programme for International Student Assessment (PISA; OECD, 2015), the Programme for the International Assessment of Adult Competencies (PIAAC; Schleicher, 2008), or the Trends in International Mathematics and Science Study (TIMSS; Mullis, Martin, Ruddock, O'Sullivan, & Preuschoff, 2009), are the objects of a growing and highly active area of research. Particular efforts are being made regarding the analysis and interpretation of corresponding results. On the one hand, LSAs provide an invaluable pool of rich data that allow for the application of complex methods to answer empirical research questions that cannot be addressed by smaller-scale studies. On the other hand, the adequate use and interpretation of these data pose unique methodological challenges. These may include issues as diverse as dealing with assessment instruments in different languages and their applicability across different cultures, establishing measurement invariance between these different assessment conditions, handling missing data, figuring out how to reduce the long computation times that are needed for complex analyses, or figuring out how to use process data in computer-based assessments.The complex structure and size of international LSA databases often cause researchers to hesitate. For example, the 2012 cycle of the PISA assessment included data from 510,000 children from 65 economies. International LSA data therefore differ in many ways from more traditional data sets. For instance, LSA data (including international surveys) are usually not sampled at random, and students are typically not given every available test item (Martin, Mullis, & Kennedy, 2007; OECD, 2009). Moreover, there are particularly challenging organizational differences that must be handled adequately as they influence the data; examples are different modalities and language barriers (Butler & Stevens, 2001). For data analysis, specific approaches that are not part of many university curricula are required. As a consequence, the complex data structure of LSAs is often neglected by researchers, thus leading to the application of inadequate analyses and methods (Rutkowski, Gonzalez, Joncas, & von Davier, 2010).Within this special issue, we aim to provide an overview of the vast array of methodological challenges that come with LSA data as well as the current state of the art in tackling them. Considering the diversity of challenges, we convinced various experts on different aspects of LSA to contribute their latest work to this special issue. However, because there were so many contributions, we needed to split the special issue into two parts in order to cover the whole range of methodological challenges posed by LSA.The first part of the special issue consists of four papers highlighting the diversity of challenges while simultaneously introducing potential solutions. authors, all established experts in the field of LSA, demonstrate exciting new ways of handling the transition to computer-based testing, maintaining maximum measurement precision, and dealing with missing data.In the first paper, titled The transition to computer-based in large-scale assessments: Investigating (partial) measurement invariance between modes, Sarah Burger, Ulf Krohne, and Frank Goldhammer illustrate how investigating (partial) measurement invariance between modes can facilitate the transition to computer-based in LSAs. authors present a multiple-group IRT model approach for analyzing mode effects on the test and item levels. In addition, they review instances where partial measurement invariance is sufficient for combining item parameters into one metric. Finally, they present an extension of the modeling approach to explain mode effects by means of item properties.The second paper is titled Differentiated assessment of mathematical competence with multidimensional adaptive testing and was contributed by Anna Mikolajetz and Andreas Frey. …
IntroductionBullying is generally defined as an intentional aggressive act characterized by repetition of actions and asymmetric power relationships (Olweus, 1999). Three decades of research on bullying around the world, including research on ijime in Japan, considered the most similar concept to bullying in the West and confirmed the extensiveness and diversity of the problem (Smith, Morita, Junger-Tas, Olweus, Catalano, & Slee, 1999). Many studies identified the serious negative consequences of being victimized, of bullying others, and of being a bystander, not only for individuals but also for the climate of classes, year groups and schools in general (Boulton, Trueman, & Murray, 2008; Obermann, 2011; Rivers, Poteat, Noret, & Ashurst, 2009; Sweeting, Young, West, & Der, 2006; Ttofi, Farrington, & Losel, 2011). Whether a child becomes a stable victim may depend on the child's ability to use internal resources to respond to the victimization. Sometimes external assistance is available, though victims often tolerate the mistreatment because of the fear of bullying getting worse or of not having enough support from others (Kanetsuna & Smith, 2002). Recently, attentions of researchers and practitioners have been directed towards bystanders for bullying prevention and intervention, as bullying mostly takes place in the playground, classrooms, or corridors, where other children are likely to be present (Whitney & Smith, 1993; Wiens & Dempsey, 2009). However, Pergolizzi, Richmond, Macario, Gan, Richmond, and Macario (2009) revealed high level of apathy among bystanders, claiming that half of their participating students did nothing when they witnessed others being bullied, and 40 % of them considered the bullying as none of their business.Indices to evaluate antibullying interventionsIn light of these issues, a wide variety of antibullying intervention projects has been developed and implemented worldwide, and a number of meta-analyses have been carried out on the effectiveness of such projects. The outcomes of some earlier metaanalyses suggested that overall effects were minimal. For example, Smith, Schneider, Smith, and Ananiadou (2004) reviewed 14 antibullying intervention studies implemented in 11 different countries, and re-evaluated the intervention effects of each study by using the change on outcome measures between pretest and posttest. They found that the effects of intervention projects fell almost exclusively into the categories of small, negligible, and negative for both victimization and bullying outcomes. Only one condition in one study was categorized as having a medium effect, and none was categorized as large. More recently, Ttofi and Farrington (2011) reviewed 53 different school-based intervention projects and meta-analyzed 44 of these, and revealed more positive outcomes. They found that, on average, the projects reduced bullying by around 20-23 % and victimization by around 17-20 %. They also reported that some individual projects, such as KiVa in Finland, have yielded reductions of around 40-50 %, at least in some age groups.These evaluation studies and meta-analyses are certainly an important source of information for developing future successful bullying prevention and intervention programs. However, it has also been noted that we should not rely too much on a single source of data for outcome measures (Smith et al., 2004), and should consider how to interpret the evaluation data very carefully (Toda, Strohmeier, & Spiel, 2008). Most of these evaluations expect reductions of the number of reported bullying and victimization incidents within the school as a whole. Although the goals of such prevention and interventions are to reduce as much bullying as possible, results can be statistically significant even when the size of the reduction is a few percentage points. Others would regard a project as successful if there is a 50 % reduction, as was reported in the Bergen project in Norway or KiVa in Finland, for instance (Olweus, 1999; Salmivalli, Karna, & Poskiparta, 2011). …
IntroductionIn psychological assessment, the speed-power-issue has existed almost since the beginnings of psychological testing. Most intelligence tests have a time-limit, which is often only due to organizational reasons (to make test administration possible for a group of testees, for instance). Apart from this, some intelligence and achievement tests involve by scoring. For example, the commonly used Wechsler tests (e.g. Wechsler Adult Intelligence Scale - Fourth Edition, WAIS IV; Wechsler, 2008) include subtests that credit quick solutions with bonus points. The desirable advantage of such a scoring procedure is an attainment of information about a testeeu0027s ability. The scores of the testees are more differentiated and thus, measurement would take place in a more precise way. Of course, the advantage of such information is only valid if the assumptions underlying the scoring procedure are correct. Particularly, using bonus points assumes that and speed are confounded but not separated traits. This means the ability to solve an item and speededness of a testee in finding the solution are assumed to be a manifestation of the same latent trait and reflect only gradual differences in the measured trait. This assumption is to be scrutinized, as there are empirical results that show and power are actually separated (Carroll, 1993; Partchev, De Boeck, u0026 Steyer, 2011). Partchev, De Boeck, and Steyer (2011) even remark that the current focus should be on avoiding a mixture between and power. However, in practice, scoring procedures combining and power do still exist, which makes the application of methods which test the validness of such a scoring procedure important.Nowadays, it is easy to record response times for each item and there are increasing attempts to use these response times not only in psychological but also in educational assessment. Large-scale tests that have been applied for years, are currently going through a transition from paper-based administration to a computer-based one. In the first part of the special topic Current Methodological Issues in Educational Large-Scale Assessments by Stadler, Greiff, and Krolak-Schwerdt (2016) in this journal Burger, Krohne, and Goldhammer (2016) give a short overview which of the broad-based international large-scale assessments have already changed their administration mode or are planning to change it in the near future. This transition to a computer-based administration provides the opportunity to record not only the response of the examinee but many other variables including the item-specific response time. Response times provide an additional information about how the testee did work on the test. So far, there are many attempts to use response times to increase the measurement accuracy and to minimize measurement errors in psychological and educational assessment. For example, a variety of studies dealt with the detection of guessing in multiple-choice items by using the response times (DeMars, 2007, 2010; Kong, Wise, u0026 Bhola, 2007; Schnipke, u0026 Scrams, 2007, Wise, Pastor, u0026 Kong, 2009). Weeks, von Davier, and Yamamoto (2016) are using response times to distinguish between missing responses which were skipped and those the testee had tried on which but did not give a response.Another option to use the information of response times in large-scale assessment is to incorporate the response times into scoring; that is analogous to intelligence tests which use some credit points for quick solutions. This approach is thought of as a means of increasing measurement accuracy which is of need especially in large-scale assessments where the number of items is limited due to organizational restrictions.A variety of approaches were introduced for incorporating response times in assessments. Van der Linden (2011) gives a fine overview of actual IRT methods modeling response times. He distinguishes between models that include the distributions of response times without any reference to the quality of the item response, and models that integrate item responses and response times (e. …