
This study has a didactic purpose to help applied investigators and practitioners to understand the roles of observed categorical data (OCD) in structural equation modelling (SEM) and the appropriate ways of analysing such data under SPSS AMOS. To that end, the study reviews types of OCD (nominal, ordinal, dichotomous and polytomous) and their incorporation into SEM under AMOS to play different roles. The study presents two applications from the health and retirement study where Bayesian statistical inference is used to analyse one set of OCD variables serving as endogenous variables with/without groups created by another OCD variable. Besides, the study demonstrates the typical ways of summarising, reporting and interpreting the results from Bayesian statistics, and compares AMOS with several other SEM programmes (Mplus, R lavaan, Stata and SAS PROC CALIS) on handling OCD. The study concludes with summaries of the findings for its intended audience.
In this paper, we explore the possibility of whether the likelihood of observing transformative learning may be predicted using information related to personal history and current and previous professions. We examine empirical data collected from a group of Indian profession changers using a machine learning method: random forest algorithm. Results indicate that the following variables play an important role in prediction: 'overall formality (previous profession)', 'community sanction (previous profession)', 'professional authority (current profession)', 'bridge course', and 'gender'. Additionally, this provides empirical support to the position that profession change may have a transformative effect. A discussion and a list of areas for further research are provided.
This study's purpose was to analyse key factors underlying social marginalisation and academic performance. The 2017 data from the Danish PLM survey (N = 42,703) were analysed which contained responses by students (grades 4-10), parents, and class teachers. Multigroup structural equation modelling was applied to explore anticipated gender differences. Two critical factors were identified that were associated with reduced levels of social marginalisation: 1) the degree of teacher support; 2) the strength of the parental community. Finally, the study indicated that girls, and students at lower grade levels, tend to experience greater social marginalisation.
This study examined the mediating effect of self-efficacy on the relationship between emotional intelligence and academic achievement. A total of 257 postgraduate diploma students from the Bangladesh Institute of Management were conveniently sampled. The Schutte Self-Report Emotional Intelligence Test (SSEIT) advanced by Schutte et al. (1998) was applied for measuring 'emotional intelligence'. In assessing self-efficacy, the study adopted the 'Generalised Self-Efficacy Scale (GSES)' formulated by Schwarzer and Jerusalem (1995). The study revealed that all the variables included in the study, i.e., emotional intelligence, self-efficacy, and academic achievement were significantly correlated to each other. Moreover, the results of the study showed that the link between emotional intelligence and academic achievement was fully mediated through self-efficacy. Based on the findings of the study, academic institutions are recommended to include emotional intelligence and self-efficacy in their curriculum.
The common item equating design requires two forms of a test which have a set of items in common in order to control for differences in examinee ability. The common set is subject to compromise when it is used repeatedly, which most likely becomes a serious threat to test fairness. If cheating occurs on common items, the equating process produces inaccurate results which might vary as a result of common item difficulty. This simulation study was conducted to evaluate the impact of the difficulty level of compromised common items on the equating process. The recovery of scaling coefficients and equated scores was assessed using bias and RMSE under various cheating conditions. The results indicated that cheating on higher-difficulty common items produced the most overestimation in the scaling coefficients; which, in turn, caused the most inflation in equating true scores for all test takers, whether they engage in cheating or not.
This study compared candidates' scores based on the normalised model and the two-parameter item response theory (2PL IRT) model using simulated multi-form exam data. Candidates' calculated scores, rankings, qualification status and score ties from the two models were compared with their true values. The results suggest that the 2PL IRT model outperformed the normalised model when the candidate ability distributions varied across forms. It was found that candidate scores based on the 2PL model were more closely related to the true scores. The qualification status of candidates belonging to the top 10% group were more accurately classified by the 2PL model than the normalised model when group abilities differed.
The CELPIP-G test is used by the Canadian federal government to screen immigration eligibility for the skilled worker class. Differential option functioning is a technique used to detect potential bias in the options of multiple-choice items. The purpose of this paper is to investigate DOF in a CELPIP-G reading test form by way of multinomial logistic regression. The results showed that 13.7% of options were flagged as gender DOF. Nonetheless, 11.2% were negligible or small DOF. In the case of uniform gender DOF, twice as many options were found to function against female immigration applicants than against their male counterparts. Female test-takers were more likely to be disadvantaged when tackling questions that asked them to make direct inferences based on factual but unfamiliar information. In contrast, male test-takers were more likely to be disadvantaged when tackling questions that asked them to develop their own interpretations over different views. Moreover, test questions that required an understanding of more sophisticated ideas in complex language structure and allowing personal interpretation tended to show more marked and non-uniform gender DOF.
Grading of short answers in an examination is a tedious exercise that takes so much of examiners' time. Fatigue could set in leading to errors. Sometimes sentiments come into play. The attendant effect of this is variations in the marks awarded to candidates even when they express the same opinion. In this study, Jaccard, Cosine, Jaro and Dice similarity measures were used to grade the answers provided by candidates in examinations of 647 questions. The similarity measures were tested with the aim of ascertaining the measure that rank closest to the average scores provided by three human examiners with the same examinations' answers and marking guides. Results showed that Jaro similarity measure ranked closest to the mean score of the examiners with a variance absolute error of 0.62% and covaried strongly by 97% with a significant level of 0.001.
There has been increasing discussion about the importance of test use and consequences of tests in language teaching, learning, and assessment. The main concern in test development and test use is establishing the valid score-based interpretations and test scores use (Bachman, 1990). Test use is considered at the center of language assessment. As Bachman (1990, p. 55) suggests, “the single most important consideration in both the development of language tests and the interpretation of their results is the purpose or purposes which the particular tests are intended to serve”. Moreover, he argues that “tests are not developed and used in a value-free psychometric test-tube; they are virtually always intended to serve the needs of an educational system or ABSTRACT
Missing data, which are the occasions where answers to certain questions have not been provided, are a common feature of most datasets. However, minimal research has been performed in relation to the existence of missing responses in low-stakes testing situations from international studies, and whether variations can be found in the non-response patterns from country to country. Therefore, the purpose of this study is to identify the variations that might exist in relation to omitting item responses in the TIMSS 2015 fourth grade data and look into their predictors based on item and student characteristics. The results of the study have found that differences do exist between countries on these issues, even when the countries have similar levels of achievement on TIMSS. Moreover, certain student and item characteristics were also found to be related to the patterns of non-responses, although these predictors varied from country to country.
Nonlinear structural equation models (SEMs), which include interactions among latent predictors, as well as quadratic or higher order terms, have been the focus of research over the last three decades, beginning with Kenny and Judd (1984). The great majority of that work has focused on the case where the indicator variables are continuous in nature. However, in practice many nonlinear SEMs will involve the use of responses to items on scales, which are categorical. The focus of the current simulation study was on comparing several methods for modelling nonlinear SEMs when indicator variables were dichotomous. Results of the study showed that a Bayesian approach, as well as a method based on 2-stage least squares, provided the most accurate parameter estimates, the highest power, and the best control over the Type I error rate for the interaction effect. Implications of these findings for practice are discussed.
This study investigated the use of different binomial logistic models as alternatives to the normal model when analysing non-normal aggregate outcomes that are sums of correlated binary responses. The outcome variables provided in the two illustrative examples were preschoolers' uppercase and lowercase letter naming knowledge with different shapes of non-normal distributions. The binomial, beta-binomial, and mixed binomial models with logit links were examined and compared to each other and to the normal linear model. Results were consistent in both examples. Among the models compared, the beta-binomial and mixed binomial models with overdispersion parameters captured interdependence among correlated binary responses. In addition, the mixed binomial model further explained remaining overdispersion and best fitted the data. Implications including advocating for the use of the binomial models with overdispersion parameters for clustered data were further discussed.
Digitally delivered tasks can provide students opportunities to interact with disciplinary content while their interactions with the tasks can be recorded as time-stamped log events in the system server. Through post-hoc analysis of log data, we can re-enact and discover patterns in students' activities. This study addresses the online Earth science module where students engaged in writing and revising scientific arguments in a structured format. We adopted natural language processing (NLP) techniques to analyse students' responses, which enabled us to provide immediate feedback to students on their responses and revisions. Cluster analyses were conducted on the action sequences in four argumentation tasks embedded in the module. For each task, the cluster analyses identified two clusters of students who showed different revision patterns with allocation of time on different items. In addition, students in those two clusters also differed in their initial item scores and item score changes after revision.
Surveys of students are among the primary data sources for research in higher education. Student surveys suffer from increasing level of non-response. This paper investigated the demographic factors that have contributed to non-responses in Mapworks student survey. The study utilised t-test and logistic regression to analyse the data collected from undergraduate students. The findings showed that high school GPA, gender, campus residency, race/ethnicity were significant predictors of survey non-response, and there were significant differences between respondents and non-respondents. Implications for academic administrators were discussed.