The purpose of this article is to explore crossing differential item functioning (DIF) in a test drawn from a national examination of mathematics for 11-year-old pupils in England. An empirical dataset was analyzed to explore DIF by gender in a mathematics assessment. A two-step process involving the logistic regression (LR) procedure for detecting uniformand nonuniform DIF was applied to identify crossing DIF. The results showed 36 uniform and 19 nonuniform statistically significant gender DIF items. Out of the 19 nonuniform DIF items, 10 items were crossing DIF. We explained nonuniform DIF using the crossing point in item characteristic curves and the LR-DIF coding scheme. This study showed that crossing DIF exists in empirical data and the findings from this study provide a potentially valuable contribution in understanding such items.
Researchers interested in exploring substantive group differences are increasingly attending to bundles of items (or testlets): the aim is to understand how gender differences, for instance, are explained by differential performances on different types or bundles of items, hence differential bundle functioning (DBF). Some previous work has modelled hierarchies in data in this context or considered item responses within persons, but here we model the bundles themselves as explanatory variables at the item level potentially explaining significant intra-class correlation due to gender differences in item difficulty, and thus explaining variation at the second item level. In this study, we analyse DBF using single-and two-level models (the latter modelling random item effects, which models responses at Level 1 and items at Level 2) in a high-stakes National Mathematics test. The models show comparable regression coefficients but the statistical significances of the two-level models are smaller due to the larger values of the estimated standard errors. We discuss the contrasting relevance of this effect for test developers and gender researchers.
The aims of this study are (a) to examine the sources of differential functioning by gender via differential bundle functioning (DBF) in mathematics assessment and (b) to use DBF to explore whether the differential functioning displayed is construct-relevant or construct-irrelevant. Three qualitatively different areas, namely curriculum domains, test modality, and problem presentation, were investigated using the Roussos and Stout (1996) multidimensionality-based differential item functioning analysis paradigm with an adaptation of the Logistic Regression approach to model DBF. An empirical dataset from a mathematics national examination of 11-year-old pupils in England was analyzed. We argue that DBF found in curriculum domains may be construct-relevant of the measure, whereas problem presentation and test modality include arguably construct-irrelevant and so may indicate gender bias.
When examinees with the same ?ability? take a test, they should have an equal chance of responding correctly to an item irrespective of group membership. This logic in assessment is known as measurement invariance. The lack of invariance of the item-, bundle-, and test-difficulty across different subgroups indicates differential functioning (DF). The aim of this study is to advance our understanding of DF by detecting, predicting and explaining the sources of DF by gender in a mathematics test. The presence of DF means that the test scores of these examinees may fail to provide a valid measure of their performance. A framework for investigating DF was proposed, moving from the item-level to a more complex random-item level, which provides a theme of critiques of limitations in DF methods and explorations of some advances. A dataset of 11-year-olds of a high-stakes National mathematics examination from England was used in this study. The results are reported in three journal publication format papers. The first paper addressed the issue of understanding nonuniform differential item functioning (DIF) at the item- level. The nonuniform DIF is investigated because it is a possible threat when common DIF statistics sensitive to uniform DIF may indicate no significant DIF. This study differentiates two different types of nonuniform DIF, namely crossing and noncrossing DIF. Two commonly used DIF detection methods, namely the Logistic Regression (LR) procedure and the Rasch measurement model were used to identify crossing and noncrossing DIF. This paper concludes that items with nonuniform DIF do exist in empirical data; hence there is a need to include statistics sensitive to crossing DIF in item analysis. The second paper investigated the sources of DF via differential bundle functioning (DBF) because this way we may get a substantive explanations of DF - without which we do not know if DF is ?valid? or ?biased?. Roussos and Stout?s (1996a) multidimensionality-based DIF paradigm was used with an extension of the LR procedure to detect DBF. Three qualitatively different content areas: test modality, curriculum domains and problem presentation were studied. This paper concludes that DBF in curriculum domains may elicit construct-relevant variance, and so may indicate 'real' differences, whereas problem presentation and test modality arguably includes construct-irrelevant variance and so may indicate gender bias. Finally, the third paper considered item-person responses as hierarchically nested within items. Hence a two-level logistic model was used to model the random item effects, because otherwise it is argued that DF might be over-exaggerated and may lead to invalid inferences. This paper aimed to explain DF via DBF comparing single-level and two-level models. The DIF effects of the single-level model were found to be attenuated in the two-level model. A discussion of why the two different models produced different results was presented. Taken together, this thesis shows how validity arguments regarding bias should not be reduced to DF at item-level but can be analysed on three different levels.