In this paper, we review the notion of aberrant responses in educational and psychological assessments. We examine the literature on aberrant responding, from the early work of Lee Cronbach, to contemporary interpretations of aberrant response processes in digital assessments. We contrast approaches that portray aberrant responses as a de-facto threat to test integrity and validity (i.e. something to be detected and removed), with an alternative paradigm, informed by the work of Cronbach, which interprets anomalous and unexpected responses as a source of formative insight, with potential to improve understanding of test content and design, to illuminate the diverse ways in which tested populations engage with tests and questionnaires. We argue that the latter approach provides stronger foundations for psychometrics and for AI-based models of item responding.
Abstract: There is no consensus among assessment researchers about many of the central problems of response process data, including what is it and what is it comprised of. The Standards for Educational and Psychological Testing ( American Educational Research Association et al., 2014 ) locate process data within their five sources of validity evidence. However, we rarely see a conceptualization of response processes; rather, the focus is on the techniques and methods of assembling response process indices or statistical models. The method often overrides clear definitions, and, as a field, we may therefore conflate method and methodology – much like we have conflated validity and validation ( Zumbo, 2007 ). In this paper, we aim to clear the conceptual ground to explore the scope of a holistic framework for the validation of process and product. We review prominent conceptualizations of response processes and their sources and explore some fundamental questions: Should we make a theoretical and practical distinction between response processes and response data? To what extent do the uses of process data reflect the principles of deliberate, educational, and psychological measurement? To answer these questions, we consider the case of item response times and the potential for variation associated with disability and neurodiversity.
There is no consensus among assessment researchers about many of the central problems of response process data, including what is it and what is it comprised of. The Standards for Educational and Psychological Testing (American Educational Research Association et al., 2014) locate process data within their five sources of validity evidence. However, we rarely see a conceptualization of response processes; rather, the focus is on the techniques and methods of assembling response process indices or statistical models. The method often overrides clear definitions, and, as a field, we may therefore conflate method and methodology - much like we have conflated validity and validation (Zumbo, 2007). In this paper, we aim to clear the conceptual ground to explore the scope of a holistic framework for the validation of process and product. We review prominent conceptualizations of response processes and their sources and explore some fundamental questions: Should we make a theoretical and practical distinction between response processes and response data? To what extent do the uses of process data reflect the principles of deliberate, educational, and psychological measurement? To answer these questions, we consider the case of item response times and the potential for variation associated with disability and neurodiversity.
Digital-first high stakes assessments invite us to rethink the standards, user experience, and vocabulary of assessment validity. We explore the disruptive potential of digital-first high stakes assessments to set higher standards in test performance and validity. In this paper we describe the new lexicon that digital-first assessments have introduced into high stakes digital assessments such as personalisation, user experience, and accessibility. The features of digital high stakes assessments have the potential to reduce sources of construct-irrelevant variation and improve test validity. In conclusion we argue that the distinctive features of digital high stakes assessment challenge our understanding of good assessment design setting new standards for assessment performance and validity.
Drawing on Kane's argument-based approach to validity and Toulmin's later work on cosmopolitanism and diversity, this paper asks whose validity arguments and evidence are being presented in International Large-Scale Assessments (ILSAs), where and when. With a case study of the OECD's PISA for Development, we demonstrate that validity arguments are assembled, negotiated and transformed by the network of actors. We claim that the challenge of ILSAs is not to establish a single authoritative argument through the displacement of plural interpretations and uses. Instead, the tasks of an argument-based approach should be to create a democratic space in which legitimately diverse arguments and intentions can be recognized, considered, assembled and displayed.
We propose a novel method called Quantitative Multimodal Interaction Analysis to understand the meaning of interactions from a set of multimodal observable behavior. We apply this method for the measurement of collaborative problem-solving skills in a dyadic online game specially designed for this purpose. We outline our assumptions and describe the machine learning approach that help us tag multimodal behaviors connecting the theoretical construct with the empirical evidence.
Purpose This paper aims to investigate small-scale, qualitative observations of interviewer–respondent interaction in the Organisation for Economic Co-operation and Development’s Programme for the International Assessment of Adult Competencies (PIAAC). Design/methodology/approach The paper uses video-ethnographic methods to document talk and gesture in assessment in Slovenian household settings. It presents an in-depth case study of interaction in a single testing situation. Findings Observing interaction in assessment captures data on assessment performance that is not available in quantitative analysis of assessment response processes. The character of interviewer–respondent interaction and rapport is shaped by the cognitive demands of assessment and the distinctive ecological setting of the household. Research limitations/implications Observational data on assessment response processes and interaction in real-life assessments can be integrated into and synthesized with other sources of “process data”. Practical implications Assessment programs such as PIAAC should consider the significance of the household setting on assessment quality and observations of interaction in assessment as a valid source of paradata. Social implications There is a place for small-scale observational studies of assessment to inform public understanding of assessment quality and validity. Originality/value The paper provides qualitative insights into the significance of interaction and “interviewer effects” in household assessment settings.
Often excluded and overlooked at the national level, poor and marginalized communities within low- and middle-income countries also frequently slip through international efforts to raise global education outcomes. When the focus is on average country-level performance, those who face the most barriers to education and learning continue to be left out. This book shifts the conversation to bring more attention to learning inequalities within countries. It is rooted in discussions on learning that first took place during an international conference held by the University of Pennsylvania, in the United States, on 2-3 March 2017. The premise is that by focusing more on marginalized communities within poorer countries equity can be improved and overall national levels of learning will increase. Quality education and learning play central roles in the 2030 UN Sustainable Development Goals (SDGs). However, the development goals are largely normative, meaning that too little attention has been given to variations within countries and to those who are performing at the lower end of the social pyramid. Dan Wagner, UNESCO Chair and senior organizer of the conference, has called for greater attention to the 'learning equity' agenda as a way of reinforcing the potential and impact of the UN goals. Featuring essays and commentaries from 36 international experts, the book is a cutting-edge resource for those seeking to understand the science of learning in low-resource settings worldwide. It explores the complex nuances of the phrase bottom of the pyramid, looks at how to measure learning in marginalized populations, and ways that new policy approaches can improve learning for all. Greater knowledge on learning in low-income societies is crucial if we are to achieve both inclusion and equity in improving the quality of education worldwide.
This paper reports on a pilot study that used eye tracking techniques to make detailed observations of item response processes in the OECD Programme for the International Assessment of Adult Competencies (PIAAC). The lab-based study also recorded physiological responses using measures of pupil diameter and electrodermal activity. The study tested 14 adult respondents as they individually completed the PIAAC computer-based assessment. The eye tracking observations help to fill an ‘explanatory gap’ by providing data on variation in item response processes that are not captured by other sources of process data such as think aloud protocols or computer-generated log files. The data on fixations and saccades provided detailed information on test item response strategies, enabling profiling of respondent engagement and response processes associated with successful performance. Much of that activity does not include the use of the keyboard and mouse, and involves ‘off-screen’ use of pen and paper (and calculator) that are not captured by assessment log-files. In conclusion, this paper points toward an important application of eye tracking in large-scale assessments. This includes insights into response processes in new domains such as adaptive problem-solving that aim to identify individuals’ ability to select and combine resources from the digital and physical environment.
Observations of real-life testing situations can provide important insights into test validation and assessment response processes. We consider observations of face-to-face interaction as a starting point to investigate how variation in assessment performance is informed by the ecological characteristics of the testing situation. This perspective offers a contrast to conventional 'think aloud' protocols. Observations of real-life testing situations capture in-vivo perspectives on assessment interaction and performance as it occurs, rather than as if occurring in-vitro. Our aim is to deliberately capture and consider the "noisy" ecological characteristics in the testing situation (e.g., interaction, setting, affect) that may influence response processes, but that are not part of the trait under investigation. Observations of real-life testing situations reveal observable structures and patterns of behaviour, but every performance is somewhat different. Like jazz, the testing situation involves elements of improvisation. This is illustrated using transcripts from the OECD Survey of Adult Skills (PIAAC).
This article discusses talk and gesture as neglected sources of process data (Maddox, 2015, Maddox and Zumbo, 2017). The significance of the article is the growing use of various sources of process data in computer-based testing (Ercikan and Pellegrino, (Eds.) 2017; Zumbo and Hubley, (Eds.) 2017). The use of process data on talk and gesture expands the sources of information about test performance and validity (Kane and Mislevy, 2017).
This article explores the potential for ethnographic observations to inform the analysis of test item performance. In 2010, a standardized, large-scale adult literacy assessment took place in Mongolia as part of the United Nations Educational, Scientific and Cultural Organization Literacy Assessment and Monitoring Programme (LAMP). In a novel form of interdisciplinary collaboration, an ethnographer worked closely with psychometric researchers to investigate the sources and explanations of differential item functioning and test item performance. The research involved detailed ethnographic observations of literacy assessment events. The results illustrate the potential for ethnography to provide insights into testing situations, group and test item performance, and to combine with psychometric analysis in cross-cultural assessment.
Informed by Goffman's influential essay on 'The neglected situation' this paper examines the contextual and interactive dimensions of performance in large-scale educational assessments. The paper applies Goffman's participation framework and associated theory in linguistic anthropology to examine how testing situations are framed and enacted as social occasions. It considers assessment as a shared focus of social activity, located in time and space, involving an assemblage of artefacts and actors. The paper presents ethnographic examples of adult literacy and numeracy assessment in Mongolia. The first part provides ethnographic description of a testing situation. The second part looks in detail at how linguistic interaction influences assessment performance.
What happens when standardised literacy assessments travel globally? The paper presents an ethnographic account of adult literacy assessment events in rural Mongolia. It examines the dynamics of literacy assessment in terms of the movement and re-contextualisation of test items as they travel globally and are received locally by Mongolian respondents. The analysis of literacy assessment events is informed by Goodwin's 'participation framework' on language as embodied and situated interactive phenomena and by Actor Network Theory. Actor Network Theory (ANT) is applied to examine literacy assessment events as processes of translation shaped by an 'assemblage' of human and non-human actors (including the assessment texts).
The concepts of literacy events and practices have received considerable attention in educational research and policy. In comparison, the question of value, that is, 'which literacy practices do people most value?' has been neglected. With the current trend of cross-cultural adult literacy assessment, it is increasingly important to recognise locally valued literacy practices. In this paper we argue that measuring preferences and weighting of literacy practices provides an empirical and democratic basis for decisions in literacy assessment and curriculum development and could inform rapid educational adaptation to changes in the literacy environment. The paper examines the methodological basis for investigating literacy values and its potential to inform cross-cultural literacy assessments. The argument is illustrated with primary data from Mozambique. The correlation between individual values and respondents' socio-economic and demographic characteristics is explored.