In this study, upper-elementary-age students used an interactive reading app to read from a classic children's novel during a summer program. Students took turns reading with an adult virtual narrator (audiobook). We use process and background data to explore factors that could predict whether a reader will read their next turn or skip it. We find that skipping quickly becomes self-perpetuating, underscoring the need to support the teacher in providing just-in-time personalized intervention to help students avoid the disengagement trap.
At a time when institutions of higher education are exploring alternatives to traditional admissionstesting, institutions are also seeking to better support students and prepare them for academicsuccess. Under such an engaged model, one may seek to measure not just the accumulatedknowledge and skills that students would bring to a new academic program, but also their ability togrow and learn through the academic program. To help prepare students for law school before theymatriculate, the JD-Next is a fully-online, non-credit, 7-10 week course to train potential juris doctor(JD) students in case reading and analysis skills. This study builds upon the work presented forprevious JD-Next cohorts by introducing new scoring and reliability estimation methodologies basedon a recent redesign of the assessment for the 2021 cohort, as well as presenting updated validity andfairness findings, using first-year grades, rather than merely first-semester grades as in prior cohorts.Results support the claim that the JD-Next exam is reliable and valid for predicting law schoolsuccess, providing a statistically significant increase in predictive power over baseline modelsincluding entrance exam scores and grade-point average. In terms of fairness across racial and ethnicgroups, smaller score disparities are found with JD-Next than with traditional admissionsassessments, and the assessment is shown to be equally predictive for students from underrepresentedminority groups and first-generation students. These findings, in conjunction with those fromprevious research, support the use of the JD-Next exam for both preparation and admissions offuture law school students.
In vertical scaling, results of tests from several different grade levels are placed on a common scale. Most vertical scaling methodologies rely heavily on the assumption that the construct being measured is unidimensional. In many testing situations, however, such an assumption could be problematic. For instance, the construct measured at one grade level may differ from that measured in another grade (e.g., construct shift). On the other hand, dimensions that involve low-level skills are usually mastered by almost all students as they progress to higher grades. These types of changes in the multidimensional structure, within and across grades, create challenges for developing a vertical scale. In this article, we propose the use of projective IRT (PIRT) as a potential solution to the problem. Assuming that a test measures a primary dimension of substantive interest as well as some peripheral dimensions, the idea underlying PIRT is to integrate out the secondary dimensions such that the model provides both item parameters and ability estimates for the primary dimension. A simulation study was conducted to evaluate the effectiveness of the PIRT as a method for vertical scaling. An example using empirical data from a measure of foundational reading skills is also presented.
The construct of reading comprehension has changed significantly in the twenty-first century; however, some test designs have not evolved sufficiently to capture these changes. Specifically, the nature of literacy sources and skills required has changed (wrought primarily by widespread use of digital technologies). Modern theories of comprehension and discourse processes have been developed to accommodate these changes, and the learning sciences have followed suit. These influences have significant implications for how we think about the development of comprehension proficiency across grades. In this paper, we describe a theoretically driven, developmentally sensitive assessment system based on a scenario-based assessment paradigm, and present evidence for its feasibility and psychometric soundness.
Traditional measures of reading ability designed for younger students typically focus on componential skills (e.g., decoding, vocabulary) and the items are often presented in a discrete and decontextualized format. The current study was designed to explore whether it was feasible to develop a more integrated, scenario-based assessment of comprehension for younger students. A secondary goal was to examine developmental differences in item performance when administration was in listening versus reading modalities. Cross-sectional differences were examined across kindergarten to third grade on a scenario-based assessment comprised of literal comprehension, inference, vocabulary, and background knowledge items. The assessment, originally targeted for third grade, was administered one-on-one to 141 third grade and 485 second grade students. It was adapted for and administered to kindergarten (n = 390) and first grade (n = 419) students by reducing the number of items and switching to a listening comprehension method of administration. Each grade was significantly more accurate than the previous grade on overall performance and background knowledge. A regression analysis showed significant variance associated with background knowledge in predicting comprehension, even after controlling for grade. A deeper analysis of item performance across grades was conducted to examine what elements worked well and where improvements should be made in adapting comprehension assessments for use with young children. ASSESSING COMPREHENSION FROM K TO 3 Assessing Comprehension in Kindergarten through Third Grade Assessing young children’s comprehension development can be challenging. Since the passing of the No Child Left Behind Act of 2001 (2002) school level reading comprehension assessments have been administered beginning in third grade. However, in earlier grades, a componential approach is most common. Following the Simple View of Reading (Gough & Tunmer, 1986; Hoover & Gough, 1990), a traditional starting point is to divide measures between those involving word-level reading recognition versus linguistic comprehension. Although the divided assessment of word reading and language comprehension may provide some understanding regarding young children’s reading and language skills, this assessment practice does not necessarily provide insights regarding how children integrate reading and language in order to access deeper meaning. The purpose of the exploratory research presented in this paper was to begin to develop and explore an early comprehension assessment that moves beyond this componential method and yet still readily evaluates students who have limited word reading and language abilities. One way to accommodate for early or persisting word reading limitations and individual language differences is to assess reading comprehension using a componential approach. There is practicality and efficiency in adopting a componential approach prior to third grade. Component measures may limit the complexity of the task environment, potentially reducing working memory and cognitive load because there are fewer task demands, in comparison to a more integrated assessment that measures multiple skills simultaneously. Historically, this componential approach was used and considered a sensible and functional way to assess the changing development of children between kindergarten and 3rd grade. This development includes the progression of children from a) limited alphabet ASSESSING COMPREHENSION FROM K TO 4 knowledge, to understanding of the alphabetic principle; b) from basic decoding of words in their listening lexicon, to acquiring a sizeable sight-word vocabulary; c) from word by word reading, to fluent oral (and silent) reading of continuous texts. Assessments that provide an indication of a child’s ability to assemble the aforementioned component skills into an integrated whole have the potential to provide valuable insights into the skills necessary in literacy activities beyond third grade. One of the aims of the larger project in which this study is embedded was to develop innovative assessments of reading for understanding, and create a new type of computer-based assessment, termed Scenario-Based Assessment (SBA). The use of SBA techniques allowed us to deliver a set of thematically related source materials in a digital environment and potentially enable us to assess reading comprehension and language processes in a more integrated way (Bennett, 2010; 2011). In order for the SBA results to be interpretable, we examined task difficulty relative to child development. While most kindergarten and first grade children would not yet have the word recognition skills to read, we sought to determine whether younger children would have the language comprehension abilities that were targeted by the SBA form. In this study we explored the performance levels associated with the texts and questions that were read to the students, as well as the impact of changing modality (listening vs. reading). Additionally, we examined the developmental differences in children’s background knowledge, memory, and reasoning skills. In short, our goal was to understand—at least in part—how to design SBA comprehension tests that target early developmental reading comprehension abilities in children, and to better understand individual differences as children learn to integrate their language with their reading skills. ASSESSING COMPREHENSION FROM K TO 5 Theoretical Background An evolving construct of reading in the 21st century While a simplified construct of reading (or listening) comprehension may be justifiable as one type of measure for young children, we agree with the position that the construct of reading comprehension has been changing significantly in the past several decades and that comprehension assessment designs have not kept pace with the changes in how people read in the 21st century (e.g., digital literacy, Coiro, 2009; Leu et al., 2013), nor advances in cognitive science and instruction (Gordon Commission, 2013). This position is aligned with various assessment reforms such as the Common Core State Standards (National Governors Association Center for Best Practices & Council of Chief State School Officers, 2010), the Partnership for 21st Century Skills (2008), and other seminal works (Bennett, 2010; 2011; Bransford, Brown, & Cocking, 2000; Pellegrino, Chudowsky, & Glaser, 2001). These sources support the argument that the typical approach to measuring comprehension, one that focuses on students’ understanding of a single text in isolation, underrepresents the complexity of a modern construct of reading comprehension that emphasizes purpose-driven, multiple document processing (Britt & Rouet, 2012). This is not to say that this form of comprehension test is not valid or should not be used, but rather an acknowledgment that there is more to comprehension than what is covered in traditional, print-based tests of reading. If the construct of comprehension is evolving, ideally, these changes should be reflected in developmentally appropriate content and tasks administered to young, as well as older, children. From a review and synthesis of these and other literatures, we have been developing a framework for assessing reading for understanding across prekindergarten through twelfth grade ASSESSING COMPREHENSION FROM K TO 6 (O’Reilly & Sabatini, 2013; Sabatini, O’Reilly, 2013; Sabatini, O’Reilly & Deane, 2013). In these publications, we provide a definition of reading, outline the constructs underlying the measures and how they might change across the school years, explain the use of scenario-based assessment, and the role of performance moderators. While the details of the framework are beyond the scope of this paper, we briefly summarize some key points. From our survey of the literature, we hold that reading is a purposeful activity (van den Broek, Linderholm, & Gustafson, 2001), that purposes are used to set standards of coherence (Linderholm, Virtue, Tzeng, & van den Broek, 2004) for determining what is relevant when reading text sources (McCrudden, Magliano & Schraw, 2011). In practical and everyday reading contexts, students must be able to integrate and evaluate multiple sources (Britt & Rouet, 2012) to satisfy their purpose for reading. This process draws upon students’ background knowledge (Shapiro, 2004), as they may be required to interpret texts from different points of view or through the lenses of different disciplines (Goldman, 2012; LaRusso et al., 2016). Skilled readers may also use reading strategies (McNamara, 2012), metacognition, and self-regulation (Hacker, Dunlosky, & Graesser, 2009) to help process text deeply. Encouraging the use of reading strategies during an assessment is one way of modeling good comprehension practices and cognitive habits (e.g., Griffin, Malone, & Kammenui, 1995; Ozuru, Best,O’Reilly, & McNamara, 2007; Meyer & Ray, 2011), as well as an effective means of collecting evidence of reading proficiency. While this portrayal of purposeful, integrative comprehension is often reserved for describing what it means to be college and career ready, one must establish the precursors of these skills in younger children to ensure a trajectory of learning that leads to proficiency by the end of secondary schooling (Goldman, 2004). ASSESSING COMPREHENSION FROM K TO 7 To address coverage of this expanded construct, we identified five knowledge and skill targets that span all developmental levels: print, verbal, discourse, conceptual, and social. Print targets address the skills needed to “get the printed words off the page” including decoding, word recognition, and all other typographical conventions of written language. Verbal targets address broader language resources such as vocabulary, morphology, syntax, and grammar, with a focus on word to sentence level processes. Moving beyond the sentences and word level, discour
In this research report, we describe the conceptual foundation and measurement properties of the Reading Inventory and Scholastic Evaluation (RISE). The RISE is a 6‐subtest, Web‐administered reading skills components battery. We review the theoretical and empirical foundations of each subtest in the battery, as well as item designs. The results included in this report feature a calibrated item pool based on a national sample of students, an extension of the vertical scale to span Grades 3–12, psychometric analyses of the data for each subtest, an item response theory scaling study for each of the subtests across the entire grade span, an evaluation of multidimensionality, an evaluation of differential item functioning for gender and race/ethnicity, and an expanded review of validity evidence.
The validity of studies investigating interventions to enhance fluid intelligence (Gf) depends on the adequacy of the Gf measures administered. Such studies have yielded mixed results, with a suggestion that Gf measurement issues may be partly responsible. The purpose of this study was to develop a Gf test battery comprising tests meeting the following criteria: (a) strong construct validity evidence, based on prior research; (b) reliable and sensitive to change; (c) varying in item types and content; (d) producing parallel tests, so that pretest-posttest comparisons could be made; (e) appropriate time limits; (f) unidimensional, to facilitate interpretation; and (g) appropriate in difficulty for a high-ability population, to detect change. A battery comprising letter, number, and figure series and figural matrix item types was developed and evaluated in three large-N studies (N = 3,067, 2,511, and 801, respectively). Items were generated algorithmically on the basis of proven item models from the literature, to achieve high reliability at the targeted difficulty levels. An item response theory approach was used to calibrate the items in the first two studies and to establish conditional reliability targets for the tests and the battery. On the basis of those calibrations, fixed parallel forms were assembled for the third study, using linear programming methods. Analyses showed that the tests and test battery achieved the proposed criteria. We suggest that the battery as constructed is a promising tool for measuring the effectiveness of cognitive enhancement interventions, and that its algorithmic item construction enables tailoring the battery to different difficulty targets, for even wider applications.
We report results of 2 studies examining the relation between decoding and reading comprehension. Based on our analysis of prominent reading theories such as the Simple View of Reading (Gough & Tunmer, 1986), the Lexical Quality Hypothesis (Perfetti & Hart, 2002) and the Self-Teaching Hypothesis (Share, 1995), we propose the Decoding Threshold Hypothesis, which posits that the relation between decoding and reading comprehension can only be reliably observed above a certain decoding threshold. In Study 1, the Decoding Threshold Hypothesis was tested in a sample of over 10,000 Grade 5–10 students. Using quantile regression, classification analysis (Receiver Operating Characteristics) and broken-line regression, we found a reliable decoding threshold value below that there was no relation between decoding and reading comprehension, and above which the two measures showed a positive linear relation. Study 2 is a longitudinal analysis of over 30,000 students’ reading comprehension growth as a function of their initial decoding status. Results showed that scoring below the decoding threshold was associated with stagnant growth in reading comprehension. We argue that the Decoding Threshold Hypothesis has the potential to explain differences in the prominent reading theories in terms of the role of decoding in reading comprehension in students at Grade 5 and above. Furthermore, the identification of decoding threshold also has implications for reading practice. (PsycINFO Database Record (c) 2019 APA, all rights reserved)
ABSTRACT Indicators of student academic growth are desired in state accountability systems in order to approximate student learning over time and attribute observed growth to schooling inputs. Through an extant analysis of five states’ assessment data, this study offers evidence about whether longitudinal match rates and measures of growth differ at the state level for students with disabilities, relative to students without disabilities. There were three main findings: 1) In states in which a modified assessment was offered, students with disabilities were more likely to have missing prior year scores, and consequently missing growth scores; 2) Low scoring students, many of whom had a disability, were more likely to have missing prior scores on the state general assessment, and consequently missing growth scores; 3) Students with and without disabilities showed similar growth using transition and gain score definitions of growth, but students with disabilities had lower growth when estimated via a regression-based model. Measurement and policy considerations are discussed.
Vertical scales are widely used in educational assessment as a basis for considering grade-to-grade changes in student performance. Typically, the underlying construct is assumed to be essentially unidimensional; however, if there is a change in the measured construct across grades, this assumption may be untenable. Developing a multidimensional vertical scale in these instances provides a potential solution to this problem. This paper uses empirical data from four parallel forms of a test designed to measure six foundational reading skills-administered to students in grades 6-9-to address issues in the development of a multidimensional vertical scale. The defensibility of the multidimensional structure, value-added subscores, and the stability of the scale are considered. Student growth based on unidimensional versus multidimensional estimates of ability is also presented with particular attention to implications associated with potential construct shift.
Achievement estimates are often based on either number correct scores or IRT-based ability parameters. Van der Linden (2007) and other researchers (e.g., Fox, Klein Entink, & van der Linden, 2007; Ranger, 2013) have developed psychometric models that allow for joint estimation of speed and item parameters using both response times and response data. This paper presents an application of this type of approach to a battery of 4 types of fluid reasoning measures, administered to a large sample of a highly educated examinees. We investigate the extent to which incorporation of response times in ability estimates can be used to inform the potential development of shorter test forms. In addition to exploratory analyses and response time data visualizations, we specifically consider the increase in precision of ability estimates given the addition of response time data relative to use of item responses alone. Our findings indicate that there may be instances where test forms can be substantially shortened without any reduction in score reliability, when response time information is incorporated into the item response model. (PsycINFO Database Record
(ProQuest: ... denotes formulae omitted.)In many assessments there is a high likelihood that some examinees will at least one answer for one reason or another. This type of may or may not be ability related. While low ability with respect to the measured construct may play a role, other reasons, such as low motivation, lack of attention, or running out of time may be likely possibilities. If the data are ignorable (i.e., at random or completely at random), estimates of item parameters and examinee ability in a latent variable model will be unbiased, but if they are not ignorable, the treatment of these values can introduce systematic error into parameter estimates (Rubin, 1976). When analyzing responses from test administrations in which data are more than rarely occurring exceptions, some principled way of treating these data is required. This is true in operational analyses using either classical test theory (which typically requires complete data without missingness) or modern test theory such as item response theory (IRT; Lord u0026 Novick, 1968) which, in principle, can handle data that are completely at random or at random. For this paper we primarily address data treatments in the context of IRT or related methods. The goal of this study is to examine whether the coding of omitted responses based on response time information from a computer-based assessment in a low-stakes context can improve results compared to ad hoc methods (e.g., treating omitted responses as incorrect by default) typically applied in estimates of item/ability parameters. This goal is accomplished using empirical data from the Programme for the International Assessment of Adult Competencies (PIAAC) literacy and numeracy cognitive tests.BackgroundTerminologyBefore proceeding, it is important to clearly define the different types of seen in large scale assessment data. We use the term nonresponse to refer to any value in a dataset of item responses that, after scoring, does not correspond to a correct or incorrect response code (or by extension for polytomous items, responses that do not correspond to a score category that influences an examineeu0027s estimate of ability). In more simple terms, if an individual does not provide an answer to a given item, it is considered a nonresponse. If an examinee has no opportunity to respond to the item, either by design or because the individual did not see the item, we refer to these as not administered2 and not reached items respectively as missing responses. On the other hand, we use the term omit to refer to values in cases where the examinee saw the item (or is believed to have seen the item) but no response was given. The reason for this distinction is that and omitted responses are treated differently for the purpose of item response modeling and/or scoring. Not reached items, not administered items, and omitted item responses all warrant a different treatment: An individual who never saw an item by design cannot be expected to respond, obviously. Similarly, an examinee who did not reach the last 2-3 items because of time constraints also had no chance to produce a response and may or may not have gotten the items correct. On the other hand, an individual who saw an item and decided not to provide a response may have done so due to an understanding that the item is too difficult, or due to other reasons such as a lack of motivation, or an intent to come back to this item later that was never acted upon.Treatment of dataTypically, data are treated in one of two ways for the purpose of item analysis and scoring: 1) the values are coded as not administered and excluded from the estimation of item and/or ability parameters or 2) the values are coded as omits and scored as incorrect or partially correct. The former approach is generally applied for responses that appear sequentially, usually at the end of a test or test section. …
Traditional measures of reading ability designed for younger students typically focus on componential skills (e.g., decoding, vocabulary), and the items are often presented in a discrete and decontextualized format. The current study was designed to explore whether it was feasible to develop a more integrated, scenario-based assessment of comprehension for younger students. A secondary goal was to examine developmental differences in item performance when administration was in listening versus reading modalities. Cross-sectional differences were examined across kindergarten to third grade on a scenario-based assessment comprising literal comprehension, inference, vocabulary, and background knowledge items. The assessment, originally targeted for third grade, was administered one-on-one to 141 third-grade and 485 second-grade students. It was adapted for and administered to kindergarten (n = 390) and first-grade (n = 419) students by reducing the number of items and switching to a listening comprehension method of administration. Each grade was significantly more accurate than the previous grade on overall performance and background knowledge. A regression analysis showed significant variance associated with background knowledge in predicting comprehension, even after controlling for grade. A deeper analysis of item performance across grades was conducted to examine what elements worked well and where improvements should be made in adapting comprehension assessments for use with young children.
The nonequivalent groups with anchor test (NEAT) design is frequently used in test score equating or linking. One important assumption of the NEAT design is that the anchor test is a miniversion of the 2 tests to be equated/linked. When the content of the 2 tests is different, it is not possible for the anchor test to be adequately representative of both tests. Lin and Dorans conducted a simulation study in 2010 to investigate the effect of content representativeness of the anchor test on linking via different linking methods when the 2 tests are nonparallel in content structure in the unique case where the groups are equivalent. The current study extends the Lin and Dorans study to the case with nonequivalent group data. Specifically, the current study investigates the impact of content representativeness and length of anchor test on linking when the 2 tests are multidimensional and nonparallel in content structure. The NEAT design was employed. The linking results from 3 classic linear equating methods—Levine observed score, Tucker equating, and chained linear—were examined. The results from the study indicated that equating the tests with different structure should be avoided. For equatings with anchor test, additional bias is likely to be introduced by using an inadequate anchor test.
This technical report describes the conceptual foundation and measurement properties of the Reading Inventory and Scholastic Evaluation (RISE). The RISE is a 6‐subtest, Web‐administered reading skills components battery. The theoretical and empirical foundations of each subtest in the battery are reviewed, as well as item designs. The results included in this report feature a vertical extension of the RISE to span Grades 5–10, psychometric analysis of parallel forms of each subtest, results of item response theory (IRT) scaling studies for each of the subtests across the entire grade span, and evaluation of differential item functioning (DIF) for gender and race/ethnicity.
When designing a reading intervention, researchers and educators face a number of challenges related to the focus, intensity, and duration of the intervention. In this paper, we argue there is another fundamental challenge-the nature of the reading outcome measures used to evaluate the intervention. Many interventions fail to demonstrate significant improvements on standardized measures of reading comprehension. Although there are a number of reasons to explain this phenomenon, an important one to consider is misalignment between the nature of the outcome assessment and the targets of the intervention. In this study, we present data on three theoretically driven summative reading assessments that were developed in consultation with a research and evaluation team conducting an intervention study. The reading intervention, Reading Apprenticeship, involved instructing teachers to use disciplinary strategies in three domains: literature, history, and science. Factor analyses and other psychometric analyses on data from over 12,000 high school students revealed the assessments had adequate reliability, moderate correlations with state reading test scores and measures of background knowledge, a large general reading factor, and some preliminary evidence for separate, smaller factors specific to each form. In this paper, we describe the empirical work that motivated the assessments, the aims of the intervention, and the process used to develop the new assessments. Implications for intervention and assessment are discussed.
Using longitudinal data for an entire state from 2004 to 2008, this article describes the results from an empirical investigation of the persistence of value-added school effects on student achievement in reading and math. It shows that when schools are the principal units of analysis rather than teachers, the persistence of estimated school effects across grades can only be reasonably identified by placing strong constraints on the variable persistence model implemented by Lockwood, McCaffrey, Mariano, and Setodji. In general, there are relatively strong correlations between the school effects estimated using these constrained models and a reference model that assumes full persistence. These correlations vary somewhat by grade and the underlying test subject. The results from this study indicate cautious support for previous findings that the assumption of full persistence for cumulative value-added effects may be untenable, and evidence is also presented, which indicates a strong interaction by test subject. However, the practical impact of violating the assumption of full persistence appears to be smaller in the context of schools than it is for teachers.
The R package plink has been developed to facilitate the linking of mixed-format tests for multiple groups under a common item design using unidimensional and multidimensional IRT-based methods. This paper presents the capabilities of the package in the context of the unidimensional methods. The package supports nine unidimensional item response models (the Rasch model, 1PL, 2PL, 3PL, graded response model, partial credit and generalized partial credit model, nominal response model, and multiple-choice model) and four separate calibration linking methods (mean/sigma, mean/mean, Haebara, and Stocking-Lord). It also includes functions for importing item and/or ability parameters from common IRT software, conducting IRT true-score and observed-score equating, and plotting item response curves and parameter comparison plots.