According to PISA 2012, mathematics self-beliefs involve three dimensions: mathematics self-efficacy, mathematics self-concept and mathematics anxiety. This study utilized the multiple indicators, multiple causes (MIMIC) method, an approach for detecting whether there is a lack of measurement invariance or differential item functioning (DIF), to deal with multiple covariates, multiple dimensions, and ordered categorical variables with threshold structures to study DIF among mathematics self-beliefs items across Shanghai-China ( N = 5177) and the USA ( N = 4978) from Program for International Student Assessment (PISA) 2012. The MIMIC approach with mediators was also applied in the study, which helped to detect variables that could account for meaningful partial or complete DIF effects. Confirmatory factor analysis (CFA), a statistic method that can verify the number of latent traits of the dataset, indicated that the three-factor structure worked for the data. Both robust weighted least square (WLSMV) estimator and robust maximum likelihood (MLR) estimator were used in the parameter estimation to identify items with DIF and quantify the DIF effect size. It was found in the study with MIMIC method that in-school mathematics class periods and out-of-school study hours in PISA 2012 had partial effects on most items with meaningful DIF effects.
Objective: The necessity for pre-injury baseline computerized neurocognitive assessments versus comparing post-concussion outcomes to manufacturer-provided normative data is unclear. Manufacturer-provided norms may not be equivalent to institution-specific norms, which poses risks for misclassifying the presence of impairment when comparing individual post-concussion performance to manufacturer-provided norms. The objective of this cohort study was to compare institutionally derived normative data to manufacturer-provided normative values provided by ImPACT (R) Applications, Incorporated. Method: National Collegiate Athletic Association Division 1 university student athletes (n = 952; aged 19.2 +/- 1.4 years, 42.5% female) from one university participated in this study by completing pre-injury baseline Immediate Post-Concussion Assessment and Cognitive Test (ImPACT) assessments. Participants were separated into 4 groups based on ImPACT's age and gender norms: males <= 18 years old (n = 186), females <= 18 years old (n = 165), males >= 19 years old (n = 361) or females >= 19 years old (n = 240). Comparisons were made between manufacturer-provided norms and institutionally derived normative data for each of ImPACT's clinical composite scores: Verbal (VEM) and Visual (VIM) Memory, Visual Motor Speed (VMS), and Reaction Time (RT). Outcome scores were compared for all groups using a Chi-squared goodness of fit analysis. Results: Institutionally derived normative data indicated above average performance for VEM, VIM, and VMS, and slightly below average performance for RT compared to the manufacturer-provided data (chi(2) >= 20.867; p < 0.001). Conclusions Differences between manufacturer- and institution-based normative value distributions were observed. This has implications for an increased risk of misclassifying impairment following a concussion in lieu of comparison to baseline assessment and therefore supports the need to utilize baseline testing when feasible, or otherwise compare to institutionally derived norms rather than manufacturer-provided norms.
We used multidimensional item response theory to test the internal structure of the Phonological Awareness Literacy Screening in Spanish for Preschool and test for item parameter drift. The measure is aligned with the simple view of reading, which defines reading as consisting of two equally important dimensions: decoding and language comprehension. It involves 134 items grouped into nine different tasks. We administered a pilot version of the measure to 677 students in 2014 and a final version to 968 students in 2015. We did not find evidence that code-related skills and oral language are distinct but correlated dimensions among preschoolers. Rather, a single general dimension of early literacy with task-related specific dimensions fit best. A benefit of our modeling approach is that we can examine the internal structure while also evaluating the difficulty and discrimination of individual items. Results suggest a stable ordering of items within tasks and that some Spanish letters are learned more easily than others.
Threat assessment has been widely endorsed as a school safety practice, but there is little research on its implementation. In 2013, Virginia became the first state to mandate student threat assessment in its public schools. The purpose of this study was to examine the statewide implementation of threat assessment and to identify how threat assessment teams distinguish serious from nonserious threats. The sample consisted of 1,865 threat assessment cases reported by 785 elementary, middle, and high schools. Students ranged from pre-K to Grade 12, including 74.4% male, 34.6% receiving special education services, 51.2% White, 30.2% Black, 6.8% Hispanic, and 2.7% Asian. Survey data were collected from school-based teams to measure student demographics, threat characteristics, and assessment results. Logistic regression indicated that threat assessment teams were more likely to identify a threat as serious if it was made by a student above the elementary grades (odds ratio 0.57; 95% lower and upper bound 0.42-0.78), a student receiving special education services (1.27; 1.00-1.60), involved battery (1.61; 1.20-2.15), homicide (1.40; 1.07-1.82), or weapon possession (4.41; 2.80-6.96), or targeted an administrator (3.55; 1.73-7.30). Student race and gender were not significantly associated with a serious threat determination. The odds ratio that a student would attempt to carry out a threat classified as serious was 12.48 (5.15-30.22). These results provide new information on the nature and prevalence of threats in schools using threat assessment that can guide further work to develop this emerging school safety practice. (PsycINFO Database Record
The United States has the highest incarceration rate in the world, and as a result, one of the largest populations of incarcerated parents. Growing evidence suggests that the incarceration of a parent may be associated with a number of risk factors in adolescence, including school drop out. Taking a developmental ecological approach, this study used multilevel modeling to examine the association of parental incarceration on truancy, academic achievement, and lifetime educational attainment using the National Longitudinal Survey of Adolescent Health (48.3 % female; 46 % minority status). Individual characteristics, such as school and family connectedness, and school characteristics, such as school size and mental health services, were examined to determine whether they significantly reduced the risk associated with parental incarceration. Our results revealed small but significant risks associated with parental incarceration for all outcomes, above and beyond individual and school level characteristics. Family and school connectedness were identified as potential compensatory factors, regardless of parental incarceration history, for academic achievement and truancy. School connectedness did not reduce the risk associated with parental incarceration when examining highest level of education. This study describes the school related risks associated with parental incarceration, while revealing potential areas for school-based prevention and intervention for adolescents.
We developed a criterion-referenced student rating of instruction (SRI) to facilitate formative assessment of teaching. It involves four dimensions of teaching quality that are grounded in current instructional design principles: Organization and structure, Assessment and feedback, Personal interactions, and Academic rigor. Using item response theory and Wright mapping methods, we describe teaching characteristics at various points along the latent continuum for each scale. These maps enable criterion-referenced score interpretation by making an explicit connection between test performance and the theoretical framework. We explain the way our Wright maps can be used to enhance an instructor’s ability to interpret scores and identify ways to refine teaching. Although our work is aimed at improving score interpretation, a criterion-referenced test is not immune to factors that may bias test scores. The literature on SRIs is filled with research on factors unrelated to teaching that may bias scores. Therefore, we also used multilevel models to evaluate the extent to which student and course characteristic may affect scores and compromise score interpretation. Results indicated that student anger and the interaction between student gender and instructor gender are significant effects that account for a small amount of variance in SRI scores. All things considered, our criterion-referenced approach to SRIs is a viable way to describe teaching quality and help instructors refine pedagogy and facilitate course development.
ASD is one of the most heritable neuropsychiatric disorders, though comprehensive genetic liability remains elusive. To facilitate genetic research, researchers employ the concept of the broad autism phenotype (BAP), a milder presentation of traits in undiagnosed relatives. Research suggests that the BAP Questionnaire (BAPQ) demonstrates psychometric properties superior to other self-report measures. To examine evidence regarding validity of the BAPQ, the current study used confirmatory factor analysis to test the assumption of model invariance across genders. Results of the current study upheld model invariance at each level of parameter constraint; however, model fit indices suggested limited goodness-of-fit between the proposed model and the sample. Exploratory analyses investigated alternate factor structure models but ultimately supported the proposed three-factor structure model.
BACKGROUND:School climate is well recognized as an important influence on student behavior and adjustment to school, but there is a need for theory-guided measures that make use of teacher perspectives. Authoritative school climate theory hypothesizes that a positive school climate is characterized by high levels of disciplinary structure and student support.METHODS:A teacher version of the Authoritative School Climate Survey (ASCS) was administered to a statewide sample of 9099 7th- and 8th-grade teachers from 366 schools. The study used exploratory and multilevel confirmatory factor analyses (MCFA) that accounted for the nested data structure and allowed for the modeling of the factor structures at 2 levels.RESULTS:Multilevel confirmatory factor analyses conducted on both an exploratory (N = 4422) and a confirmatory sample (N = 4677) showed good support for the factor structures investigated. Factor correlations at 2 levels indicated that schools with greater levels of disciplinary structure and student support had higher student engagement, less teasing and bullying, and lower student aggression toward teachers.CONCLUSIONS:The teacher version of the ASCS can be used to assess 2 key domains of school climate and associated measures of student engagement and aggression toward peers and teachers.
Preface. Acknowledgements. Chapter 1: Data Management. Chapter 2: Item Scoring. Chapter 3: Test Scaling. Chapter 4: Item Analysis. Chapter 5: Reliability. Chapter 6: Differential Item Functioning. Chapter 7: Rasch Measurement. Chapter 8: Polytomous Rasch Models. Chapter 9: Plotting Item and Test Characteristics. Chapter 10: IRT Scale Linking and Score Equating. References. Appendix
Researchers often use generalizability theory to estimate relative error variance and reliability in teaching observation measures. They also use it to plan future studies and design the best possible measurement procedures. However, designing the best possible measurement procedure comes at a cost, and researchers must stay within their budget when designing a study. In this study, we applied the LaGrange multiplier method to obtain facet sample size equations that minimize relative error variance (hence maximize reliability) under budget constraints. We did this for a crossed design and three nested designs that are more typical of data collection in teaching observation studies. Using an example budget and variance components similar to those found in practice, we demonstrate the use of these equations. We also show the way variance components from fully crossed designs can be combined to use our equations for a nested design.
Observational methods are increasingly being used in classrooms to evaluate the quality of teaching. Operational procedures for observing teachers are somewhat arbitrary in existing measures and vary across different instruments. To study the effect of different observation procedures on score reliability and validity, we conducted an experimental study that manipulated the length of observation and order of presentation of 40-minute videotaped lessons from secondary grade classrooms. Results indicate that two 20-minute observation segments presented in random order produce the most desirable effect on score reliability and validity. This suggests that 20-minute occasions may be sufficient time for a rater to observe true characteristics of teaching quality assessed by the measure used in the study, and randomizing the order in which segments were rated may reduce construct irrelevant variance arising from carry over effects and rater drift.
The Authoritative School Climate Survey was designed to provide schools with a brief assessment of 2 key characteristics of school climate-disciplinary structure and student support-that are hypothesized to influence 2 important school climate outcomes-student engagement and prevalence of teasing and bullying in school. The factor structure of these 4 constructs was examined with exploratory and confirmatory factor analyses in a statewide sample of 39,364 students (Grades 7 and 8) attending 423 schools. Notably, the analyses used a multilevel structural approach to model the nesting of students in schools for purposes of evaluating factor structure, demonstrating convergent and concurrent validity and gauging the structural invariance of concurrent validity coefficients across gender. These findings provide schools with a core group of school climate measures guided by authoritative discipline theory.
The purpose of this study is to introduce a measure of standards-based mathematics teaching practices, the Mathematics Scan (M-Scan), and to examine its validity and score reliability. First, we define standards-based mathematics teaching practices based on eight dimensions that have emerged in recent conceptualizations by researchers and in the context of existing observational measures. Second, we present three sources of validity evidence: content review by experts, analysis of response processes of coders, and convergent and discriminant patterns with existing observational measures. Third, we provide evidence of inter-coder (or inter-rater) reliability through analyses of variance components and calculation of reliability coefficients, using the framework of generalizability theory. Results show the M-Scan holds promise as a useful tool in mathematics education research, measuring indicators of standards-based teaching practices unique to the subject of mathematics.
Universal Design for Learning (UDL) is a framework that is commonly used for guiding the construction and delivery of instruction intended to support all students. In this study, we used a related model to guide creation of a multimedia-based instructional tool called content acquisition podcasts (CAPs). CAPs delivered vocabulary instruction during two concurrent social studies units to 32 SWD and 109 students without disabilities. We created CAPs using a combination of evidence-based practices for vocabulary instruction, UDL, and Mayer's instructional design principles. High school students with and without learning disabilities completed weekly curriculum-based measurement (CBM) probes (vocabulary matching) over an 8-week period along with two corresponding posttests. Students were nested within sections of world history and randomly assigned to alternating treatments (CAPs and business as usual) that were administered sequentially to each group. Results revealed that students with and without disabilities made significant growth on CBMs and scored significantly higher on the posttests when taught using CAPs.
The authors generated exact probability distributions for sample sizes up to 35 in each of three groups (n ≤ 105) and up to 10 in each of four groups (n ≤ 40). They compared the exact distributions to the chi-square, gamma, and beta approximations. The beta approximation was best in terms of the root mean square error. At specific significance levels, either the gamma or beta approximation was best. These results suggest that the most common approximation, the chi-square approximation, is not a good choice. The authors demonstrate this point using an applied example. Critical value tables for the exact distribution are available online at http://faculty.virginia.edu/kruskal-wallis. The portion of these tables that provides critical values for equal sample sizes appears in this article. The authors recommend that researchers use critical values from the exact distribution whenever possible. If sample sizes exceed those included in the authors' exact probability tables, they recommend using the beta approximation instead of the chi-square and gamma approximations.
jMetrik and WINSTEPS are two Rasch measurement software applications that implement joint maximum likelihood estimation of Rasch, partial credit, and rating scale model parameters via a proportional curve fitting algorithm. We describe this algorithm in this paper and explain the handling of missing data and extreme cases. Results from a simulation study that manipulated sample size and the number of test items indicate that both programs produce similar bias and root mean squared error values. In addition, root mean squared difference values indicate that estimates from each program are within 0.001 and 0.004 logits of each other depending on the model in question.
To examine the current state of reading attitudes among middle school students in the United States, a survey was developed and administered to 4,491 students in 23 states plus the District of Columbia. The instrument comprised four subscales measuring attitudes toward: recreational reading in print settings, recreational reading in digital settings, academic reading in print settings, and academic reading in digital settings. Factor analysis confirmed the factor structure corresponding to the four subscales, and reliability coefficients for these subscales ranged from 0.78 to 0.86. Correlations among the subscales varied considerably, due largely to the recreational digital subscale. Analyses of variance subsequently confirmed a pattern for the recreational digital subscale that differed from that of the others. For academic digital, recreational print, and academic print, the attitudes of females were more positive than those of males; however, for attitudes toward recreational reading in digital settings, the pattern was reversed. In addition, results for three of the subscales showed a gradual worsening of attitudes from 6th to 8th grade. The exception was academic print, for which attitudes did not differ by grade. No interactions were observed between grade and gender for any of the subscales. Results are discussed in the context of attitude theory and the rapid evolution of digital literacy and its social uses by adolescents.