
This paper describes a comparative analysis of (ADL) and (IADL) items administered to two samples, 4,430 persons representative of older Americans, and 605 persons representative of patients with rheumatoid arthrisit (RA). Responses are scored separately using both Likert and Rasch measurement models. While Likert scoring seems to provide information similar to Rasch, the descriptive statistics are often contrary if not contradictory, and estimates of reliability from Likert are inflated. The test characteristic curves derived from Rasch are similar despite differences between the levels of disability with the two samples. Correlations of Rasch item calibrations across three samples were .71, .76, and .80. The fit between the items and the samples, indicating the compatibility between the test and subjects, is seen much more clearly with Rasch with more than half of the general population measuring the extremes. Since research on disability depends on measures with known properties, the superiority of Rasch over Likert is evident.
Considerable uncertainty exists over the benefit that patients receive from surgical decompressive treatment for cervical spondylotic myelopathy (CSM). Such difficulties might be addressed by accurate quantification of CSM severity as part of a trial determining the outcome of surgery in different patient groups. This study compares the applicability of various existing quantitative severity scales to measurement of CSM severity and the effects on severity of surgical decompression. Scores on the following scales were determined on 100 patients with CSM preoperatively and then again six months following surgical decompression: Odom's Criteria, Nurick grade, Ranawat grade, Myelopathy Disability Index (MDI), Japanese Orthopaedic Association (JOA) Score, European Myelopathy Score (EMS) and Short Form-36 Health Survey (SF36). All the scales showed significant improvement following surgery. However, each had differing qualities of reliability, validity and responsiveness that made them more or less suitable. The MDI showed the greatest sensitivity between different severity levels, sensitivity to operative change and reliability. However, analysis of all the questionnaire scales into components that looked at different aspects of function revealed potential problems with redundancy and a lack of consistency. This prospective observational study provides a rational basis for determining the advantages and disadvantages of different existing scales in measurement of CSM severity and for making adaptations to develop a scale more specifically suited to a comprehensive surgical trial.
Constructed-response or open-ended tasks are increasingly used in recent years. Since these tasks cannot be machine-scored, variability among raters cannot be completely eliminated and their effects, when they are not modeled, can cast doubts on the reliability of the results. Besides rater effects, the estimation of student ability can also be impacted by differentially weighted tasks/items that formulate composite scores. This simulation study compares student ability estimates with their true abilities under different rater scoring designs and differentially weighted composite scores. Results indicate that the spiraled rater scoring design without modeling rater effects works as well as the nested design in which rater tendencies are modeled. As expected, differentially weighted composite scores have a confounding effect on student ability estimates. This is particularly true when open-ended tasks are weighted much more than the multiple-choice items and when rater effects interact with weighted composite scores.
This paper describes a comparative analysis of (ADL) and (IADL) items administered to two samples, 4,430 persons representative of older Americans, and 605 persons representative of patients with rheumatoid arthrisit (RA). Responses are scored separately using both Likert and Rasch measurement models. While Likert scoring seems to provide information similar to Rasch, the descriptive statistics are often contrary if not contradictory, and estimates of reliability from Likert are inflated. The test characteristic curves derived from Rasch are similar despite differences between the levels of disability with the two samples. Correlations of Rasch item calibrations across three samples were .71, .76, and .80. The fit between the items and the samples, indicating the compatibility between the test and subjects, is seen much more clearly with Rasch with more than half of the general population measuring the extremes. Since research on disability depends on measures with known properties, the superiority of Rasch over Likert is evident.
This article raises and tries to answer questions concerning what objectivity in psychosocial measurement is, why it is important, and how it can be achieved. Following in the tradition of the Socratic art of maiuetics, objectivity is characterized by the separation of meaning from the geometric, metaphoric, or numeric figure carrying it, allowing an ideal and abstract entity to take on a life of its own. Examples of objective entities start from anything teachable and learnable, but for the purposes of measurement, the meter, gram, volt, and liter are paradigmatic because of their generalizability across observers, instruments, laboratories, samples, applications, etc. Objectivity is important because it is only through it that distinct conceptual entities are meaningfully distinguished. Seen from another angle, objectivity is important because it defines the conditions of the possibility of shared meaning and community. Full objectivity in psychosocial measurement can be achieved only by attending to both its methodological and its social aspects. The methodological aspect has recently achieved some notice in psychosocial measurement, especially in the form of Rasch's probabilistic conjoint models. Objectivity's social aspect has only recently been noticed by historians of science, and has not yet been systematically incorporated in any psychosocial science. An approach to achieving full objectivity in psychosocial measurement is adapted from the ASTM Standard Practice for Conducting an Interlaboratory Study to Determine the Precision of a Test Method (ASTM Committee E-11 on Statistical Methods, 1992).
We present an approach to constructing an aggregate index of health at the population level with data from Medicare beneficiaries using the 1991 (N = 12,667), 1995 (N = 15,590), and 1997 (N=17,058) Medicare Current Beneficiary Survey (MCBS). Similar to other work with survey data, we develop a weighted health status index from which one can calculate a point in time health status score for any beneficiary. Scores range from 1.0, representing "excellent health and no activity limitation", to 0.0, representing deceased. Sequences of numerically weighted health states experienced over time can be summed to calculate years of healthy life for beneficiaries. We test both the stability of the scoring system when developed on independent samples, as well as the sensitivity of years of healthy life calculations to changes in scoring assumptions. Findings suggest that, in addition to mortality, morbidity appears to play a significant role in the years of healthy life accrued by Medicare beneficiaries since entry into the Medicare program. Further, the index scoring system is highly stable when derived on independent samples. Finally, calculations of years of healthy life are robust to changes in scoring assumptions. The weighted health index for Medicare current beneficiaries (WHIMCBS) is a stable overall index of health and may be a useful ongoing indicator of health within the Medicare population.
The purpose of this study is to evaluate the measurement properties of the Symptom Impact Inventory using both psychometric and Rasch analyses. This inventory is designed for generally healthy midlife women. The sample included 340 midlife women aged 45-65 representing two studies. The first study involved Black and White employed sedentary women (n = 161) who volunteered for a walking intervention. The second study of migration and health included women who were recent immigrants from the former Soviet Union (n = 179). The women reported experiencing an average of 13.44 symptoms (S.D.=7.88) with a range of 1 to 32. Principal components analysis identified 5 components in this sample. Rasch measurement analysis found excellent model fit for the Symptom Impact Inventory with only 2 symptoms, Decreased appetite and Decreased sexual desire or interest, unstable in scale dimensionality analyses. Person and item parameters were reliable, and comparisons with groups known to differ on symptom reporting provided substantial validity. Although the two sample groups differed significantly on most demographic characteristics, a cross-cultural comparison found the scale structure remarkably robust.
Univocal definition and classification of Asthma have always been a matter of discussion, and that is reflected in the difficulty of constructing a measure of pathology severity. The European Community Health Survey is a multinational survey designed to compare the prevalence of asthma in subjects, aged 20 to 44 years, in several European areas. In each participating center a sample of 3000 adults filled a self-administered screening questionnaire composed by 9 dichotomous items. Aim of the present study is to investigate unidimensionality of the ECRHS screening questionnaire and to determine and validate a scoring of asthma-like symptoms seriousness. Dimensionality and scoring was determined through a Homogeneity Analysis by Alternating Least Square; while scoring validation was assessed by a cross validation technique. This study found the existence of a sole dimension underlying the screening questionnaire; furthermore a scoring of asthma-like symptoms seriousness was determined with the indication of a cut-off in order to distinguish between asthma symptomatic and non symptomatic subjects.
This report describes the use of Rasch analysis of paired comparisons in measuring physician work. The method examines the ranking of a series of related medical or surgical procedures by survey of physicians experienced in the provision of such services. Each service in the group is paired with every other service. The physicians select which of a pair of services from the group under study represents the greater amount of physician work. When the results are analyzed by Rasch method, the resulting measures (in logits) provide a scale (yardstick) upon which each service can be placed relative to all the other services in the group. In this way, a rank ordering of the services is accomplished that is both linear and objective. The Rasch measures can than be charted against existing work values to refine the assigned relative values for such services. The result of applying this method to a series of spinal operations indicates that this technique could be expanded to many other groups of related medical services, with improved relative valuation of physician work.
We describe the use of a mathematical/statistical method (i.e., Rasch analysis) to elucidate biological patterns of disability present in the functional ability of persons undergoing medical rehabilitation. Two measures chosen for illustration are the FIM Instrument for inpatients and the Body Movement and Control (BMC) measure for outpatients. In order to meet the assumptions necessary for application of linear statistics to clinical measurement studies, Rasch analysis was used to transform ordinal scales into linear measures. Another unique feature of Rasch analysis is that it allows evaluation of the difficulty of items and the abilities of persons being tested, separately, on the same metric. Also, the difficulty represented by each item may be arranged along a hierarchy from easy to hard. The hierarchies of functional ability items are dependent upon the specific patterns of disability related to underlying pathophysiology. For inpatients, initial analyses of the 18 items of the FIM Instrument demonstrated separate hierarchies for the 13 motor items and for the 5 cognition items. Subsequent analyses demonstrated five distinct patterns for the 13 motor items of: brain dysfunction, orthopedic conditions, pain conditions, ambulatory spinal cord dysfunction, and wheelchair users with spinal cord dysfunction. Two patterns were identified for cognition: stroke with right body hemiparesis and all others. For outpatients, the BMC measure of physical functioning is used to demonstrate that pathophysiologic conditions are expected to affect the hierarchial pattern of items differently. This was noted to be the case for persons with lower body dysfunction, low back pain, and neck pain/upper limb dysfunction. Based upon the item responses, sitting, reaching and standing appear to represent items most useful for discriminating between the three conditions in terms of the functional consequences. Rasch analysis, among other advantages, enables investigation of the subtle relationships among items and is a useful method to evaluate underlying biological patterns of disability. A clinician, using a map that shows the expected relationships between item scores, may observe that a particular patient matches or does not match the expected pattern. Such insights may help the clinician in monitoring the responses of the patient to treatment efforts.
This study tests the stability of health status measurement (SF-36) in a working population. A total of 4,225 employees from two sectors (one state agency, one private company) enrolled in three health plans at Trigon BlueCross/BlueShield of Virginia. An eight-dimension short-form health survey (SF-36) was first tested on a cross-sectional basis for its validity. Then, a panel study was established to test for the stability of health status instrument over time. Structural equation modeling built on equality constraint conditions was the statistical technique for this study. Data were collected through two-wave mail surveys. Both comprehensive (original eight scales) and parsimonious (revised five scales) models of health status were found fit into the data quite well. Furthermore, the revised parsimonious model was shown highly stable over time. Within a working population aged 18 to 64, people are relatively healthy. Their perception of health issues is reflected mainly on "physical health status," as indicated by physical functionings or role limitations. The high stability of revised health status model warrants the possibility of using a more concise health status instrument for the majority of people in working force.
An estimation method is proposed for the Rasch model on the basis of the pseudolikelihood theory of Arnold and Strauss (1988). A simulation study was conducted to compare the proposed maximum pseudolikelihood estimates with the well known conditional maximum likelihood and unconditional maximum likelihood estimates for the item parameters of the Rasch model. The results show great similarity between the methods.
The Center on Rehabilitation Effectiveness (CRE) was created in 1998 at Boston University's Sargent College of Health and Rehabilitation Sciences. An important reason for the creation of the Center was demand by purchasers of health services and patients for high quality yet cost-effective rehabilitation programs, particularly with respect to pediatric services. This demand for accountability has created many pressures and challenges for the psychometric community. These challenges include: pediatric rehabilitation assessments that are conceptually grounded in rehabilitation theory; new instruments that are short yet sensitive enough to detect meaningful disability restrictions and that are sensitive enough to measure meaningful change; and new scales that offer real meaningful comparisons across patients. The purpose of this paper is to explain how the Center for Rehabilitation Effectiveness will meet these growing challenges.
Data collected on rating scales have generally been analyzed without verifying that the scales have functioned as intended. The FIM levels are precisely conceptualized and meticulously defined. Their effective empirical functioning as ordinal categories merits continual monitoring. Ordinality implies that each succeeding level represents a higher level of functioning. Further, as a patient improves in functioning each ordinal level in turn is expected to be observed. Taking advantage of the clarity of Rasch theory, guidelines are suggested that prompt the analyst to investigate whether the rating categories are cooperating to produce observations on which useful measurement and prudent inference about patient status can be based.
The Principal Components subroutine of the BIGSTEPS computer program (Rasch analysis) was used to identify primary and secondary dimensions within an item set composed of variables pertaining to family capacity to provide assistance to a family member with disability. Data were obtained through interviews with family caregivers of patients in a major rehabilitation hospital, both during patients' inpatient rehabilitation stay and 3 and 6 months after inpatient discharge. Rasch analysis revealed the primary dimension within these data to be family capacity to provide unpaid instrumental help (18 items) and the strongest secondary dimension to be stressors on the caregivers as assessed during the inpatient stay (5-11 items). A scale consisting of the caregiver stressor items significantly predicted patients' functional gain at 3 months after discharge from inpatient rehabilitation (r = -.21) and at 6 months (r = -.20), and also the number of days that the patient had spent in a nursing facility as of 6 months after discharge (r = +.17). Caregiver stress, burnout, and quality of life at 3 months and 6 months were also significantly predicted. These findings strongly suggest the importance of more definitive research into family stressors that affect long-term patient outcomes.
Physical functioning is a common construct of interest for patients receiving rehabilitation. This report describes the assessment of hierarchial structure, unidimensionality and reproducibility of item calibrations along the continuum of physical functioning defined by the PF-10 of the MOS SF-36. Three new questions specific to patients with upper extremity impairments were added, and item calibrations were compared across several groups of patients with different musculoskeletal impairments. Reproducibility of item calibrations over testing times was supported. Item order was dependent on impairment in a clinically logical pattern. Construct validity of the physical functioning scale was supported and improved with the new questions for patients with upper extremity impairments as well as for patients with some lower level extremity impairments.
SERVPEFR, the performance component of the Service Quality Scale (SERVQUAL), has been shown to measure five underlying dimensions corresponding to Tangibles, Reliability, Responsiveness, Assurance, and Empathy (Parasuraman, Zeithaml, & Berry, 1988). This paper describes three separate studies employing SERVPERF in an Australian context. In the first of these studies (N = 113), a shortened 15-item version of the SERVPERF scale (SERVPERF-R) was found to be suitable for use in an Australian small business setting. A five-factor structure was identifiable but the factors were highly correlated, suggesting that they were not clearly distinct. The tendency for marked negative skewness observed by other researchers was also noted here. A follow-up study involving three other small businesses (N = 212) used Rasch analysis to test assumptions about the spread of items on the underlying continuum. These analyses indicated that there is an even, though narrow, spread of items across the continuum. The Rasch analysis suggested that the items in both SERVPERF and SERVPERF-R are too easy to rate highly and that more "difficult" items need to be added to the scale. The third study (N = 122) was conducted using a version of SERVPERF-R that included seven new items intended to extend the range of the scale. The new items, however, did not achieve this desirable outcome. The implications for service quality assessment are discussed.
The School Assessment of Motor and Process Skills (School AMPS) is an assessment tool designed to be used by occupational therapists to measure the effectiveness of a student's ability to perform school tasks in naturalistic classroom settings. Rater reliability, internal scale validity, and person response validity of the School AMPS was investigated by examining the goodness-of-fit of raters, motor and process skill items, and students to the many-faceted Rasch model used in the development of the School AMPS. Five of six raters demonstrated acceptable goodness-of-fit (MnSq < or = 1.4 and z < 2). All 36 motor and process skill items demonstrated acceptable goodness-of-fit. Of the 208 students in the study, 93.7% demonstrated acceptable goodness-of-fit on the School AMPS motor scale and 88.9% demonstrated acceptable goodness-of-fit on the School AMPS process scale. The results of this study support the rater reliability, scale validity, and person response validity for the School AMPS as a tool to be used to evaluate the effectiveness of student performance of school tasks in the classroom setting.