To compare two methods for estimation of Work Limitations Questionnaire scores (WLQ, 8 items) from the Role Physical (RP, 4 items) and Role Emotional scales (RE, 3 items) of the SF-36 Health survey. These measures assess limitations in role performance attributed to health (emotional, physical, or both) and different breadth of impact (work vs. work and other activities). We compared WLQ estimates based on an item response theory crosswalk (Method1) and a regression imputation (Method 2). Such estimates can expand the information from studies using only the SF-36 measure, and can inform future data collection strategies. We used data from two independent cross-sectional panel samples (Sample1, n=1382, 51% female, 72% Caucasian, 49% with preselected chronic conditions, 15% with fair/poor health; Sample2, n=301, 45% female, 90% Caucasian, 47% with preselected chronic conditions, 21% with fair/poor health). Method 1 used previously developed and validated IRT based calibration tables. Method 2 used regression models to develop aggregate imputation weights as described in the literature. We evaluated the agreement of observed and estimated WLQ scale scores from the two methods and their ability to discriminate among known groups of patients. Estimated scores from the two methods were strongly correlated (r=.99). Estimated and observed scale scores had strong correlations (r= .68 for RE and r=.76 for RP). Observed and estimated WLQ from both methods successfully differentiated between levels of self reported general health and between patients with and without chronic conditions. For both methods the estimated WLQ means from SF36RP score were closer (and not statistically different for Method1) to the observed WLQ means than estimated WLQ scores from the SF36 RE scale. Our results suggest that both methods provide useful WLQ estimates for group level analysis. Method 1 appears slightly more accurate than Method 2, but is computationally more complex.
Despite the American Heart Association’s interest in research on the determinants of health-related quality of life (HRQoL) among Acute Coronary Syndrome (ACS) survivors, little is known about trajectories of HRQoL post-ACS. We sought to identify such longitudinal patterns, and their predictors, over the 6 months post-ACS discharge. We used data from the Transitions, Risks, and Actions in Coronary Events – Center for Cardiovascular Outcomes and Education (TRACE-CORE) prospective cohort of patients hospitalized with ACS. HRQoL was measured using the Seattle Angina Questionnaire (SAQ) at the index hospitalization and at 1-, 3 -, and 6-months post-discharge. The quality of life subscore of the SAQ ranges 1-100 with higher scores indicating better HRQoL. We used trajectory analysis to identify subgroups of patients with distinctive 6-month post-discharge HRQoL patterns, and predictors of different trajectories. Participants (N=920) had mean age 63 (SD 11) years, 34% were female, and 83% non-Hispanic white. We identified 3 HRQoL trajectories (FAIR, GOOD, and EXCELLENT HRQoL) consisting of 12.1%, 53.8% and 34.1% of participants, respectively: FAIR (baseline average HRQoL = 38.8, and remaining low over the 6-month follow-up); GOOD (baseline average = 62.6 and increasing modestly over time); and, EXCELLENT (baseline average = 87.0 and remaining high). With FAIR HRQoL as the referent, we found that older age predicted better HRQoL (OR per year for GOOD = 1.09 and for EXCELLENT = 1.18) as did non-Hsipanic white race/ethnicity (OR for GOOD = 3.01; for EXCELLENT = 4.21) and male sex (OR for GOOD=1.17 p=0.56, for EXCELLENT = 2.57). [All p < 0.05, except as noted.] On average, HRQoL was relatively stable over time, and no trajectories with decreasing HRQoL were found. Early HRQoL scores could direct resources to improve HRQoL over the first 6 months post-ACS.
To summarize the magnitude of placebo effects on patient-reported physical and mental health outcomes in well-controlled drug trials. We conducted a systematic review of randomized, double-blind, placebo-controlled drug trials via PubMed and supplemental sources published over a 12-year period (1/1/1995-12/31/2011) with documentation of full Medical Outcomes Study short-form 36 (SF-36) scores by treatment group. Standardized effects (SE=change/general population SD) for SF-36 physical and mental component summary score (PCS, MCS) changes were computed from baseline to endpoint for each treatment arm. SE was evaluated versus Cohen’s dcutoffs (small ≤0.30, medium 0.30 to 0.80, and large SE ≥0.80), as well as currently recommended minimum important difference (MID) thresholds of 0.3 and 0.5 SD units. The search identified 805 publications, 150 of which were well-controlled and sufficiently documented drug trials for 35 heterogeneous medical conditions. Sample sizes varied with IQR: 67-448 (median= 205). The majority of trials, 107 (71.3%) reported small SE for PCS or MCS. Medium placebo effects were reported in 37 trials (25%). Large SE were reported in just 6(4%) trials. In contrast, clinically effective drug treatment arm effects were more often medium (51/103;50%) and large(19/104;18%). Median SE for placebo arms across all trials for PCS was 0.11(IQR: 0.00 to 0.26), for MCS (median=0.06, IQR: -0.05 to 0.21). Mean placebo changes met the 0.3 MID threshold in 43 (29%) trials; 17(12%) met the 0.5 point threshold. Although placebo effects are most often small for patient-reported physical and mental outcomes, a noteworthy proportion of well-controlled trials report placebo effects achieving the 0.3 SD MID threshold. Therefore, we recommend evaluating MID for single group comparisons (of controlled or observational data), with a 0.5 SD cutoff, and for controlled trial findings expressed as net of placebo changes, with a smaller 0.3 SD cutoff.
To evaluate the validity and participants’ acceptance of an online assessment of role function using computer adaptive test (RF-CAT).
Item response theory (IRT) modeling evaluated the metric underlying generic physical function (PF) surveys including SF-36, PROMIS and new PF categorical rating items and tested whether scores could be estimated more efficiently while maintaining forward-backward comparability. Generalized Partial Credit Model (GPCM) estimates of parameters for MOS SF-36 PF-10, PROMIS 6-item PF and new (easy-hard) PF categorical rating items and model fit were tested in a probability sample representing the general US population (N = 625). Analyses included: (a) fit of GPCM for 35-item bank; (b) item utilization in computerized adaptive tests (CAT), (c) % at ceiling and floor; (d) % for whom reliability > 0.90 (reliable range); (e) equivalence of mean norm–based scores (mean=50, SD=10) for all measures across mild, moderate and severe chronically-ill groups, and (f) validity in predicting physical and emotional health general summary measures at a 9-month follow-up. The GPCM fit the data and item parameter estimates agreed very well with those previously reported for MOS and PROMIS PF items. In tests of discriminant validity, group means differed substantially across severity groups (RV = 0.81 to 1.00) and score equivalence across methods within each group was confirmed (all differences < 1 point). RV's for standardized PF scores estimated from new E-H items were equivalent to PF-10 and PROMIS PF estimations. Predictive validity was equivalent and substantial (across methods) for physical and significant, but lower, for emotional outcomes at 9 months, as hypothesized. The most efficient (reduced respondent burden, comparable or improved reliability and validity) measure was a new 6-item PF using E-H items and an improved adaptive survey logic. Findings support the standardization of the metric underlying PF measures and extend choice of methods to include more efficient categorical rating scales that maintain forward-backward score comparability. Improved adaptive survey logic reduces respondent burden and increases the reliable range for estimates of scores for familiar legacy measures. This approach warrants application to other generic health domains and tests of translated items and standardized parameters across countries and languages.
To evaluate bootstrap techniques in comparing the validity of PRO measures in discriminating among CKD patients and responding to longitudinal changes. The Kidney Disease Impact Scale (KDIS), CKD-specific legacy (KDQOL Burden, Symptom, and Effect) and generic health (SF-12) scales were administrated to 453 patients and re-administered to 110 patients after three months. ANOVA-based relative validity (RV) coefficients were used to compare how well each scale discriminated between three clinically-defined groups ordered in terms of severity (Dialysis > Stage 3-5 > Transplant), and how responsive each scale was to changes over time for self-evaluated Better, Same and Worse groups. Bootstrap was used to construct confidence intervals (CIs) to determine whether the differences in RVs were significant in comparisons between each scale and the best legacy measure - KDQOL Burden. Sample size, number of bootstrap iterations, and type of CIs were varied to evaluate their impacts on CI using real and artificial data. The sample size played a substantial role. 300 people for 3 groups were suggested as the minimum number to make meaningful comparisons between RVs using CI. Number of bootstrap replications (100 to 10,000) did not show an obvious effect on bootstrap standard error, although 300 showed improvement over 100 on CI. The bias-corrected and accelerated (BCa) type of CI was preferred for correcting both bias and skewness in bootstrap distribution and for producing narrower CIs. Using 95% CI and 300 sample size, differences in RVs were non-significant in comparisons with KDQOL-Burden (RV=1) for the following scales: SF-12 PCS (RV=.6), PF (RV=.7), RP (RV = .77), KDQOL-Effect (RV=.99), and KDIS (RV=1.13). Bootstrapping appears to be valuable in testing the significance of differences in the relative validity of these PRO measures from a statistical perspective. Samples of 100 per group compared and 300 bootstrap replications are recommended.
Objective. This study examined the effect of abatacept, a costimulation modulator, on the health-related quality of life (HRQOL) of patients with rheumatoid arthritis (RA). Methods. Three hundred thirty-nine patients with RA on a background of methotrexate (MTX), who participated in a multicenter, double-blind, placebo-controlled trial, were randomized to abatacept 2 mg/kg, abatacept 10 mg/kg, or placebo. HRQOL was assessed at pretreatment, and at 3, 6, and 12 months posttreatment using the SF-36 Health Survey (SF-36). Changes in SF-36 scores from baseline to 12 months were compared across treatment and placebo groups to examine HRQOL benefits of abatacept. A link between American College of Rheumatology improvement and changes in SF-36 scores was established to demonstrate the association between HRQOL outcomes and clinical response. Results. After 12 months of treatment, patients randomized to abatacept 10 mg/kg showed significantly better HRQOL outcomes overall versus patients randomized to placebo (MANOVA F = 4.71, p < 0.001) or to abatacept 2 mg/kg (MANOVA F = 1.97, p = 0.05). Differences in SF-36 change scores between abatacept 10 mg/kg and placebo groups reached statistical significance on all 8 domain scales, the 2 summary measures, and the SF-36 utility index (SF-6D). Differences in SF-36 change scores between abatacept 10 mg/kg and abatacept 2 mg/kg reached statistical significance on 5 of the 8 domain scales, the physical summary measure, and the SF-6D. Improvement in HRQOL was highly related to clinical response. Conclusion. Abatacept 10 mg/kg plus MTX demonstrated a stronger HRQOL response than placebo plus MTX. The abatacept 2 mg/kg arm showed a very weak and transient response. (First Release Mar 1, 2006; J Rheumatol 2006; 33:681–9) Key Indexing Terms: HEALTH-RELATED QUALITY OF LIFE RHEUMATOID ARTHRITIS SF-36 HEALTH OUTCOMES HEALTH ASSESSMENT ABATACEPT From the Rheumatology and Rehabilitation Research Unit, University of Leeds, Leeds, UK; QualityMetric Incorporated, Lincoln, Rhode Island, USA; Bristol-Myers Squibb, Princeton, New Jersey, USA; Health Assessment Lab, Boston, Massachusetts, USA; Tufts University School of Medicine, Boston, MA, USA; University of Massachusetts, Worcester, MA, USA; and Department of Medicine, University of Alberta, Edmonton, Alberta, Canada. Supported by a grant from Bristol-Myers Squibb. P. Emery, MA, MD, FRCP, Professor, Rheumatology and Rehabilitation Research Unit, University of Leeds; M. Kosinski, MA, Senior Scientist, QualityMetric Inc.; T. Li, PhD, Associate Director, Immunology, BristolMyers Squibb; M. Martin, PhD, Director, Consulting Services, QualityMetric Inc.; G.R. Williams, ScD, Director, Outcomes Research; J-C. Becker, MD, Clinical Director, Bristol-Myers Squibb; B. Blaisdell, MA, Outcomes Research, QualityMetric Inc.; J.E. Ware Jr, PhD, CEO, Chief Science Officer, Chairman of the Board, QualityMetric Inc., Health Assessment Lab, Tufts University School of Medicine; C. Birbara, MD, Rheumatologist, University of Massachusetts; A.S. Russell, MD, Professor of Rheumatology, Department of Medicine, University of Alberta. Address reprint requests to Prof. P. Emery, Department of Rheumatology and Rehabilitation, University of Leeds, 36 Clarendon Road, Leeds LS2 9N2, UK. E-mail: p.emery@leeds.ac.uk Accepted for publication November 25, 2005. A chronic and progressive disease such as rheumatoid arthritis (RA), characterized by joint pain, stiffness, and joint deformity as well as by varying degrees of physical impairment, fatigue, fever, and reactive depression1, places a tremendous burden on the patient, their families, the healthcare system, and society at large. For patients, early RA (average disease duration < 18 mo) places a substantial burden on their physical functioning and emotional well-being that is comparable to diabetes or congestive heart failure2. As the disease progresses, patients experience increasing functional impairment, which may lead to work disability and lost wages. For the families of patients with RA, the progression of RA is likely to place a significant burden, particularly to the extent that the patient is physically impaired, in pain, emotionally distressed, and unable to work. For the healthcare system that attends to patients with RA, the costs of office visits, medications, surgeries, hospitalizations, occupational therapy, and other social services amount to billions of dollars per year, with direct medical care costs in 2001 for an RA patient estimated at US $9519 per year3. For society at large, the functional impairment that characterizes later stages of RA often results in work disability and lost opportunity to gain from the RA patient’s contribution to the workforce. Estimates place the amount of lost wages due to RA at US $2.5 billion per year4. Personal non-commercial use only. The Journal of Rheumatology Copyright © 2006. All rights reserved. Rheumatology The Journal of on July 10, 2020 Published by www.jrheum.org Downloaded from Pharmacologic interventions that improve physical function so that patients are able to work longer and perform daily activities have the potential to improve patients’ quality of life as well as reduce the burden that RA places on society as a whole. Abatacept, a novel selective T cell costimulation modulator, has been shown to be safe, well tolerated, and able to produce a significant dose-dependent reduction in disease activity for RA patients who experienced inadequate responses to disease modifying antirheumatic drugs (DMARD)5,6. Recent results from a randomized, controlled study showed that abatacept significantly improved the signs and symptoms of disease in patients who had active RA despite ongoing methotrexate (MTX) treatment5. In this study, we examined the effect of abatacept therapy (10 mg/kg) in combination with MTX on a broad range of health-related quality of life (HRQOL) domains in patients who had inadequate response to MTX treatment. MATERIALS AND METHODS Study population. Three hundred thirty-nine patients with RA participated in a multicenter, multinational, double-blind, randomized, placebo-controlled trial comparing the efficacy of MTX plus placebo (n = 119), abatacept 2 mg/kg (n = 105), or abatacept 10 mg/kg (n = 115). Abatacept was administered by intravenous infusions at baseline, every 2 weeks for the first month, and monthly thereafter. To participate, patients were required to meet several criteria: (1) the American Rheumatism Association criteria for RA while meeting functional class I, II, or III according to the revised criteria of the American College of Rheumatology (ACR)7,8; (2) have > 10 swollen, > 12 tender joints, and C-reactive protein level > 1 mg/dl signifying active disease; (3) have been treated with MTX for at least 6 months and on a stable dose for 28 days prior to enrollment; and (4) be washed-out of all DMARD other than MTX for at least 28 days before treatment. Provided that the prescribed dose remained stable for the first 6 months of the study, participants were permitted to continue on low-dose corticosteroids (≤ 10 mg/day) and nonsteroidal antiinflammatory drugs. This study was carried out in accordance with the ethical principles of the Declaration of Helsinki. General health status measures. The Medical Outcomes Study Short Form36 Health Survey (SF-36), a well-validated measure of general health status911, was self-administered at baseline (pretreatment) and 3, 6, and 12 months posttreatment to measure HRQOL. The SF-36 measures 8 health dimensions [physical functioning (PF), role limitations due to physical health (RP), bodily pain (BP), general health perceptions (GH), vitality (VT), social functioning (SF), role limitations due to emotional health (RE), mental health (MH)], which are aggregated to produce physical (PCS) and mental (MCS) summary measures. All SF-36 scales and both summary measures were scored using norm-based methods that standardize the scores to a mean of 50 and a standard deviation of 10 in the general US population, with higher scores indicative of better health12,13. A health utility index (SF-6D) was also derived from 11 items of the SF-3614. The SF-6D is a preference-based measure of health that places the observable states of health and functioning on a preference continuum with a value of 0 for death through to 1 for completely well or optimal health. These values are known as health state utilities, which are primarily used to adjust life-years saved by quality for use in economic evaluations and decision models. Effect of treatment on HRQOL. HRQOL outcomes were evaluated in 2 ways. First, changes in SF-36 scale and summary measure scores and the SF-6D from baseline to 12 months were evaluated and compared between placebo and treatment groups. To account for multiple comparisons, multivariate analysis of variance (MANOVA) was conducted to test for overall differences in change scores across all 8 SF-36 scales between placebo and abatacept groups. MANOVA analyses were followed with independent pairwise t tests of statistically significant differences in mean change scores on each SF-36 scale, summary measure, and the SF-6D between placebo and abatacept groups. For these analyses, the last observation was carried forward for those patients who dropped out early from the study. Since presenting mean changes in scores can mask the underlying variability in HRQOL outcomes, the second way that HRQOL outcomes were evaluated was to determine the percentage of patients in each group whose 12 month score on each SF-36 scale was “better,” “the same,” or “worse” than the baseline score. Given that standards for the minimal clinically important difference for the SF-36 have not been well established, the standard error of measurement (SEM) was calculated for each SF-36 scale and summary measure, which provided a boundary of measurement error that could be used to categorize changes in individual patient scores as better, t
The DYNHA Headache Impact Test (HIT)™ uses computerized adaptive testing (CAT) to provide practical and precise measurement at the individual patient level, building on a large bank of headache impact items. This study documents expansion of the HIT™ item pool (from k=47 to k=75 items); and presents results from an initial pilot test of the improved tool, the HEADACHE-CAT™. The HIT item pool covers domains of pain, role functioning, social functioning, fatigue, cognition, and mental health. Modifications were made to existing items in the pool (e.g., revised item stem and response option text), and existing item content was evaluated for redundancy. Expert panels were convened to identify new items with potential clinical relevance. The expanded item pool was fielded through an Internet-based survey of 1,103 headache sufferers. Items were analyzed using factor analytic methods for categorical data and item response models. Analyses of item thresholds and information functions indicated that the new items covered the same range as previous items; however, total test information and measurement precision were increased, with the most notable improvements occurring at the high and low end of the scale. A field study was then conducted with patients (N=50) in two primary care clinics. Measurement coverage, precision and response burden of a dynamically-administered HEADACHE-CAT versus a static fixed-form version were compared. Patients’ acceptance and experiences with the items and a subsequent report produced by the tool were also assessed. Results from the first site (n=25) indicate that the HEADACHE-CAT is very precise, reduces respondent burden, and is well accepted by patients. The DYNHA Headache Impact Test (HIT)™ uses computerized adaptive testing (CAT) to provide practical and precise measurement at the individual patient level, building on a large bank of headache impact items. This study documents expansion of the HIT™ item pool (from k=47 to k=75 items); and presents results from an initial pilot test of the improved tool, the HEADACHE-CAT™. The HIT item pool covers domains of pain, role functioning, social functioning, fatigue, cognition, and mental health. Modifications were made to existing items in the pool (e.g., revised item stem and response option text), and existing item content was evaluated for redundancy. Expert panels were convened to identify new items with potential clinical relevance. The expanded item pool was fielded through an Internet-based survey of 1,103 headache sufferers. Items were analyzed using factor analytic methods for categorical data and item response models. Analyses of item thresholds and information functions indicated that the new items covered the same range as previous items; however, total test information and measurement precision were increased, with the most notable improvements occurring at the high and low end of the scale. A field study was then conducted with patients (N=50) in two primary care clinics. Measurement coverage, precision and response burden of a dynamically-administered HEADACHE-CAT versus a static fixed-form version were compared. Patients’ acceptance and experiences with the items and a subsequent report produced by the tool were also assessed. Results from the first site (n=25) indicate that the HEADACHE-CAT is very precise, reduces respondent burden, and is well accepted by patients.
PAR10 IMPROVING THE SENSITIVITY OF PHYSICAL FUNCTION MEASURES IN RHEUMATOID ARTHRITIS: USE OF ITEM RESPONSE THEORY IN PATIENTS TREATED WITH ABATACEPT (CTLA4IG) Martin M, Emery P, Kosinski M,Ware J, Li T,Williams R, Maclean R, Bjorner J Quality Metric, Inc, Lincoln, RI, USA; University of Leeds, Leeds, United Kingdom; Bristol-Myers Squibb, Princeton, NJ, USA; Bristol-Myers Squibb, Brussels, Belgium OBJECTIVES: The Health Assessment Questionnaire (HAQ) and the Modified HAQ (MHAQ) are examples of common short-forms of physical function used to measure improvement in the treatment of rheumatoid arthritis (RA) which have been associated with ceiling problems. These problems, inherent to short-form surveys, pose risks of failing to detect a treatment response in clinical trials. Item Response Theory (IRT) methods were used to examine the properties of two physical function measures and construct a combined measure to better detect changes in disease activity and treatment response. METHODS: Data were from a 12-month, double-blind, multi-center study of 339 RA patients on a background of methotrexate randomized to Abatacept at 2mg/kg, at 10mg/kg, or placebo. MHAQ and SF-36 (with its Physical Functioning scale, PF10) were administered at pretreatment and 3, 6, and 12 months post-treatment. IRT methods were used to examine the surveys’ measurement properties and compute new IRT-based physical function scores. Analyses of variance were used to assess sensitivity to changes in disease severity and treatment response. Relative validity coefficients were used to compare the measures. RESULTS: A Rasch IRT model fit the data. IRT-based scores successfully lowered the floor and raised the ceiling of the physical function measured. IRT-based scores were 30% more efficient than MHAQ and 50% more efficient than PF10 in discriminating among ACR groups. In discriminating among treatment groups, IRT-based scores were 25% more efficient than MHAQ and 12% more efficient than PF10 at 6-months; and 16% and 17% more efficient at 12-months based on observed effect sizes. CONCLUSIONS: Using IRT methodology to estimate a combined score for physical functioning lead to greater range of the construct measured. The improved measure, with greater measurement precision and sensitivity to treatment response, further confirmed the beneficial effect of Abatacept on physical function in the treatment of RA.
Diminished work productivity, which is common among patients with headache, has economic, social, and personal costs. Lost workplace productivity attributed to headache is due to both absenteeism and to reduced performance and output while on the job. The concept of reduced effectiveness while working is often denoted as presenteeism. Health-care providers recognize lost workplace productivity as a relevant marker of the functional consequences of headache syndromes.
OBJECTIVES The authors examine the data quality and measurement performance of the Primary Care Assessment Survey (PCAS), a patient-completed questionnaire that operationalizes formal definitions of primary care, including the definition recently proposed by the Institute of Medicine Committee on the Future of Primary Care. METHODS The PCAS measures seven domains of care through 11 summary scales: accessibility (organizational, financial), continuity (longitudinal, visit-based), comprehensiveness (contextual knowledge of patient, preventive counseling), integration, clinical interaction (clinician-patient communication, thoroughness of physical examinations), interpersonal treatment, and trust. Data from a study of Massachusetts state employees (n = 6094) were used to evaluate key measurement properties of the 11 PCAS scales. Analyses were performed on the combined population and for each of the 16 subgroups defined according to sociodemographic and health characteristics. RESULTS The 11 PCAS scales demonstrated consistently strong measurement characteristics across all subgroups of this adult population. Tests of scaling assumptions for summated rating scales were well satisfied by all Likert-scaled measures. Assessment of data completeness, scale score dispersion characteristics, and inter-scale correlations provide strong evidence for the soundness of all scales, and for the value of separately measuring and interpreting these concepts. CONCLUSIONS With public and private sector policies increasingly emphasizing the importance of primary care, the need for tools to evaluate and improve primary care performance is clear. The PCAS has excellent measurement properties, and performs consistently well across varied segments of the adult population. Widespread application of an assessment methodology, such as the PCAS, will afford an empiric basis through which to measure, monitor, and continuously improve primary care.