ObjectivesTo evaluate the psychometric properties of the Diary for Irritable Bowel Syndrome Symptoms-Constipation (DIBSS-C), which was developed to support primary and secondary endpoints in irritable bowel syndrome (IBS) with predominant constipation (IBS-C) clinical trials.MethodsObservational data were collected from 108 adults with IBS-C using a smartphone-type device for 17 days. DIBSS-C data regarding bowel movements (BMs) were collected for each event (along with the Bristol Stool Form Scale); abdominal symptoms were rated each evening. Global status items and the Gastrointestinal Symptom Rating Scale-IBS were completed on day 10 and day 17 and the IBS-Symptom Severity Scale on day 17. Item-level performance, internal consistency reliability, test-retest reliability, and construct validity were evaluated.ResultsThe Abdominal Symptoms Domain score demonstrated high internal consistency reliability (Cronbach’s alpha week 1 = 0.98; week 2 = 0.96) and test-retest reliability (intraclass correlation coefficient [ICC] = 0.93). Test-retest reliability was stronger for abdominal symptoms (ICC = 0.91-0.94) than for the frequency-based BM-related outcomes (ICC = 0.54-0.66). Key construct validity hypotheses were supported by moderate to strong correlations with the corresponding Gastrointestinal Symptom Rating Scale-IBS, IBS-Symptom Severity Scale, and Bristol Stool Form Scale items. All known-groups comparisons were statistically significant for the abdominal symptom items and domain score; evidence for known-groups validity of BM-related outcomes was supportive when based on constipation severity.ConclusionsThe results of this study provided key psychometric evidence for the DIBSS-C, ultimately contributing to its qualification by the US Food and Drug Administration for use in IBS-C clinical trials.
The ISPOR Task Force on measurement comparability between modes of data collection for patient-reported outcome measures (PROMs) has updated the good practice recommendations from the 2009 ISPOR electronic patient-reported outcome and 2014 patient-reported outcome mixed modes Good Research Practices Task Force reports in light of accumulated evidence of measurement comparability among different modes of PROM data collection. Furthermore, with the increasing use of electronic formats of clinical outcome assessments in clinical trials and the US Food and Drug Administration's encouragement of electronic data collection, this new task force report provides stakeholders with best practice recommendations reflecting the current body of evidence and enables them to respond to future developments in research and technology. This task force recommends an evidence-based approach to determine whether new research is needed to evaluate measurement comparability for a given questionnaire or technology. The suitability of existing evidence depends upon whether it satisfactorily demonstrates that the change in data collection mode has not affected the PROM's measurement properties. In cases where sufficient evidence of measurement comparability exists and best practices for faithful migration are followed, this task force concludes that further testing of measurement comparability among the data collection modes is unnecessary, including cases of "mixing modes" within clinical trials such as bring your own device designs.
In evaluating the clinical benefit of new therapeutic interventions, it is critical that the treatment outcomes assessed reflect aspects of health that are clinically important and meaningful to patients. Performance outcome (PerfO) assessments are measurements based on standardized tasks actively undertaken by a patient that reflect physical, cognitive, sensory, and other functional skills that bring meaning to people's lives. PerfO assessments can have substantial value as drug development tools when the concepts of interest being measured best suit task performance and in cases where patients may be limited in their capacity for self-report. In their development, selection, and modification, including the evaluation and documentation of validity, reliability, usability, and interpretability, the good practice recommendations established for other clinical outcome assessment types should continue to be followed, with concept elicitation as a critical foundation. In addition, the importance of standardization, and the need to ensure feasibility and safety, as well as their utility in patient groups, such as pediatric populations, or those with cognitive and psychiatric challenges, may enhance the need for structured pilot evaluations, additional cognitive interviewing, and evaluation of quantitative data, such as that which would support concept confirmation or provide ecological evidence and other forms of construct evidence within a unitary approach to validity. The opportunity for PerfO assessments to inform key areas of clinical benefit is substantial and establishing good practices in their selection or development, validation, and implementation, as well as how they reflect meaningful aspects of health is critical to ensuring high standards and in furthering patient-focused drug development.
OBJECTIVES:Evaluating the clinical benefit of interventions for conditions with heterogeneous symptom and impact presentations is challenging. The same condition can present differently across and within individuals over time. This occurs frequently in rare diseases. The purpose of this review was to identify (1) assessment approaches used in clinical trials to address heterogeneous manifestations that could be relevant in rare disease research and (2) US Food and Drug Administration (FDA)-approved labeling claims that used these approaches.METHODS:A targeted literature review was conducted examining peer-reviewed publications and FDA-approved labeling claims from January 2002 to July 2020, focusing on claims incorporating clinical outcome assessments. Approaches were then assessed for their potential application in rare diseases.RESULTS:A total of 6 assessment approaches were identified: composite or other multicomponent endpoints, multidomain responder index, most bothersome symptom (MBS), goal attainment scaling, sliding dichotomy, and adequate relief. A total of 59 FDA-approved labeling claims associated with these approaches were identified: composite or other multicomponent endpoints (n=49), MBS (n=9), and adequate relief (n=1). A total of 10 FDA-approved labeling claims, all using multicomponent endpoints, were identified for rare diseases.CONCLUSIONS:Multicomponent, MBS, and adequate relief have been included in FDA-approved labeling claims. Multicomponent endpoints, including composite endpoints, were the most frequent way to address heterogeneous manifestations of both common and rare diseases. MBS may be acceptable to regulators, whereas multidomain responder index is unlikely to be. The goal attainment scaling and adequate relief approaches may have potential utility in rare disease trials, assuming the theoretical and statistical challenges inherent in each approach are managed.
Introduction:The NSCLC Symptom Assessment Questionnaire (NSCLC-SAQ) was developed to assess NSCLC symptom severity in accordance with Food and Drug Administration evidentiary expectations leading to Food and Drug Administration qualification in 2018. This study evaluated the NSCLC-SAQ's measurement properties within a clinical trial.Methods:The KEYNOTE-598 phase 3 study of participants with stage IV metastatic NSCLC with programmed death-ligand 1 tumor proportion score greater than or equal to 50% was used to assess the NSCLC-SAQ's reliability, construct validity, responsiveness, and estimate clinically meaningful within-person change. Other patient-reported outcome measures included patient global impression items of severity and change in lung cancer symptoms, and the European Organisation for Research and Treatment of Cancer Quality of Life Questionnaire core 30 and lung cancer module, LC13.Results:Participants (N = 560) were mostly men (70%), had a mean age of 64 years, and had Eastern Cooperative Oncology Group performance status of 1 (64%) or 0 (36%). Internal consistency at baseline (Cronbach's α = 0.74) and test-retest reliability after 3 weeks (intraclass correlation coefficient = 0.79) were satisfactory. NSCLC-SAQ items, domains, and total score correlated moderately to highly with patient-reported outcome measures capturing similar content, and the total score differentiated among patient global impression of severity groups (p < 0.001). The total score detected improvement over time and the estimated clinically meaningful within-person change threshold for improvement ranged from three to five points on the 0 to 20 scale. Few participants exhibited symptom worsening (n = 38), limiting inferences in this group.Conclusions:The NSCLC-SAQ was found to be reliable, valid, responsive, and interpretable for assessing symptom improvement in NSCLC. Further evaluation is recommended in trial participants whose symptoms worsen over time.
Score reproducibility is an important measurement property of fit-for-purpose patient-reported outcome (PRO) measures. It is commonly assessed via test–retest reliability, and best evaluated with a stable participant sample, which can be challenging to identify in diseases with highly variable symptoms. To provide empirical evidence comparing the retrospective (patient global impression of change [PGIC]) and current state (patient global impression of severity [PGIS]) approaches to identifying a stable subgroup for test–retest analyses, 3 PRO Consortium working groups collected data using both items as anchor measures. The PGIS was completed on Day 1 and Day 8 + 3 for the depression and non-small cell lung cancer (NSCLC) studies, and daily for the asthma study and compared between Day 3 and 10. The PGIC was completed on the final day in each study. Scores were compared using an intraclass correlation coefficient (ICC) for participants who reported “no change” between timepoints for each anchor. ICCs using the PGIS “no change” group were higher for depression (0.84 vs. 0.74), nighttime asthma (0.95 vs. 0.53) and daytime asthma (0.86 vs. 0.68) compared to the PGIC “no change” group. ICCs were similar for NSCLC (PGIS: 0.87; PGIC: 0.85). When considering anchor measures to identify a stable subgroup for test–retest reliability analyses, current state anchors perform better than retrospective anchors. Researchers should carefully consider the type of anchor selected, the time period covered, and should ensure anchor content is consistent with the target measure concept, as well as inclusion of both current and retrospective anchor measures.
Background: The Non-Small Cell Lung Cancer Symptom Assessment Questionnaire (NSCLC-SAQ) was developed to incorporate the patient's perspective into evaluation of clinical benefit in advanced non-small cell lung cancer trials and meet regulatory expectations for doing so. Qualitative evidence supported 7 items covering 5 symptom concepts. Objective: This study evaluated measurement properties of the NSCLC-SAQ's items, overall scale, and total score. Methods: In this observational cross-sectional study, a purposive sample of patients with cliniciandiagnosed advanced non-small cell lung cancer, initiating or undergoing treatment, provided sociodemographic information and completed the NSCLC-SAQ, National Comprehensive Cancer Network/Functional Assessment of Cancer Therapy Lung Symptom Index (FLSI-17), and a Patient Global Impression of Severity item. Rasch analyses, factor analyses, and assessments of construct validity and reliability were completed. Results: The 152 participants had a mean age of 64 years, 57% were women, and 87% where White. The majority were Stage IV (83%), 51% had an Eastern Cooperative Oncology Group performance status of 1 (32% performance status 0 and 17% performance status 2), and 33% were treatment naive. Rasch analyses showed ordered thresholds for response options. Factor analyses demonstrated that items could be combined for a total score. Internal consistency (Cronbach alpha= 0.78) and test-retest reliability (intraclass correlation coefficient = 0.87) were quite satisfactory. NSCLC-SAQ total score correlation was 0.83 with the National Comprehensive Cancer Network/Functional Assessment of Cancer Therapy Lung Symptom Index-17. The NSCLC-SAQ was able to differentiate between symptom severity levels and performance status (both P values <.001). Conclusions: The NSCLC-SAQ generated highly reliable scores with substantial evidence of construct validity. The Food and Drug Administration's qualification supports the NSCLC-SAQ as a measure of symptoms in drug development. Further evaluation is needed on its longitudinal measurement properties and interepretation of meaningful within-patient score change. (C) 2021 The Authors. Published by Elsevier Inc.
While the EQ-5D-5L has been migrated to several electronic modes, evidence supporting the measurement equivalence of the original paper-based instrument to the electronic modes is limited. This study was designed to comprehensively examine the equivalence of the paper and electronic modes (i.e., handheld, tablet, interactive voice response [IVR], and web). As part of the foundational work for this study, the test–retest reliability of the paper-based, UK English format of the EQ-5D-5L was assessed using a single-group, single-visit, two-period, repeated-measures design. To compare paper and electronic modes, three independent samples were recruited into a three-period crossover study. Each participant was assigned to one of six groups to account for order effects. Descriptive statistics, mean differences (i.e., split-plot analysis of variance [ANOVA]), and intraclass correlation coefficients (ICCs) were calculated. The test–retest results showed mean differences near zero and ICC values > 0.90 for both the index and the EQ VAS scores. For the electronic comparisons, mean difference confidence intervals (CIs) for the EQ-5D index scores and EQ VAS scores reflected equivalence of the means across all modes, as the CIs were wholly contained inside the equivalence interval. Further, the ICC 95% lower CIs for the index and EQ VAS scores showed values above the thresholds for denoting equivalence across all comparisons in each sample. No significant mode-by-order interactions were present in any ANOVA model. Overall, our comparisons of the paper, screen-based, and phone-based formats of the EQ-5D-5L provided substantial evidence to support the measurement equivalence of these modes of data collection.
Background: The Symptoms of Major Depressive Disorder Scale (SMDDS) was expressly developed on the basis of qualitative data to directly incorporate patients' voices into evaluation of treatment benefit in major depressive disorder (MDD) clinical trials. Objectives: To collect quantitative data necessary to refine/optimize the SMDDS and document its psychometric properties. Methods: In this multicenter, observational study, participants with clinically diagnosed MDD completed questionnaires in 2 waves. Wave 1 was designed to refine the SMDDS using Rasch measurement evaluations and item reduction analyses. On a subset of wave 1 subjects, 7 to 12 months later, wave 2 further examined item performance and measurement properties. Exploratory factor analyses and assessments of construct validity and reliability (internal consistency and reproducibility) were completed. Results: Using wave 1 data (N = 315; females = 71%, white = 81%, mean age = 44 years), the SMDDS was revised from 36 to 16 items. The Rasch item threshold map indicated that all but 1 item (suicidal ideation) were appropriately ordered. The 207 wave 2 participants were 74% females, 82% white, with a mean age of 45 years. The exploratory factor analyses resulted in a single component (all standardized factor loadings>0.46). Cronbach alpha was 0.93 and the 7-day test-retest intraclass correlation coefficient (n = 93) was 0.84 (95% confidence interval 0.77-0.89). SMDDS scores discriminated between MDD severity levels. Conclusions: The 16-item SMDDS generated highly reliable scores with substantial evidence of construct validity. On the basis of the evidence of appropriate content validity and sound psychometric performance, the Food and Drug Administration qualified the SMDDS as an outcome measure to support exploratory efficacy endpoints in MDD clinical trials.
Assessment of clinical benefit in treatment trials can be made through report by a clinician, a patient, or a nonclinician observer (eg, caregiver) or through a performance-based assessment. The US Food and Drug Administration (FDA) published a final guidance for industry for one type of clinical outcome assessment (COA)-patient-reported outcome (PRO) measures-in 2009 that described how FDA reviews PRO measures for their adequacy to support medical product-labeling claims. Many of the principles described in the PRO Guidance could be applicable to the other types of COAs, including instruments completed by clinicians (ie, clinician-reported outcome assessments) and nonclinician observers (ie, observer-reported outcome assessments). FDA guidance describing the regulatory expectations for all COA types including performance outcome assessments, which are based on the patient's performance of a defined task or activity, is in progress to meet requirements described within the 21st Century Cures Act and PDUFA VI. This communication highlights potential ways in which existing instruments might be modified or used "as is" to conform to good measurement principles. An industry and a regulatory perspective are described.
To evaluate the psychometric properties of the Diary for Irritable Bowel Syndrome Symptoms-Constipation (DIBSS-C), one of three subtype-specific measures developed by the Patient-Reported Outcome Consortium’s Irritable Bowel Syndrome (IBS) Working Group, for qualification by the United States (US) Food and Drug Administration to support primary and secondary endpoints in IBS clinical trials. Observational data were collected from 108 adults with IBS-C at 10 clinical sites across the US. Using a handheld electronic device, participants completed the DIBSS-C for 17 consecutive days; data related to bowel movements (BMs) were collected on an event-driven basis and abdominal symptoms were rated each evening. Global status items and the Gastrointestinal Symptom Rating Scale-IBS (GSRS-IBS) were completed on Days 10 and 17 and the IBS-Symptom Severity Scale (IBS-SSS) on Day 17. Data collected on Days 4 through 17 (Week 1 and Week 2) were analyzed to assess item-level performance, internal consistency, test-retest reliability, construct validity, and discriminating ability. The abdominal symptom items were highly correlated with each other and the subscale comprised of these items produced high estimates of internal consistency (Cronbach’s alpha Week 1: 0.98 and Week 2: 0.96) and test-retest reliability (intraclass correlation coefficient: 0.93). As anticipated, test-retest reliability evidence was not as strong for the frequency-based BM-related outcomes. Key construct validity hypotheses were supported for both abdominal and BM-related items (i.e., correlations with GSRS-IBS and IBS-SSS items addressing similar concepts were higher than others). While all comparisons were statistically significant for the abdominal symptoms, evidence for the discriminating ability of BM-related outcomes was strongest for known-groups comparisons based on constipation severity. Overall, the results provide strong support for the reliability and validity of the DIBSS-C. Data from a recent clinical trial are expected to confirm these findings, support the measure’s responsiveness, and identify thresholds to support the interpretation of changes in scores.
The use of performance outcome (PerfO) assessments to measure cognitive or physical function in drug trials presents several challenges for both sponsors and regulators, owing in part to a relative lack of scientific guidance on their development, implementation, and interpretation. In December 2016, the Duke-Margolis Center for Health Policy convened a 2-day workshop to explore the evidentiary, methodologic, and operational challenges associated with PerfO measures, and to identify potential paths to addressing these challenges. This paper presents both a summary of the discussion as well as additional input from a working group of experts from FDA, industry, academia, and public-private consortia. It is intended to advance the discussion around the development and use of PerfO measures to assess patient functioning in clinical trials intended to support registration of new treatments, and to highlight the key gaps in knowledge where additional research, collaboration, and discussion are needed.
PURPOSE:The US Food and Drug Administration (FDA) 2009 guidance for industry on patient-reported outcome (PRO) measures describes how the Agency evaluates the psychometric properties of measures intended to support medical product labeling claims. An important psychometric property is test-retest reliability. The guidance lists intraclass correlation coefficients (ICCs) and the assessment time period as key considerations for test-retest reliability evaluations. However, the guidance does not provide recommendations regarding ICC computation, nor is there consensus within the measurement literature regarding the most appropriate ICC formula for test-retest reliability assessment. This absence of consensus emerged as an issue within Critical Path Institute's PRO Consortium. The purpose of this project was to generate thoughtful and informed recommendations regarding the most appropriate ICC formula for assessing a PRO measure's test-retest reliability.METHODS:Literature was reviewed and a preferred ICC formula was proposed. Feedback on the chosen formula was solicited from psychometricians, biostatisticians, regulators, and other scientists who have collaborated on PRO Consortium initiatives.RESULTS AND CONCLUSIONS:Feedback was carefully considered and, after further deliberation, the proposed ICC formula was confirmed. In conclusion, to assess test-retest reliability for PRO measures, the two-way mixed-effect analysis of variance model with interaction for the absolute agreement between single scores is recommended.
BACKGROUND:The purpose of this literature review was to examine the existing patient-reported outcome measurement literature to understand the empirical evidence supporting response scale selection in pain measurement for the adult population.METHODS:The search strategy involved a comprehensive, structured, literature review with multiple search objectives and search terms.RESULTS:The searched yielded 6918 abstracts which were reviewed against study criteria for eligibility across the adult pain objective. The review included 42 review articles, consensus guidelines, expert opinion pieces, and primary research articles providing insights into optimal response scale selection for pain assessment in the adult population. Based on the extensive and varied literature on pain assessments, the adult pain studies typically use simple response scales with single-item measures of pain-a numeric rating scale, visual analog scale, or verbal rating scale. Across 42 review articles, consensus guidelines, expert opinion pieces, and primary research articles, the NRS response scale was most often recommended in these guidance documents. When reviewing the empirical basis for these recommendations, we found that the NRS had slightly superior measurement properties (e.g., reliability, validity, responsiveness) across a wide variety of contexts of use as compared to other response scales.CONCLUSIONS:Both empirical studies and review articles provide evidence that the 11-point NRS is likely the optimal response scale to evaluate pain among adult patients without cognitive impairment.
The collection of electronic patient-reported outcome (ePRO) data in clinical trials presents an opportunity to minimize missing data by requiring subjects to respond to all items in order to complete the questionnaire. However, implementation of this data entry rule can have unintended consequences. The purpose of this report is to share considerations around requiring subjects to respond to items and provide data on the prevalence of skipped items in three therapeutic areas and ePRO modes. Three quantitative pilot studies conducted by the PRO Consortium allowed participants to skip items on the draft questionnaires, one of three scenarios described by O’Donohoe et al. (2015) on considerations for requiring completion. Use of an “active skip” ensured that participants indicated they were choosing to skip an item, and that it was not missed accidentally. Data on skipped items were analysed from the Non-Small Cell Lung Cancer Symptom Assessment Questionnaire (NSCLC-SAQ) on a tablet device, the Symptoms of Major Depressive Disorder Scale (SMDDS) on a web-based system, and the Asthma Daily Symptom Diary (ADSD) on a handheld device. Diverse samples were recruited for the NSCLC-SAQ (N=152), SMDDS (Wave 1=315; Wave 2=207), and ADSD (N=219) studies. No items were skipped on the NSCLC-SAQ, while rates of item-level skipping ranged from 0.09% to 2% of possible completions on the SMDDS and ADSD, respectively. Missing data appeared to be at random and did not indicate problems with the items skipped. Requiring completion of items may reduce missing data but can result in questionable data. Careful implementation of skipping rules and the use of well-designed questionnaires assessing relevant and appropriate concepts for the context of use may reduce respondents’ desire to skip items when allowed to do so, as evidenced by the low rates of missing item-level data seen in three PRO Consortium studies.