Valid and reliable patient-reported outcome measures are vital for assessing disease impact, responsiveness to healthcare and the cost-effectiveness of interventions. A recent review has questioned the ability of existing measures to assess hypoglycaemia-related impacts on health-related quality of life for people with diabetes. This mixed-methods project was designed to produce a novel health-related quality of life patient-reported outcome measure in hypoglycaemia: the Hypo-RESOLVE QoL. Three studies were conducted with people with diabetes who experience hypoglycaemia. In Stage 1, a comprehensive health-related quality of life framework for hypoglycaemia was elicited from semi-structured interviews (N=31). In Stage 2, the content validity and acceptability of draft measure content were tested via three waves of cognitive debriefing interviews (N=70 people with diabetes; N=14 clinicians). In Stage 3, revised measure content was administered alongside existing generic and diabetes-related measures in a large cross-sectional observational survey to assess psychometric performance (N=1246). The final measure was developed using multiple evidence sources, incorporating stakeholder engagement. A novel conceptual model of hypoglycaemia-related health-related quality of life was generated, featuring 19 themes, organised by physical, social and psychological aspects. From a draft version of 76 items, a final 14-item measure was produced with satisfactory structural (χ2=472.27, df=74, p<0.001; comparative fit index =0.943; root mean square error of approximation =0.069) and convergent validity with related constructs (r=0.46–0.59), internal consistency (α=0.91) and test–retest reliability (intraclass correlation coefficient =0.87). The Hypo-RESOLVE QoL is a rigorously developed patient-reported outcome measure assessing the health-related quality of life impacts of hypoglycaemia. The Hypo-RESOLVE QoL has demonstrable validity and reliability and has value for use in clinical decision-making and as a clinical trial endpoint. All data generated or analysed during this study are included in the published article and its online supplementary files ( https://doi.org/10.15131/shef.data.23295284.v2 ).
Purpose:The objective of the current study was to conduct a rigorous assessment of the psychometric properties of the Victorian Institute of Sports Assessment-patellar tendinopathy (VISA-P).Methods:Rasch analysis, confirmatory factor analysis (CFA), and multivariable linear regression were used to assess the psychometric properties of the VISA-P questionnaire in 184 Danish patients with patellar tendinopathy who had symptoms ranging from under 3 months to over 1 year. A group of 100 healthy Danish persons was included as a reference for known-group validation.Results:The analyses revealed that the 8-item VISA-P did not fit a unidimensional model, yielded at best a 3-factor model, and exhibited differential item functioning (DIF) across healthy subjects versus people with patellar tendinopathy.Conclusion:VISA-P in its present form does not satisfy a measurement model and is not a robust scale for measuring patellar tendinopathy. A new PROM for patellar tendinopathy should be developed and appropriately validated, and meanwhile, simple pain scoring (e.g., numeric rating scales) and functional tests are suggested as more appropriate outcome measures for studies of patellar tendinopathy.
BackgroundPatient reported outcome measures (PROMs) are essential for evaluating treatment of ankle instability (AI). The aim was to assess the content validity and the measurement properties of all relevant PROMs for AI.MethodsRelevant PROMs were identified from PubMed and SCOPUS. The development and validation quality of the PROMs was assessed according to established scientific standards.ResultsSeventeen PROMs and 56 validation studies were analyzed. Content validity, which ensures the PROM measures what is relevant, is obtained by involving target patients in the development process. Only three PROMs identified had some degree of patient involvement (Cumberland Ankle Instability Tool (CAIT), Lower Extremity Function Scale (LEFS), and the Foot and Ankle Ability Measure (FAAM)). Of these, only FAAM was somewhat rigorously validated using modern psychometric validation methods, and exhibited superior measurement properties (construct validity).ConclusionNo existing PROM is completely adequate to evaluate AI. However, FAAM is the best choice.
BackgroundAssessment of patient-reported outcome measures (PROMs), including quality of life (QoL), is essential in diabetes research and care. However, a recent review concluded that current hypoglycaemia-specific PROMs have limited evidence of validity, reliability and responsiveness for assessing the impact of hypoglycaemia on QoL in people living with diabetes. None of the PROMs identified could be used directly to inform the cost-effectiveness of treatments and interventions. There is a need for a new hypoglycaemia-specific QoL PROM, which can be used directly to inform economic evaluations. AimsThis project has three aims: (a) To develop draft PROM content for measuring the impact of hypoglycaemia on QoL in adults with diabetes. (b) To refine the draft content using cognitive debriefing interviews and psychometrics. This will result in a condition-specific PROM that can be used to quantify the impact of hypoglycaemia upon QoL. (c) To generate a preference-based measure (PBM) that will enable utility values to be calculated for economic evaluation. MethodsA mixed-methods, three-stage design is used: (a) Qualitative interviews will inform the draft PROM content. (b) Cognitive debriefing interview data will be used to refine the draft PROM content. The PROM will be administered in a large-scale survey to enable psychometric validation. Final item selection for the PROM will be informed by psychometric performance, translatability assessment and input from stakeholder groups. (c) A classification system will be generated, comprising a reduced number of items from the PROM. A valuation survey will be conducted to derive a value set for the PBM.
Commentary In their study, Johnson et al. aimed to create a crosswalk between 2 patient-reported outcome measures (PROMs) commonly used to assess the effect of anterior cruciate ligament (ACL) reconstruction: the Knee Injury and Osteoarthritis Outcome Score (KOOS) and the International Knee Documentation Committee-Subjective Knee Form (IKDC-SKF). Crosswalking involves converting scores from one measure to another, in this case to enable the pooling of data for comparisons in meta-analyses and large-scale national and international ACL reconstruction registry studies, irrespective of which of the 2 PROMs was used. Crosswalking can be useful when scores from the same types of patients are available, yet for different measures and settings. For results from crosswalked PROMs to be trustworthy and valid, certain fundamental conditions must be satisfied. In this commentary, we point out that crosswalking the KOOS and the IKDC-SKF for use in registry studies is problematic for a series of reasons. Among these are the facts that: Neither the KOOS nor the IKDC-SKF was subjected to robust content validation for patients with an ACL injury during their creation1. Three of the 5 domains (33 of 42 items) in the KOOS are copied from the Western Ontario and McMaster Universities Osteoarthritis Index (WOMAC), which is a PROM developed in the mid-1980s to target patients with end-stage osteoarthritis. The 3 WOMAC-derived domains of the KOOS have inadequate construct validity for patients with ACL injury2. Crosswalking requires the involved PROMs to be unidimensional. However, the KOOS is not unidimensional and the IKDC-SKF is questionable. Johnson et al. refer to the KOOS homepage3 and contend that the mean scores from the 5 domains of the KOOS can be summed to yield a single composite score, but fail to mention that even the KOOS homepage opposes the use of a composite KOOS score (e.g., KOOS4 [the mean of 4 subscales of the KOOS, in which the subscale of Activities of Daily Living was not included] and KOOS5 [the mean of 5 subscales of the KOOS]), because neither the WOMAC nor the KOOS consists of a singular measurement construct. Simply adding subscale scores together and ignoring the underlying scaling properties of each subdomain is unacceptable in a measurement context. Therefore, using composite scores, as described in this study, carries the risk that information from important domains is blurred by scores from extraneous domains. This is a particular threat in relation to the KOOS, in which several domains are irrelevant to patients with ACL injury2. Furthermore, in contrast to statements by the authors, numerous robust psychometric analyses confirm that the KOOS is a multidimensional instrument2,4. To circumvent this, the authors have resorted to computing indices that reflect what is called “essential unidimensionality” to a theoretical global factor. This factor is assumed to represent the patients’ overall description of their condition, whereupon the authors maintain that the derived composite score is “unidimensional enough.” They base this on an explained common variance of between 0.55 and 0.73, meaning that 27% to 45% of scores in the KOOS4, KOOS5, and IKDC-SKF are not explained by the common global factor. Hence, 2 hypotheses should be tested: (1) that the KOOS domains measure the same unidimensional latent variable, and (2) that this KOOS composite score and the IKDC-SKF score measure the same unidimensional latent variable. Both of these could have been tested by confirmatory factor analysis using the reported data, and essential unidimensionality could similarly have been addressed using a bifactor model. When the multidimensional structure of the KOOS has already been demonstrated using the most robust methods, it makes no sense, scientifically, to ignore this fact until convincing bifactor analyses have been reported. Another problem that arises from crosswalking inadequate PROMs across national ACL reconstruction databases is that only about one-third of patients complete the PROMs at the time of follow-up5. The lowest generally accepted follow-up rate for valid data analysis from registries is 60%6. Moreover, differences between treatment strategies are best assessed through randomized controlled studies, and testing for clinically relevant differences rarely requires more than a few hundred patients. In addition, when data are pooled from various countries, it is necessary to adjust for differential item functioning (DIF) between the different language versions of a PROM4 to obtain valid results. DIF for language and cultural discrepancies is unaccounted for in the crosswalked scores, which further adds to the inaccuracy of the pooled data. Given these limitations, the justification for pooling or comparison of crosswalked data from ACL reconstruction registries is questionable. Pooled PROM data from the ACL reconstruction registries can be subject to regression analyses, but the statement that they have “great potential to improve our understanding of recovery after ACL reconstruction” is highly optimistic. The creation of a crosswalk between the KOOS and IKDC-SKF as presented in this study does not acknowledge the previously well-documented weak measurement properties of the 2 questionnaires. Results derived from crosswalked data must be interpreted with the utmost care, as the responsiveness and precision of such data are questionable. Perhaps the greatest risk of the current study is that future authors may use it as a reference to imply unidimensionality of the KOOS4, KOOS5, and IKDC-SKF. This would be a major scientific setback for the discipline of psychometric validation of PROMs for use in clinical trials. Composite scores reduce the validity of data from clinical trials. Such studies are increasingly used to establish health-care strategies, and the results should thus be based on data of the highest validity. Johnson et al. state that their study is a solution to the concerns about using the KOOS in athletic populations, as well as to the hesitancy and difficulty regarding the use of PROMs across national ACL reconstruction registries. Also, the authors believe that crosswalking KOOS data to IKDC-SKF data can eliminate such concerns. However, it is impossible to acquire valid data from a questionnaire that has been proven to be inadequate. The crosswalking of invalid data simply creates more invalid data. There is only one solution to the concerns raised in the introduction of the current study: switch to adequate outcome scores in ACL reconstruction registries as soon as possible and, thus, produce valid PROM data.
Introduction The aim was to evaluate content validity and measurement properties of patient reported outcome measures (PROMs) to assess patients with chronic ankle instability (CAI). Materials and Methods Potential PROMs for CAI and validity studies of these were identified in PubMed and SCOPUS. Development and validation methods for all PROMs were analyzed. Results Seventeen PROMs were relevant for CAI, and 56 validity studies were identified for the quality assessment. Only three PROMs had been developed with inputs from patients and were potentially adequate: the Cumberland Ankle Instability Tool (CAIT), the Lower-Extremity Functional Scale (LEFS) and the Foot and Ankle Ability Measure (FAAM). Measurement properties of CAIT has never been validated by modern test theory models (MTT), which are optimal for this purpose. In addition, CAIT is used to identify the presence of instability and not to evaluate the condition. Four analyses of LEFS with MTT methods for patients with an CAI have shown inadequate fit to the statistical model. For FAAM one study including CAI patients found adequate fit to the statistical model. Conclusion Fourteen (of seventeen) PROMs had been developed without involvement of patients and must be considered as inadequate measurement instruments. Of the three PROMs developed with patient involvement, only FAAM exhibited fit to the statistical model for patients with CAI. However, for other conditions evidence for construct validity for FAAM is inconsistent. No existing PROM possesses adequate content and construct validity for patients with CAI, but FAAM is suggested to be the best choice.
Introduction Content validity is the most important property of PROMs. The COSMIN guidelines are often referred to as gold standard to evaluate PROM properties. The aim of this study was by use of the COSMIN checklist to evaluate the content validity of five PROMs, all highly relevant in musculoskeletal research; the modified Harris’ Hip Score (mHHS), the Copenhagen Hip and Groin Outcome Score (HAGOS), the International Knee Documentation Committee Subjective Knee evaluation Form (IKDC-SKF), the Knee injury and Osteoarthritis Outcome Score (KOOS) and the Knee Numeric-Entity Evaluation Score ACL (KNEES-ACL). Materials and Methods Development articles were identified in PubMed and SCOPUS. A secondary literature search identified studies assessing content validity of the PROMs. Missing information was obtained from the five developers after direct request. To evaluate the quality of the development studies and rate the content validity, the COSMIN Risk of Bias checklist was applied to all relevant studies by two independent researchers. Results The development of mHHS, IKDC-SKF, and KOOS was rated inadequate, and these PROMs possess insufficient content validity. KOOS was in particular inappropriate to evaluate patients with ACL injury, but it is, despite this, the primary outcome in the Scandinavian ACL-reconstruction registries. The development of HAGOS was rated inadequate, although the insufficiency aspects can be regarded as minor. KNEES-ACL possessed sufficient content validity. Conclusion Out of five highly relevant orthopaedic PROMs, only KNEES-ACL possessed sufficient content validity according to COSMIN guidelines. There is an urgent need in musculoskeletal research for condition-specific PROMs developed with adequate methods.
Background: The Foot and Ankle Ability Measure (FAAM) was developed by involvement of patients with chronic ankle instability (CAI) and has acceptable measurement properties, but is not available in Danish.Methods: FAAM was translated and culturally adapted into Danish, and its measurement properties were assessed using Rasch analyses.Results: A Danish version was produced with small adaptations, and content relevance was confirmed by Danish patients. The 21-item ADL domain showed misfit to the Rasch model, but after removing six items, the resulting 15-item scale displayed adequate fit. The Sports domain also exhibited misfit, but after removing one item and adjusting due to differential item functioning related to age for another item, a 7-item scale showed good fit. This resulted in a 22-item 2-dimensional Danish version of FAAM.Conclusion: The 22-item Danish FAAM exhibits robust measurement properties for patients with various conditions of the lower leg, ankle, and foot, including CAI.(c) 2021 The Author(s). Published by Elsevier Ltd on behalf of European Foot and Ankle Society. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
Translating patient‐reported outcome measures (PROMs) can alter the meaning of items and undermine the PROM's psychometric properties (quantified as cross‐cultural differential item functioning [DIF]). The aim of this paper was to present the theoretical background for PROM translation, adaptation, and cross‐cultural validation, and assess how PROMs used in sports medicine research have been translated and adapted. We also assessed DIF for the Knee Injury and Osteoarthritis Outcome Score (KOOS) across Danish, Norwegian, and Swedish versions. We conducted a search in PubMed and Scopus to identify the method of translation, adaptation, and validation of PROMs relevant to musculoskeletal research. Additionally, 150 preoperative KOOS questionnaires were obtained from the Scandinavian knee ligament reconstruction registries, and cross‐cultural DIF was evaluated using confirmatory factor analysis and Rasch analysis. There were 392 studies identified, describing the translation of 61 PROMs. Ninety‐four percent were performed with forward‐backward technique. Forty‐nine percent used cognitive interviews to ensure appropriate wording, understandability, and adaptation to the target culture. Only two percent were validated according to modern test theory. No study assessed cross‐cultural DIF. One KOOS subscale showed no cross‐cultural DIF, two had DIF with respect to some (but not all) items, and thus conversion tables could be constructed, and two KOOS subscales could not be pooled. Most PROM translations are of undocumented quality, despite the common conclusion that they are valid and reliable. Scores from three of five KOOS subscales can be pooled across the Danish, Norwegian, and Swedish versions, but two of these must be adjusted for DIF.
Choosing the most adequate PROM for a study is a non‐trivial process. The aim of this study was to provide a catalogue with analyses of content and construct validity of PROMs relevant to research in sports science, including all published local translations. The most commonly used PROMs in sports research were selected from a PubMed search “patient reported outcome measures sports”, identifying 439 articles and 194 different PROMs. Articles describing development of the 61 selected PROMs were assessed for content validity, and all articles regarding construct validity of each PROM and all published translations (in total 622 articles) were analyzed. A catalogue with assessments of the 61 PROMs was produced. The majority were of inferior validity, with few exceptions. The most common reason for this was that the PROM had not been developed by methods that ensure high content validity. Another major reason for inferior validity was that construct validity had not been secured by adequate statistical methods. In conclusion, this catalogue provides a tool for researchers to facilitate choosing the most valid PROM for studies in sports research. Furthermore, it shows for popular PROMs where further validation is needed, and for fields in musculoskeletal medicine where valid PROMs are lacking. It is suggested that a targeted effort is made to develop valid PROMs for major conditions in musculoskeletal research. The current method is easier to practice compared with assessment after COSMIN guidelines.
The purpose of this article was to introduce the reader to the nature of patient-reported outcome measures (PROMs) and pitfalls in their use. PROMs collect subjective information directly from the patient regarding specific or general conditions and add to clinical and functional outcomes, and turn unmeasurable subjective qualities into quantitative measures. PROMs are questionnaires consisting of items: questions or statements with predefined response options. The items in an adequate PROM have been developed by involvement of patients with the condition in focus, and the PROM has been validated for these patients using suitable statistical methods. An adequate well-targeted PROM is more responsive than an inadequate PROM. Unfortunately, many studies use inadequate PROMs as outcomes. The methods used to generate PROMs should be described as thoroughly as those used to develop any other types of measurement instruments, and the choice of PROM should always be explained and thereby justified. If the PROM used is not adequate, the consequences for the interpretation of the results should be discussed. In many cases, an adequate PROM does not exist. If the best available PROM is chosen, there are methods to validate the adequacy of the chosen PROM, which make an interpretation of the study results possible.
A recent COSMIN review found that the Victorian Institute of Sports Assessment-Achilles tendinopathy questionnaire (VISA-A) has flawed construct validity. The objective of the current study was to assess specifically the process of how VISA-A was constructed and validated, and whether the Danish version of VISA-A is a valid patient-reported outcome measure (PROM) for measuring the perceived impact of Achilles tendinopathy. The original item generation strategy for content validity and the process for confirming the scaling properties (construct validity) were examined. In addition, construct validity was evaluated directly using several psychometric methods (Rasch analysis, confirmatory factor analysis (CFA), and multivariable linear regression) in a cohort of 318 persons with Achilles tendinopathy with symptom duration groups ranging from less than 3 months to more than 1 year of chronicity, and a group of 120 healthy persons. We found that the item generation and item reduction in the original construction of VISA-A was based on literature review and clinician consensus with little or no patient involvement. We determined that 1) VISA-A consists of ambiguous conceptual item themes and thus lacks content validity, 2) there was no thorough investigation of the psychometric properties of the original version of VISA-A, which thus lacks construct validity, and 3) rigorous direct assessment of the psychometric properties of the Danish VISA-A revealed inadequate psychometric properties. In agreement with the COSMIN study, we conclude that when used as a single score, VISA-A is not an adequate scale for measuring self-reported impact of Achilles tendinopathy.
Several terms are used to describe changes in PROM scores in relation to treatments. Whether the change is small, large, or relevant is defined in different ways, yet these change scores are used to recommend or oppose treatments. They are also used to calculate the necessary number of patients for a study. This article offers a theoretical explanation behind the terms responsiveness, minimal important difference (MID), minimal important change (MIC), minimal relevant difference (MIREDIF), and threshold of clinical importance. It also gives instructions on how these and the optimal number of patients for a study are calculated. Responses to two domains of the Knee Injury and Osteoarthritis Outcome Score (KOOS), before and 1 year after reconstruction of the anterior cruciate ligament of 164 patients, are used to illustrate the calculations. This paper presents the most common methods used to calculate and interpret MID. Results vary substantially across domains, patient location on the scale, and health conditions. The optimal number of patients depends on the minimal relevant difference (MIREDIF), the standard error of the measure (SEM), the desired statistical power for the measurement, and the responsiveness of the measurement instrument (the PROM). There is often uncertainty surrounding the calculation and interpretation of responsiveness, MID, and MIREDIF, as these concepts are complex. When MID is used to evaluate research results, authors should specify how the MID was calculated, and its relevance for the study population. These measures should only be used after thorough consideration to justify healthcare decisions.
Choosing the most appropriate patient‐reported outcome measure (PROM) for a clinical study is essential in order to achieve trustworthy results. This choice will depend on (a) the objective of the study and hence the research question; (b) the choice of a theoretical framework, such as the World Health Organization's International Classification of Functioning, Disability, and Health (ICF); (c) whether there currently is a PROM that possesses high content validity and high construct validity for the specific patient group and objective, and if not; (d) the decision on whether to use a suboptimal PROM or develop and validate a new PROM. This paper presents the steps that should be followed in order to assess the relevance of PROMs and suggests ways to enhance the choice depending on the goal of the study.
Results by patient‐reported outcome measures (PROMs) from randomized controlled trials (RCTs) in musculoskeletal research often influence healthcare strategies. We aimed to evaluate to which extent these RCTs use adequate PROMs, and how this influences the results and conclusions. We identified RCTs of sports research relevance with PROMs as primary outcomes published in 13 preselected journals between January 1, 2008, and November 1, 2019; all journals regularly publish results from musculoskeletal research. Five journals have a high impact factor (>15), and eight with lower impact factors are widely read journals. It was assessed whether the RCTs had used PROMs with high content validity and whether the most adequate PROMs were used (ie, the most well developed and well validated for the patients enrolled in the study). We registered journal impact factor, year of publication, existence of a registered protocol, and whether the study showed significant difference between interventions. A total of 54 RCTs with 56 primary outcomes comprising 26 different PROMs were identified. For 13 RCTs (24%), a protocol was not published. In only 24 of RCTs (44%), the most appropriate PROM had been used as primary outcome, independent of a registered protocol, ranking of the journal, and year of publication. In seven cases, PROMs were used to evaluate a condition that they had not been developed for. RCTs that used the most adequate PROM showed significantly more often (46%) difference in outcomes in contrast to RCTs that used inadequate PROMs (22%) ( P = 0.0483). In conclusion, in the majority of RCTs, the most adequate PROM had not been used. Studies, in which the most adequate PROM had been used as outcome, were significantly more likely to show significant difference between interventions. The extent to which protocols were not available was surprisingly high. Journals should request that adequate PROMs are used in RCTs, and if this is not the case that it is discussed how it might influence the results and conclusions. Likewise, it should be requested that a protocol is published or registered.
PURPOSE:Content validity is the most important property of PROMs. The COSMIN initiative has published guidelines for evaluating the content validity of PROMs, but they have only sparsely been applied to relevant PROMs for musculoskeletal conditions. The aim of this study was to use the COSMIN Risk of Bias checklist to evaluate the content validity of five PROMs, that are highly relevant in musculoskeletal research and used by the arthroscopic surgery community: the modified Harris' Hip Score (mHHS), the Copenhagen Hip and Groin Outcome Score (HAGOS), the International Knee Documentation Committee Subjective Knee evaluation Form (IKDC-SKF), the Knee injury and Osteoarthritis Outcome Score (KOOS) and the Knee Numeric-Entity Evaluation Score ACL (KNEES-ACL).METHODS:The development articles for the five PROMs were identified through searches in PubMed and SCOPUS. A literature search was performed to identify additional studies assessing content validity of the PROMs. Additional information, necessary for the assessments, was obtained from the PROM developers after direct request. To evaluate the quality of the development studies and rate the content validity, the COSMIN Risk of Bias checklist was applied to all studies.RESULTS:All five development studies were identified. Three subsequent content validity studies were identified, all evaluating KOOS and one also IKDC. One content validity study was of inadequate quality and excluded from further analysis. The development of mHHS, IKDC-SKF, and KOOS was rated inadequate and possess insufficient content validity for their target populations. Due to the irrelevance of multiple items, KOOS was in particular inappropriate to evaluate patients with an ACL injury. The development of HAGOS was also rated inadequate, although the insufficiency aspects can be regarded as minor. KNEES-ACL possessed sufficient content validity.CONCLUSION:Out of five PROMs, only KNEES-ACL possessed sufficient content validity. Particularly, KOOS should not be used as an outcome for patients with an ACL injury. There is an urgent need for condition-specific PROMs for musculoskeletal conditions, developed with adequate methods.LEVEL OF EVIDENCE:III.
Study Design. Registry-based repeated-measures psychometric validation of the Danish Oswestry Disability Index (ODI). Objective. The goal was to use classical and modern psychometric validation methods to assess the measurement properties and the minimally clinical important difference (MCID) of the ODI in a Danish cohort of patients with chronic low back pain being treated with spinal surgery. Summary of Background Data. Scores for the ODI, EQ-5D, SF-36, leg pain, back pain, and a general rating of pain item from 800 patients with chronic low back pain were extracted from the National Danish Spine Registry (DaneSpine) at baseline and 1-year postspine surgery. Methods. Confirmatory factor analysis and item response theory (IRT) models were used to assess the psychometric properties of the ODI. MCID was also calculated based on generic legacy PROMs (EQ-5D and SF-36) and follow-up pain scores. Results. While ODI did not fit a Rasch model, adequate fit to a confirmatory factor analysis and a two-parameter item response theory model was found when accounting for differential item functioning across diagnostic subgroups (degenerative spondylolisthesis, spondylosis, spinal stenosis, and herniated intervertebral disc). In addition, each group exhibited substantially different MCID values. Conclusion. The Danish version of the ODI is valid and responsive, but only within each of the four major diagnosis subgroups: degenerative spondylolisthesis, spondylosis, spinal stenosis, and herniated intervertebral disc. Level of Evidence : 4