Objectives To explore priorities, barriers to and experiences of palliative and end-of-life care from the perspectives of people living with HIV.Design Cross-sectional online survey conducted in the UK between September 2024 and November 2024.Setting Online survey of people living with HIV.Participants The sample (N=90) was adults living with HIV in the UK. The majority of participants were male (82.4%), gay men (77.8%) and white (88.1%).Results The majority of participants (58.9%) reported knowing what palliative care was and that they could explain it to someone else; however, a misconception about palliative care being only for the end of life was evident. Over a quarter of respondents (27.8%) reported that their HIV status ‘Sometimes’ negatively affected their experiences of care in general practitioner, hospital and dental settings. The top three priorities for end of life were (1) being in a calm atmosphere, (2) being free of pain and (3) support with psychological well-being. Not being judged was also identified as a priority.Conclusion To promote integration of palliative and end-of-life care into care pathways for people living with HIV, partnerships with HIV services and charities may be needed as well as tailored messaging and training for staff in generalist services.
Introduction Colorectal cancer has low survival rates when diagnosed late-stage. We previously developed sex-specific dynamic risk prediction models utilising trends in the full blood count (FBC), a blood test commonly performed in primary care, to support early detection. We aimed to externally validate these prediction models. Methods We performed a hybrid case-control and cohort study of patients with at least one FBC test. We first excluded FBCs within two years before diagnosis (cases) or study exit (controls) and selected the most recent FBC as the baseline test per patient from the resulting data. Patients were aged at least 40 years at baseline and had no history of colorectal cancer. The models included age (years) at baseline and simultaneous trends over historical haemoglobin, mean corpuscular volume (MCV), and platelet measurements measured over five years before baseline to inform two-year risk of colorectal cancer diagnosis. Performance measures included the c-statistic and calibration slope. Results We included 2,956,977 males and 3,561,349 females, with 0.4% (n=12,578) and 0.3% (n=11,939) diagnosed with colorectal cancer, respectively. The c-statistic (95% CI) was 0.73 (0.72-0.73) for males and 0.74 (0.74-0.75) for females. The calibration slope (95% CI) was 0.92 (0.89-0.94) for males and 0.95 (0.93-0.98) for females. Calibration was good in subgroups of patient data, except under-predicted risk in those aged 70+ years, White individuals, and those with higher IMD. The c-statistic (95% CI) was similar regardless of the number of FBCs used to define trend and increased as the longitudinal trend window increased until around 2.5-3.0 years for men (0.73 (0.71-0.74)) and 3.0-3.5 years for women (0.73 (0.72-0.75)) and decreased with increasing longitudinal windows thereafter. Conclusion Utilising temporal changes in the FBC test could enhance risk stratification for colorectal cancer. Further research may highlight approaches for improving predictive performance further. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was funded by a National Institute for Health and Care Research (NIHR) School for Primary Care Research (SPCR) Post-doctoral fellowship for this work (award number: C092) and the NIHR Policy Research Programme (Policy Research Unit on Cancer Awareness, Screening and Early Diagnosis, reference PR-PRU-NIHR206132). This report presents independent research and the views expressed are those of the authors and not necessarily those of the NIHR, SPCR, or Department of Health and Social Care. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Research Data Governance Committee of the Clinical Practice Research Datalink gave ethical approval for this work I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The dataset used is available from the authors but is subject to access approval by the CPRD[[33][1]]. [1]: #ref-33
Background: Early detection of colorectal cancer confers substantial prognostic benefit. Most symptoms are non-specific and easily missed. The ColonFlag algorithm identifies risk of undiagnosed colorectal cancer using age, sex and changes in full blood count (FBC) indices. The aim of this study was to investigate whether the ColonFlag detects undiagnosed colorectal cancer prior to the recording of symptoms in general practice. Methods: We conducted case-control and cohort studies by linking primary care data from the Clinical Practice Research Datalink with colorectal cancer diagnoses from the National Cancer Registry. A ColonFlag score was derived for each FBC. We assessed the prevalence of symptoms at six-monthly intervals prior to index date (diagnosis date for cases, randomly selected date for controls). We then derived odds ratios (ORs) and area under the receiver operating characteristic (AUROC) curve for the ColonFlag, and for symptoms using logistic regression at each interval (primary outcome 18-24 months). Results: We included 1,893,641 patients, 10,875,556 FBCs and 8,918,037 ColonFlag scores. ColonFlag scores began to increase in cases compared with controls around 3-4 years before diagnosis. The AUROC for a diagnosis 18-24 months following the ColonFlag score was 0.736 (95% CI 0.715-0.759), falling to 0.536 (95% CI 0.523-0.548) with adjustment for age. ORs for individual symptoms became non-significant prior to 12 months before index date, except for abdominal pain (females OR=1.29, p<0.0001 at 12-18 months) and rectal bleeding (females OR=2.09, males OR=1.92, p<0.0001 at 18-24 months). Conclusions: Symptoms appear relatively late in the colorectal cancer process and are limited for supporting early stage detection. The ColonFlag can discriminate usefully at 18-24 months before diagnosis, suggesting a role for this algorithm in primary care, although some of its discriminatory ability comes from the age variable.
Background Simple blood tests can play an important role in identifying patients for cancer investigation. The current evidence base is limited almost entirely to tests used in isolation. However, recent evidence suggests combining multiple types of blood tests and investigating trends in blood test results over time could be more useful to select patients for further cancer investigation. Such trends could increase cancer yield and reduce unnecessary referrals. We aim to explore whether trends in blood test results are more useful than symptoms or single blood test results in selecting primary care patients for cancer investigation. We aim to develop clinical prediction models that incorporate trends in blood tests to identify the risk of cancer. Methods Primary care electronic health record data from the English Clinical Practice Research Datalink Aurum primary care database will be accessed and linked to cancer registrations and secondary care datasets. Using a cohort study design, we will describe patterns in blood testing (aim 1) and explore associations between covariates and trends in blood tests with cancer using mixed-effects, Cox, and dynamic models (aim 2). To build the predictive models for the risk of cancer, we will use dynamic risk modelling (such as multivariate joint modelling) and machine learning, incorporating simultaneous trends in multiple blood tests, together with other covariates (aim 3). Model performance will be assessed using various performance measures, including c-statistic and calibration plots. Discussion These models will form decision rules to help general practitioners find patients who need a referral for further investigation of cancer. This could increase cancer yield, reduce unnecessary referrals, and give more patients the opportunity for treatment and improved outcomes.
Background: Recent studies show that adults with intellectual disabilities (ID) have high incidence of major osteoporotic fracture, especially hip fracture. In those ≥ 50 years, women and men with ID have an approximately two and four times higher rate of hip fracture than women and men without ID. Increased awareness of osteoporotic fracture risk in ID may lead to wider use of antiresorptive drugs (bisphosphonates and denosumab) in this population. We aimed to compare, between people with and without ID, the incidence of 1) major side effects, namely medication related osteoporosis of the jaw (ONJ) and oesophagitis; 2) oral pathology, which can be a risk factor for ONJ. Methods: Exploratory study investigating safety of first line osteoporosis medication within the population of a previous study comparing fracture incidence in people with and without ID in the GOLD database of the Clinical Practice Research Datalink 1998–2017. Results: The percentage of people on antiresorptive drugs was identical in the ID and non ID group (1.4%). The number of individuals who developed ONJ and oesophagitis during the study was too low to allow an accurate estimate of incidence of the events and a comparison between the two groups. The incidence of any oral pathology was 119.31 vs 64.68/10000 person year in the ID vs non ID group. Conclusions: Medication related ONJ and oesophagitis are rare in people with and without ID. There is no reason based on our findings to use antiresorptives differently in people with ID as in the rest of the population. However, the potential for side effects of antiresorptives will inherently increase with wider use of these drugs. Given the higher incidence of oral pathology in people with ID, which could put them at higher risk of ONJ, precautions should be taken to prevent this complication by attention to oral health.
Background The full blood count (FBC) is a common blood test performed in general practice. It consists of many individual parameters that may change over time due to colorectal cancer. Such changes are likely missed in practice. We identified trends in these FBC parameters to facilitate early detection of colorectal cancer. Methods We performed a retrospective, case-control, longitudinal analysis of UK primary care patient data. LOWESS smoothing and mixed effects models were derived to compare trends in each FBC parameter between patients diagnosed and not diagnosed over a prior 10-year period. Results There were 399,405 males (2.3%, n = 9,255 diagnosed) and 540,544 females (1.5%, n = 8,153 diagnosed) in the study. There was no difference between cases and controls in FBC trends between 10 and four years before diagnosis. Within four years of diagnosis, trends in many FBC levels statistically significantly differed between cases and controls, including red blood cell count, haemoglobin, white blood cell count, and platelets (interaction between time and colorectal cancer presence: p <0.05). FBC trends were similar between Duke’s Stage A and D colorectal tumours, but started around one year earlier in Stage D diagnoses. Conclusions Trends in FBC parameters are different between patients with and without colorectal cancer for up to four years prior to diagnosis. Such trends could help earlier identification.
Background:Current osteoporosis guidelines do not identify individuals with intellectual disabilities (ID) as at risk of fracture, potentially missing opportunities for prevention. We aimed to assess the incidence of fractures in people with ID over the life course. Methods:Descriptive analysis of open cohort study using anonymised electronic health records from the UK Clinical Practice Research Datalink, linked to the Hospital Episode Statistics database (Jan 1, 1998-Dec 31, 2017). All individuals with ID were matched on age and sex to five individuals without ID. We calculated the incidence rate (95% CI) per 10000 person-years (py) and incidence rate ratio (IRR, 95% CI) to compare fractures between individuals with and without ID (age 1-17 and ≥18 years) for any fracture, and in those aged 18-49 and ≥ 50 years for major osteoporotic fracture (vertebra, shoulder, wrist, hip), and for hip fracture. Findings:43176 individuals with ID (15470 children aged 1-17 years; 27706 adults aged ≥ 18 years) were identified and included (40.4% females) along with 215733 matched control individuals. The median age at study entry was 24 (10th-90th centiles 3-54) years. Over a median (10th-90th centile) follow-up of 7.1 (0.9-17.6) and 6.5 (0.8-17.6) years, there were 5941 and 24363 incident fractures in the ID and non ID groups respectively. Incidence of any fracture was 143.5 (131.8-156.3) vs 120.7 (115.4-126.4)/10000 py (children), 174.2 (166.4-182.4)/10000 py vs 118.2 (115.3-121.2)/10000 py (adults) in females. In males it was 192.5 (182.4-203.2) vs 228.5 (223.0-234.1)/10000 py (children), 155.6 (149.3-162.1)/10000 py vs 128.4 (125.9-131.0)/10000 py (adults). IRR for major osteoporotic fracture was 1.81 (1.50-2.18) age 18-49 years, 1.69 (1.53-1.87) age ≥ 50 years in women. In men it was 1.56 (1.36-1.79) age 18-49 years, 2.45 (2.13-2.81) age ≥ 50 years. IRR for hip fracture was 7.79 (4.14-14.65) age 18-49 years, 2.28 (1.91-2.71) age ≥ 50 years in women. In men it was 6.04 (4.18-8.73) age 18-49 years, 3.91 (3.17-4.82) age ≥ 50 years. Comparable rates of major osteoporotic fracture and of hip fracture occurred approximately 15 and 20 years earlier respectively in women and 20 and 30 years earlier respectively in men with ID than without ID. Fracture distribution differed profoundly, hip fracture 9.9% vs 5.0% of any fracture in adults with ID vs without ID. Interpretation:The incidence, type, and distribution of fractures in people with intellectual disabilities suggest early onset osteoporosis. Prevention and management strategies are urgently required, particularly to reduce the incidence of hip fracture. Funding:National Institute for Health and Care Research.
Background The use of minimally and non-invasive monitoring systems (including continuous glucose monitoring) has increased rapidly over recent years. Up to now, it remains unclear how accurate devices can detect hypoglycaemic episodes. In this systematic review and meta-analysis, we assessed the diagnostic accuracy of minimally and non-invasive hypoglycaemia detection in comparison to capillary or venous blood glucose in patients with type 1 or type 2 diabetes. Methods Clinical Trials.gov, Cochrane Library, Embase, PubMed, ProQuest, Scopus and Web of Science were systematically searched. Two authors independently screened the articles, extracted data using a standardised extraction form and assessed methodological quality using a review-tailored quality assessment tool for diagnostic accuracy studies (QUADAS-2). The diagnostic accuracy of hypoglycaemia detection was analysed via meta-analysis using a bivariate random effects model and meta-regression with regard to pre-specified covariates. Results We identified 3416 nonduplicate articles. Finally, 15 studies with a total of 733 patients were included. Different thresholds for hypoglycaemia detection ranging from 40 to 100 mg/dl were used. Pooled analysis revealed a mean sensitivity of 69.3% [95% CI: 56.8 to 79.4] and a mean specificity of 93.3% [95% CI: 88.2 to 96.3]. Meta-regression analyses showed a better hypoglycaemia detection in studies indicating a higher overall accuracy, whereas year of publication did not significantly influence diagnostic accuracy. An additional analysis shows the absence of evidence for a better performance of the most recent generation of devices. Conclusion Overall, the present data suggest that minimally and non-invasive monitoring systems are not sufficiently accurate for detecting hypoglycaemia in routine use. Systematic review registration PROSPERO 2018 CRD42018104812
Objective Most current cardiovascular disease (CVD) risk stratification tools are for people without CVD, but very few are for prevalent CVD. In this study, we developed and validated a CVD severity score in people with coronary heart disease (CHD) and evaluated the association between severity and adverse outcomes. Methods Primary and secondary care data for 213 088 people with CHD in 398 practices in England between 2007 and 2017 were used. The cohort was randomly divided into training and validation datasets (80%/20%) for the severity model. Using 20 clinical severity indicators (each assigned a weight=1), baseline and longitudinal CVD severity scores were calculated as the sum of indicators. Adjusted Cox and competing-risk regression models were used to estimate risks for all-cause and cause-specific hospitalisation and mortality. Results Mean age was 64.5±12.7 years, 46% women, 16% from deprived areas, baseline severity score 1.5±1.2, with higher scores indicating a higher burden of disease. In the training dataset, 138 510 (81%) patients were hospitalised at least once, and 39 944 (23%) patients died. Each 1-unit increase in baseline severity was associated with 41% (95% CI 37% to 45%, area under the receiver operating characteristics (AUROC) curve=0.79) risk for 1 year for all-cause mortality; 59% (95% CI 52% to 67%, AUROC=0.80) for cardiovascular (CV)/diabetes mortality; 27% (95% CI 26% to 28%) for any-cause hospitalisation and 37% (95% CI 36% to 38%) for CV/diabetes hospitalisation. Findings were consistent in the validation dataset. Conclusions Higher CVD severity score is associated with higher risks for any-cause and cause-specific hospital admissions and mortality in people with CHD. Our reproducible score based on routinely collected data can help practitioners better prioritise management of people with CHD in primary care.
BackgroundThe complexity of general practice consultations may be increasing and varies in different settings. A measure of complexity is required to test these hypotheses.AimTo develop a valid measure of general practice consultation complexity applicable to routine medical records.Design and settingDelphi study to select potential indicators of complexity followed by a cross-sectional study in English general practices to develop and validate a complexity measure.MethodThe online Delphi study over two rounds identified potential indicators of consultation complexity. The cross-sectional study used an age–sex stratified random sample of patients and general practice face-to-face consultations from 2013/2014 in the Clinical Practice Research Datalink. The authors explored independent relationships between each indicator and consultation duration using mixed-effects regression models, and revalidated findings using data from 2017/2018. The proportion of complex consultations in different age–sex groups was assessed.ResultsA total of 32 GPs participated in the Delphi study. The Delphi panel endorsed 34 of 45 possible complexity indicators after two rounds. After excluding factors because of low prevalence or confounding, 17 indicators were retained in the cross-sectional study. The study used data from 173 130 patients and 725 616 face-to-face GP consultations. On defining complexity as the presence of any of these 17 factors, 308 370 consultations (42.5%) were found to be complex. Mean duration of complex consultations was 10.49 minutes, compared to 9.64 minutes for non-complex consultations. The proportion of complex consultations was similar in males and females but increased with age.ConclusionThe present consultation complexity measure has face and construct validity. It may be useful for research, management and policy, and for informing decisions about the range of resources needed in different practices.
The original article [1] contains an omitted grant acknowledgement and affiliation as relates to the contribution of co-author, Rafael Perera-Salazar. As such, the following two amendments should apply to the original article.
Objective Clinically applicable diabetes severity measures are lacking, with no previous studies comparing their predictive value with glycated hemoglobin (HbA 1c ). We developed and validated a type 2 diabetes severity score (the DIabetes Severity SCOre, DISSCO) and evaluated its association with risks of hospitalization and mortality, assessing its additional risk information to sociodemographic factors and HbA 1c . Research design and methods We used UK primary and secondary care data for 139 626 individuals with type 2 diabetes between 2007 and 2017, aged ≥35 years, and registered in general practices in England. The study cohort was randomly divided into a training cohort (n=111 748, 80%) to develop the severity tool and a validation cohort (n=27 878). We developed baseline and longitudinal severity scores using 34 diabetes-related domains. Cox regression models (adjusted for age, gender, ethnicity, deprivation, and HbA 1c ) were used for primary (all-cause mortality) and secondary (hospitalization due to any cause, diabetes, hypoglycemia, or cardiovascular disease or procedures) outcomes. Likelihood ratio (LR) tests were fitted to assess the significance of adding DISSCO to the sociodemographics and HbA 1c models. Results A total of 139 626 patients registered in 400 general practices, aged 63±12 years were included, 45% of whom were women, 83% were White, and 18% were from deprived areas. The mean baseline severity score was 1.3±2.0. Overall, 27 362 (20%) people died and 99 951 (72%) had ≥1 hospitalization. In the training cohort, a one-unit increase in baseline DISSCO was associated with higher hazard of mortality (HR: 1.14, 95% CI 1.13 to 1.15, area under the receiver operating characteristics curve (AUROC)=0.76) and cardiovascular hospitalization (HR: 1.45, 95% CI 1.43 to 1.46, AUROC=0.73). The LR tests showed that adding DISSCO to sociodemographic variables significantly improved the predictive value of survival models, outperforming the added value of HbA 1c for all outcomes. Findings were consistent in the validation cohort. Conclusions Higher levels of DISSCO are associated with higher risks for hospital admissions and mortality. The new severity score had higher predictive value than the proxy used in clinical practice, HbA 1c . This reproducible algorithm can help practitioners stratify clinical care of patients with type 2 diabetes.
Introduction: A full blood count (FBC) blood test includes 20 components. We systematically reviewed studies that assessed the association of the FBC and diagnosis of colorectal cancer to identify components as risk factors. We reviewed FBC-based prediction models for colorectal cancer risk. Methods: MEDLINE, EMBASE, CINAHL, and Web of Science were searched until 3 September 2019. We meta-analysed the mean difference in FBC components between those with and without a diagnosis and critically appraised the development and validation of FBC-based prediction models. Results: We included 53 eligible articles. Three of four meta-analysed components showed an association with diagnosis. In the remaining 16 with insufficient data for meta-analysis, three were associated with colorectal cancer. Thirteen FBC-based models were developed. Model performance was commonly assessed using the c-statistic (range 0.72–0.91) and calibration plots. Some models appeared to work well for early detection but good performance may be driven by early events. Conclusion: Red blood cells, haemoglobin, mean corpuscular volume, red blood cell distribution width, white blood cell count, and platelets are associated with diagnosis and could be used for referral. Existing FBC-based prediction models might not perform as well as expected and need further critical testing.
Background: The complexity of general practice consultations may be increasing and vary in different settings. Testing these hypotheses requires a measure of complexity. Aim: To develop a valid measure of general practice consultation complexity applicable to routine medical records. Design: Delphi study to select potential indicators of complexity followed by cross-sectional study to develop and validate a complexity measure. Setting: English general practices. Method: An online Delphi study over two rounds involved 32 general practitioners to identify potential indicators of consultation complexity. The cross-sectional study used an age-sex stratified random sample of 173,130 patients and 725,616 general practice face-to-face consultations from 2013/14 in the Clinical Practice Research Datalink. We explored independent relationships between each indicator and consultation duration using mixed effects regression models, and revalidated findings using data from 2017/18. We assessed the proportion of complex consultations in different age-sex groups. Results: After two rounds, the Delphi panel endorsed 34 of 45 possible complexity indicators. In the cross-sectional study, after excluding factors because of low prevalence or confounding, 17 indicators were retained. Defining complexity as the presence of any of these factors, 308,370 consultations (42.5%) were complex. Mean duration of complex consultations was 10.49 minutes, compared to 9.64 minutes for non-complex consultations. The proportion of complex consultations was similar in men and women but increased with age. Conclusion: Our consultation complexity measure has face and construct validity. It may be useful for research, management and policy, informing decisions about the range of resources needed in different practices. problems discussed within consultations. 16 Three studies have asked general practitioners about features that make patients complex, and we build on this by considering aspects of consultations as well as patients. 12,14,15,17 A few previous authors have devised case-mix measures applicable to primary care, but these have either not taken account of clinicians’ perceptions of the complexity of different factors 18-21 or not been designed for analysis of routine medical records. 13 22 There is some overlap 23 24 resource These measures are based information, other psychological factors 11 within consultations 12,15,17,26 and captured our complexity measure.
© 2020 The Author(s). The original article [1] contains an omitted grant acknowledgement and affiliation as relates to the contribution of co-author, Rafael Perera-Salazar. As such, the following two amendments should apply to the original article: 1) Rafael Perera-Salazar should also be affiliated to NIHR Oxford Biomedical Research Centre, Oxford University Hospitals NHS Foundation Trust (as shown in the affiliations of this Correction article). 2) The following Acknowledgement statement should take precedence over that in the original article.
A Full Blood Count (FBC) is a common blood test including 20 parameters, such as haemoglobin and platelets. FBCs from Electronic Health Record (EHR) databases provide a large sample of anonymised individual patient data and are increasingly used in research. We describe the quality of the FBC data in one EHR. The Test dataset from the Clinical Research Practice Datalink (CPRD) was accessed, which contains results of tests performed in primary care, such as FBC blood tests. Medical codes and entity codes, two coding systems used within CPRD to identify FBC records, were compared, with levels of mismatched coding, and number that could be rectified reported. The reliability of units of measurement are also described and missing data discussed. There were 14 entity codes and 138 medical codes for the FBC in the data. Medical and entity codes consistently corresponded to the same FBC parameter in 95.2% (n=217,752,448) of parameters. In the 4.8% (n=10,955,006) mismatches, the most common parameter rectified was mean platelet volume (n=2,041,360) and 1,191,540 could not be rectified and were removed. Units of measurement were often either missing, partially entered, or did not appear to correspond to the blood value. The final dataset contained 16,537,017 FBC tests. Applying mathematical equations to derive some missing parameters in these FBCs resulted in 15 of 20 parameters available per FBC on average, with 0.3% of FBCs having all 20 parameters. Performing data quality checks can help to understand the extent of any issues in the dataset. We emphasise balancing large sample sizes with reliability of the data.
MethodsStudy designA population-based propensity-score-matched cohort study to control for confounding at baseline has been conducted. Young adults (aged 16–35 years) who presented with a first-time TASD were selected from two computerised NHS databases (CPRD and HES). Figure 8 shows a detailed illustration of the study plan, which is described in more detail in the following paragraphs.
To the Editor: In a recent issue, Fernando et al.1 present a systematic review and meta-analysis assessing the prognostic accuracy of the HEART score for prediction of major adverse cardiac events (MACE) in adult patients presenting with chest pain at the emergency department (ED). The authors conclude that the HEART score has excellent performance for prediction of MACE (particularly mortality and myocardial infarction) in chest pain patients and should be the primary clinical decision instrument used for risk stratification of this patient population. While we agree with the utility of the HEART score, we take an alternate evidence-based medicine perspective on how to best demonstrate this utility to patients, clinicians, and other stakeholders. In this letter, we explore the concepts of prognosis, diagnosis, test accuracy, and risk estimation as they relate to the HEART score. Increasingly, clinicians and health researchers are recognizing the burden overdiagnosis can have on individual patients and health care systems. The usefulness of diagnostic decisions should be judged by whether patients classified with diagnosed disease do better than those classified without disease.2 This requires information about patient prognosis—in this case, the likelihood of future patient-centered outcomes in patients with a given HEART score. There is a case for prognosis to replace diagnosis as the framework for clinical decision making in patients presenting to the ED with chest pain. As the authors of this review acknowledge, there is a tendency for clinicians to overinvestigate chest pain patients, resulting in increased resource utilization without improved outcomes.3 In part, this is due to the lack of a good reference standard in the diagnosis of acute coronary syndrome (ACS). The diagnostic test accuracy literature for ACS in chest pain patients presenting to the ED most often uses clinical follow-up (e.g., future MACE) or a panel adjudicated final diagnosis based on clinical history and hospital course as the diagnostic reference standard.4 Both of these standards are flawed as they are subject to important incorporation, verification, and spectrum biases likely to misrepresent reported diagnostic test accuracy measures.5 As a result, ACS remains at its core a clinical diagnosis, with certain investigative tools (e.g., electrocardiogram, troponin) providing additional information supporting the diagnosis. At present, the Cochrane Collaboration has guidance on the performance of systematic reviews of diagnostic test accuracy,6 while a working group (Cochrane Prognosis Methods Group) attempts to reach consensus on how prognostic systematic reviews are best conducted.7 We highlight this because the authors of this review opted to perform a meta-analysis of diagnostic test accuracy rather than predictive ability. The rationale for this decision is that the HEART score is primarily used to “rule out” MACE in low-risk patients and so clinicians (and presumably patients) will be most interested in the accuracy of their screening decision. The authors state that when evaluating a decision instrument in the context of screening, the most important test characteristics are sensitivity, specificity, and likelihood ratios. We advocate that predictive values are the more clinically useful measures for this review's clinical question and should be routinely presented alongside other measures. Authors of prognostic reviews should also be challenged by peer reviewers to present their data in subgroups according to disease prevalence (e.g., low, intermediate, and high). The interested reader can then more reliably assess the performance of test characteristics across study populations with variable baseline risks of future MACE. Sensitivity and specificity are commonly taught to be fixed properties of a test that do not vary with disease prevalence. However, in clinical practice, the sensitivity and specificity of a test can vary with disease prevalence.8, 9 This phenomenon is known as spectrum or case-mix bias. Spectrum bias acknowledges the possibility that shifts in test performance may be at least in part due to case mix variation among study populations. Case mix variation can impact the sensitivity, specificity, and likelihood ratios of a given test. Among the studies included in this review, future MACE prevalence ranges from 1.1%10 to 29.4%.11 This is a wide prevalence range. It is reasonable to question the appropriateness of applying the reported pooled sensitivity (95.9%) for HEART score > 3 to such diverse populations. Assessing the external validity of the pooled sensitivity to variable populations is made more difficult by the fact that the review authors do not report a pooled prevalence of future MACE, though this can be calculated using data presented in Figure 2. To illustrate this concept further, the reported sensitivity in the low MACE prevalence study by Mahler et al.10 was 58%. Why is the sensitivity in this study so far from the pooled sensitivity of 95.9% reported in this review? The likely answer is spectrum bias. In fact, this American study involved patients in an observation unit who were already determined to be low risk by the treating clinician (on the basis of clinical assessment, non-diagnostic electrocardiogram and negative troponin).10 This population does not represent the chest pain patient we would typically be applying the HEART score to as observation units for low risk chest pain do not exist where we work. The study by Mahler et al.10 also presented predictive values, demonstrating that 0.6% (5/904) of patients with low-risk HEART scores suffered a future MACE. In this context, we believe that a negative predictive value of 99.4% is more helpful to the treating clinician in reaching a shared clinical decision than a sensitivity of 58%. When the HEART score is viewed as a test, there is also a risk that clinicians will attempt to “rule in” or “rule out” ACS using the HEART score. And can you blame them? After all, the population (ED patients with chest pain), index test (HEART score), reference standard (MACE), and target condition (“clinically significant” cardiac ischemia) alluded to in this review appear to provide the basic framework for an ACS diagnostic accuracy study. It is important to highlight that the HEART score was not designed to be applied to patients with a proven ACS. Patients with definite ACS at presentation are usually excluded from HEART score studies. These patients should be treated in the usual manner and referred for ongoing medical management and/or revascularization. The HEART score should be applied to those patients with chest pain, in which the diagnosis is uncertain, yet the clinician considers a cardiac etiology possible. Conceptually, we do not view the HEART score as a test. Many HEART score studies do not report measures of test accuracy, including those by authors involved with the HEART score's development.12 The HEART score is a set of clinical (history, risk factors, age) and investigative (electrocardiogram, troponin) criteria permitting estimation of a patient's future short-term risk of death, myocardial infarction, or revascularization. We liken it to the Wells’ score for pulmonary embolism, which is used to estimate a patient's short-term risk of pulmonary embolism. This pretest probability subsequently informs the decision to evaluate a patient with a D-dimer or diagnostic imaging.13 The clinical utility of the Wells’ score is not in ruling in or ruling out future thromboembolic events. Likewise, the clinical utility of the HEART score is not in ruling in or ruling out future MACE. Using terms like rule in or rule out encourages emergency clinicians to strive for perfection or no “misses” in the evaluation of these patients. These terms also perpetuate a view that for the HEART score to be valuable at the bedside, researchers must demonstrate the HEART score “outperforms” clinician gestalt. We need to accept that it is not possible to get ED patients with possible cardiac chest pain to a 30-day to 6-week MACE rate of zero, whether using the HEART score, clinician gestalt, or any other chest pain risk score. Rather than attempting to rule out future MACE, the clinician should ask the following question: is this patient's estimated risk modifiable by more observation and testing (e.g., serial troponin, cardiac stress test) over outpatient management of risk factors for coronary artery disease? We think Fernando and colleagues should be congratulated for their substantial work in this systematic review and meta-analysis of the HEART score. We appear to have the same intentions in mind but offer alternate views on how to evaluate and apply the HEART score literature.
BACKGROUND:Shoulder dislocations are the most common joint dislocations seen in emergency departments. Most traumatic cases are anterior and cause recurrent dislocations. Management options include surgical and conservative treatments. There is a lack of evidence about which method is most effective after the first traumatic anterior shoulder dislocation (TASD). OBJECTIVES:To produce UK age- and sex-specific incidence rates for TASD. To assess whether or not surgery within 6 months of a first-time TASD decreases re-dislocation rates compared with no surgery. To identify clinical predictors of recurrent dislocation. DESIGN:A population-based cohort study of first-time TASD patients in the UK. An initial validation study and subsequent propensity-score-matched analysis to compare re-dislocation rates between surgery and no surgery after a first-time TASD. Prediction modelling was used to identify potential predictors of recurrent dislocation. SETTING:UK primary and secondary care data. PARTICIPANTS:Patients with a first-time TASD between 1997 and 2015. INTERVENTIONS:Stabilisation surgery within 6 months of a first-time TASD (compared with no surgery). Stabilisation surgery within 12 months of a first-time TASD was also carried out as a sensitivity analysis. MAIN OUTCOME MEASURE:Re-dislocation rate up to 2 years after the first TASD. METHODS:Eligible patients were identified from the Clinical Practice Research Datalink (CPRD) (1997-2015). Accuracy of shoulder dislocation coding was internally validated using the CPRD General Practitioner questionnaire service. UK age- and sex-specific incidence rates for TASD were externally validated against rates from the USA and Canada. A propensity-score-matched analysis using linked CPRD and Hospital Episode Statistics (HES) data compared re-dislocation rates for patients aged 16-35 years, comparing surgery with no surgery. Multivariable Cox regression models for predicting re-dislocation were developed for the surgical and non-surgical cohorts. RESULTS:Shoulder dislocation was coded correctly for 89% of cases in the CPRD [95% confidence interval (CI) 83% to 95%], with a 'primary' dislocation confirmed for 76% of cases (95% CI 67% to 85%). Far fewer patients than expected received stabilisation surgery within 6 months of a first TASD, leading to an underpowered study. Around 20% of re-dislocation rates were observed for both surgical and non-surgical patients. The sensitivity analysis at 12 months also showed little difference in re-dislocation rates. Missing data on risk factors limited the value of the prediction modelling; however, younger age, epilepsy and sex (male) were identified as statistically significant predictors of re-dislocation. LIMITATIONS:Far fewer than the expected number of patients had surgery after a first-time TASD, resulting in an underpowered study. This and residual confounding from missing risk factors mean that it is not possible to draw valid conclusions. CONCLUSIONS:This study provides, for the first time, UK data on the age- and sex-specific incidence rates for TASD. Most TASD occurs in men, but an unexpected increased incidence was observed in women aged > 50 years. Surgery after a first-time TASD is uncommon in the NHS. Re-dislocation rates for patients receiving surgery after their first TASD are higher than previously expected; however, important residual confounding risk factors were not recorded in NHS primary and secondary care databases, thus preventing useful recommendations. FUTURE WORK:The high incidence of TASD justifies investigation into preventative measures for young men participating in contact sports, as well as investigating the risk factors in women aged > 50 years. A randomised controlled trial would account for key confounders missing from CPRD and HES data. A national TASD registry would allow for a more relevant data capture for this patient group. STUDY REGISTRATION:Independent Scientific Advisory Committee (ISAC) for the Medicines and Healthcare Products Regulatory Agency (ISAC protocol 15_0260). FUNDING:The National Institute for Health Research Health Technology Assessment programme.