Background:Ovarian cancer survival is stage-dependent: Stage I patients have 90% 5-year survival versus 15% for stage IV. Over 70% of patients worldwide are diagnosed at advanced stages. Ovarian cancer presents with non-specific symptoms (abdominal bloating, early satiety, discomfort/pain, bowel/urinary changes). Current National Institute for Health and Care Excellence guidelines recommend that symptomatic women presenting to primary care are tested with cancer antigen 125 and ultrasound, then referred to secondary care for further triage if these tests are abnormal. Current standard of care risk prediction model used to triage women in National Health Service secondary care is Risk of Malignancy Index 1 combining cancer antigen 125 and simple ultrasound features, which at 250 threshold has 70% sensitivity and 90% specificity. Newer models offer potential for improved sensitivity, earlier diagnosis and better survival outcomes. Objectives:To evaluate diagnostic strategies for ovarian cancer in women with non-specific symptoms through systematic review, United Kingdom Collaborative Trial of Ovarian Cancer Screening data set analysis, prospective studies and health economic evaluation comparing Risk of Malignancy Index 1 against newer approaches including Risk of Ovarian Malignancy Algorithm, Ovarian-Adnexal Reporting and Data System and International Ovarian Tumour Analysis models, including International Ovarian Tumour Analysis Assessment of Different NEoplasias in the adneXa. Methods:Four concurrent work packages: (1) Cochrane systematic review; (2) United Kingdom Collaborative Trial of Ovarian Cancer Screening data set model development; (3) prospective multicentre diagnostic accuracy study (ROCkeTS) with parallel pre/postmenopausal cohorts; and (4) cost-consequence analysis. Allied analyses investigated psychological impact and cancer outcomes from symptom-triggered pathways. ROCkeTS recruited 2453 women across 23 hospitals (2015-23) with symptoms, raised cancer antigen 125 and/or abnormal imaging. Women completed questionnaires, donated blood and underwent transvaginal ultrasound scored by International Ovarian Tumour Analysis terminology by certified National Health Service sonographers with quality assurance. Reference standard was histology for surgical cases or 12-month wellbeing ascertainment. Primary outcome: primary invasive ovarian cancer versus benign or normal. Results:The Cochrane systematic review (58 studies, 30,121 patients and 9061 ovarian cancer cases) demonstrated that most published diagnostic test accuracy studies failed to differentiate between pre- and postmenopausal women, and all were conducted in high-prevalence settings, limiting applicability to routine practice. In the ROCkeTS prospective study in premenopausal women, in the initial cohort recruited prior to protocol change (n = 857), Risk of Malignancy Index 1 at threshold 250 showed poor sensitivity (42.6%, 95% confidence interval 28.3 to 57.8) but high specificity (96.5%, 95% confidence interval 94.7 to 97.8). All other tests improved sensitivity but dropped specificity. International Ovarian Tumour Analysis Assessment of Different NEoplasias in the adneXa at 10% threshold achieved significantly higher sensitivity (89.1%, 95% confidence interval 76.4 to 96.4), higher than all other tests with acceptable specificity (73.2%, 95% confidence interval 69.9 to 76.4). In the ROCkeTS prospective cohort study in postmenopausal women (n = 1242), Risk of Malignancy Index 1 at 250 demonstrated better performance (82.9%, 95% confidence interval 76.7 to 88.0), but International Ovarian Tumour Analysis Assessment of Different NEoplasias in the adneXa at 10% had the best sensitivity at 96.1% (95% confidence interval 92.2 to 98.4) compared to Risk of Malignancy Index 1 with the least drop of specificity. Risk of Ovarian Malignancy Algorithm at manufacturer recommended threshold and Ovarian-Adnexal Reporting and Data System did not improve on Risk of Malignancy Index 1 sensitivity in postmenopausal women. Cancer prevalence differed between premenopausal (5.7%) and postmenopausal (17%) cohorts. Early-stage cancer (I/II) were diagnosed in 60.2% of premenopausal and 41% of postmenopausal cohorts. Cancer diagnosis rates were very low (1.6%) in women under 40 years. High anxiety and distress were noted, particularly in younger women. One in four women with high-grade serous ovarian cancers were diagnosed at early stage (I/II). Complete cytoreduction was achieved in 61.3% of cases, with optimal cytoreduction (≤ 1 cm residual disease) in an additional 15.1%. Cost-consequence analysis demonstrated that a two-step strategy deployed at the same ultrasound sitting, initially triaging out benign looking tumours on ultrasound, then calculating ovarian cancer risk with International Ovarian Tumour Analysis Assessment of Different NEoplasias in the adneXa ultrasound model at 10% demonstrated the best balance across cost, diagnostic yield and cancer deaths compared to other diagnostic strategies. Limitations:Cohort study required key changes to protocol and post-pandemic recruitment was slow. Conclusions:International Ovarian Tumour Analysis Assessment of Different NEoplasias in the adneXa ultrasound at 10% threshold, delivered by trained National Health Service sonographers demonstrated superior diagnostic performance compared to Risk of Malignancy Index 1 and should be considered as new standard of care for suspected ovarian cancer in pre- and postmenopausal women. A two-step strategy using International Ovarian Tumour Analysis Assessment of Different NEoplasias in the adneXa offers optimal balance across cost, diagnostic yield and cancer death reduction. Implementation requires sonographer training investment and quality assurance. Future research:International Ovarian Tumour Analysis Assessment of Different NEoplasias in the adneXa implementation in primary care/community settings, artificial intelligence-enabled quality assurance, reconfiguration of referral pathways in primary care to reduce unnecessary referrals in younger women and consequent harm are important research areas. Systematic symptom elicitation capitalising on routine health interactions to reach underserved communities warrants further research. Funding:This synopsis presents independent research funded by the National Institute for Health and Care Research (NIHR) Health Technology Assessment programme as award number 13/13/01.
OBJECTIVES:To review the evidence base, clinical performance claims, and usability and safety of self-tests available for sale on the UK high street. DESIGN:Cross sectional review of self-tests-regulation, evidence of performance, usability, and safety. SETTING:Tests were identified from supermarkets, pharmacies, and health and wellbeing shops within a 10 mile radius of the University of Birmingham Edgbaston Campus in 2023. MAIN OUTCOME MEASURES:Accuracy claims of self-tests, samples used to derive accuracy measures, and regulatory requirements were summarised. Ergonomics, usability and safety concerns about the equipment and instructions, including interpretability and readability, were evaluated. Details of clinical and lay person study reports (population, sample size, reference or comparator tests, test process) were summarised, and methods were assessed using the Quality Assessment of Diagnostic Studies 2 (QUADAS-2) tool. RESULTS:Thirty five self-tests were identified (30 obtained), which used seven different sample types and tested for 20 different biomarkers. Accuracy claims were made in instructions for use documents for 24/30 tests: accuracy for 19, sensitivity for 17, and specificity for 16. Performance claims of ≥98% were made on accuracy for 53% (10/19) of tests, 41% (7/17) on sensitivity, and 63% (10/16) on specificity. Where reference standards were reported in instructions for use documents, 29% (5/17) evaluated the accuracy of self-tests against similar rapid tests. For usability or safety, 18/30 self-tests had at least one high risk concern, 11 because of equipment, 10 because of the sampling process, and 15 owing to instructions or interpretation. Nine sets of clinical and lay person study reports were obtained (covering 12 tests). Across documents (nine clinical study reports and six lay person study reports) and QUADAS-2 domains, 73% were rated as having unclear risk of bias owing to poor reporting, and 58% were rated as having high applicability concerns because of inappropriate study designs. Participant descriptions were particularly inadequate in clinical study reports. Even within lay person study reports, few demographics (up to four) were presented. Some populations were unrepresentative of the intended user, inappropriate reference standards and thresholds were used, and mentions of blinding were scarce. CONCLUSIONS:This investigation highlights the need for improved regulatory oversight and clearer standards to ensure the safety and reliability of self-tests available on the UK market. Concerns about their ergonomics and usability might lead to test errors. Manufacturers' unwillingness to provide public access to study documents raises ethical concerns. Additionally, inadequate study design and reporting in available documentation hinders the ability to assess the evidence base supporting the use of self-tests. As the availability and use of self-tests continues to rise, improved regulatory oversight is urgently needed to protect the public from the effects of poor performing diagnostic self-tests.
BACKGROUND:Postpartum haemorrhage (excessive bleeding after birth) is a leading cause of maternal mortality and morbidity worldwide. However, there is no global consensus on which clinical markers best define excessive bleeding or reliably predict adverse maternal outcomes. The aim of this study was to assess the prognostic accuracy of clinical markers of postpartum bleeding in predicting maternal mortality or severe morbidity. METHODS:In this individual participant data meta-analysis, eligible datasets were identified through a global call for data issued by WHO and systematic searches of PubMed, MEDLINE, Embase, the Cochrane Library, and WHO trial registries (from database inception to Nov 6, 2024). Studies were eligible if they included at least 200 participants with objectively measured blood loss or other clinical markers of haemodynamic instability, and reported at least one clinical outcome of interest. Individual participant data were requested for all eligible studies. For each dataset, we computed the prognostic accuracy of each clinical marker to predict a composite outcome of maternal mortality or severe morbidity (blood transfusion, surgical interventions, or admission to intensive care unit). Five clinical markers were assessed: measured blood loss, pulse rate, systolic blood pressure, diastolic blood pressure, and shock index. Results were meta-analysed through two-level mixed-effects logistic regression models, with a bivariate normal model used to generate summary accuracy estimates. Clinical marker and threshold selections were informed by a WHO expert consensus process, which placed emphasis on maximising prognostic sensitivity (preferably >80%) over prognostic specificity (preferably ≥50%). This meta-analysis was registered on PROSPERO (CRD420251034918). FINDINGS:We identified 33 potentially eligible datasets and successfully obtained and analysed full data for 12 datasets, comprising 312 151 women. At the conventional threshold of 500 mL, measured blood loss had a summary prognostic sensitivity of 75·7% (95% CI 60·3-86·4) and specificity of 81·4% (95% CI 70·7-88·8) for predicting the composite outcome. The preferred sensitivity threshold was reached at 300 mL (83·9% [95% CI 72·8-91·1]), although at the expense of reduced specificity (54·8% [95% CI 38·0-70·5]). Prognostic performance improved with a decision rule that combined the use of either blood loss thresholds less than 500 mL (≥300 mL to ≥450 mL) and any abnormal haemodynamic sign (pulse rate >100 beats per min, systolic blood pressure <100 mm Hg, diastolic blood pressure <60 mm Hg, or shock index >1·0) or 500 mL or more of blood loss, with sensitivities ranging from 86·9% to 87·9% and specificities from 66·6% to 76·1%. INTERPRETATION:Measured blood loss below the conventional threshold, combined with abnormal haemodynamic signs, accurately predicts women at risk of death or life-threatening complications from postpartum bleeding and could support earlier postpartum haemorrhage diagnosis and treatment. FUNDING:The Gates Foundation and UNDP/UNFPA/UNICEF/WHO/World Bank Special Programme of Research, Development and Research Training in Human Reproduction.
Objectives Minor salivary gland (MSG) biopsy has an important role in Sjögren's disease diagnosis and research. MSGs show within-patient variation in number of lymphocytic foci per unit area, but the optimal number of MSGs required to balance reproducibility and clinical acceptability has not been determined. Methods Monte Carlo simulations were performed to investigate impact of MSG number on (i) diagnosis based on focus score (FS) ≥1; (ii) reproducibility, defined as the extent to which 2 FS measurements obtained from 2 within-patient biopsies are the same, assuming no systematic differences have occurred in between biopsies; and (iii) smallest sample size required to detect a clinically meaningful difference in FS. Data simulation was repeated for different MSG numbers (range, 2-7). Results Higher reproducibility was noted for every unit increase in MSG number, with the median absolute difference between 2 within-patient FS measurements decreasing from 1.05 (SD = 0.25) with 2 glands to 0.52 (SD = 0.12) with 7 glands. MSG number influenced the probability of a simulated patient receiving a FS ≥1, increasing from a median of 0.67 with 2 glands to 0.77 with ≥5 glands. MSG number influenced clinical trial sample sizes. For example, 80% statistical power to detect a 40% FS reduction required a sample size per group of 62 with 2 glands and 25 with 7 glands. Conclusions For a diagnostic threshold of FS ≥1, a minimum of 5 glands should ideally be targeted. For continuous FS values, a larger number of MSGs (eg, 6) will increase reproducibility further and reduce clinical trial sample size requirements.
OBJECTIVE:To assess the costs and consequences of seven diagnostic strategies for ovarian cancer in pre- and post-menopausal women with symptoms in secondary care. DESIGN:Economic evaluation alongside a prospective single-arm diagnostic accuracy study. SETTING:NHS secondary care outpatients (2-week referrals, clinics, GP referrals, cross-specialty referrals) and inpatients (emergency presentations to secondary care). SAMPLE:Two cohorts of 857 pre-menopausal and 1242 post-menopausal women newly presenting to secondary care with symptoms of suspected ovarian cancer. METHODS:A model-based cost-consequence analysis (CCA) was conducted using a decision tree simulating patient pathways over 12 months. Diagnostic accuracy data were sourced from the ROCkeTS study and supplemented by literature. MAIN OUTCOME MEASURES:Cancer deaths, correct diagnosis proportion, and diagnostic yield. RESULTS:No diagnostic strategy was optimal across all outcomes. Across both cohorts, the Risk of Malignancy Index (RMI) 200 was least expensive but had poor cancer death and diagnostic yield outcomes. The ADNEX 3% strategy had the highest diagnostic yield and lowest cancer mortality but was the most expensive. For pre-menopausal women, the IOTA ADNEX 10% strategy outperformed ORADS, ROMA, and CA125 in cost and outcomes. For post-menopausal women, the high cancer prevalence required a trade-off. In sensitivity analysis, a two-step IOTA ADNEX 10% strategy outperformed ORADS, ROMA, and CA125 across all three outcomes, making the strategy a more balanced choice in both cohorts. CONCLUSION:At 12 months, no single diagnostic strategy was superior. Early diagnosis requires balancing cancer mortality, diagnostic yield, and cost. The IOTA ADNEX two-step strategy at a 10% threshold provided the best trade-off across these factors and is recommended for practice.
OBJECTIVE:Symptom-triggered testing for ovarian cancer was introduced to the UK whereby symptomatic women undergo an ultrasound scan and serum CA125, and are referred to hospital within 2 weeks if these are abnormal. The potential value of symptom-triggered testing in the detection of early-stage disease or low tumor burden remains unclear in women with high grade serous ovarian cancer. In this descriptive study, we report on the International Federation of Gynecology and Obstetrics (FIGO) stage, disease distribution, and complete cytoreduction rates in women presenting via the fast-track pathway and who were diagnosed with high grade serous ovarian cancer. METHODS:We analyzed the dataset from Refining Ovarian Cancer Test accuracy Scores (ROCkeTS), a single-arm prospective diagnostic test accuracy study recruiting from 24 hospitals in the UK. The aim of ROCkeTS is to validate risk prediction models in symptomatic women. We undertook an opportunistic analysis for women recruited between June 2015 to July 2022 and who were diagnosed with high grade serous ovarian cancer via the fast-track pathway. Women presenting with symptoms suspicious for ovarian cancer receive a CA125 blood test and an ultrasound scan if the CA125 level is abnormal. If either of these is abnormal, women are referred to secondary care within 2 weeks. Histology details were available on all women who underwent surgery or biopsy within 3 months of recruitment. Women who did not undergo surgery or biopsy at 3 months were followed up for 12 months as per the national guidelines in the UK. In this descriptive study, we report on patient demographics (age and menopausal status), WHO performance status, FIGO stage at diagnosis, disease distribution (low/pelvic confined, moderate/ extending to mid-abdomen, high/extending to upper abdomen) and complete cytoreduction rates in women who underwent surgery. RESULTS:Of 1741 participants recruited via the fast-track pathway, 119 (6.8%) were diagnosed with high grade serous ovarian cancer. The median age was 63 years (range 32-89). Of these, 112 (94.1%) patients had a performance status of 0 and 1, 30 (25.2%) were diagnosed with stages I/II, and the disease distribution was low-to-moderate in 77 (64.7%). Complete and optimal cytoreduction were achieved in 73 (61.3%) and 18 (15.1%). The extent of disease was low in 43 of 119 (36.1%), moderate in 34 of 119 (28.6%), high in 32 of 119 (26.9%), and not available in 10 of 119 (8.4%). Nearly two thirds, that is 78 of 119 (65.5%) women with high grade serous ovarian cancer, underwent primary debulking surgery, 36 of 119 (30.3%) received neoadjuvant chemotherapy followed by interval debulking surgery, and 5 of 119 (4.2%) women did not undergo surgery. CONCLUSION:Our results demonstrate that one in four women identified with high grade serous ovarian cancer through the fast-track pathway following symptom-triggered testing was diagnosed with early-stage disease. Symptom-triggered testing may help identify women with a low disease burden, potentially contributing to high complete cytoreduction rates.
Objectives To review the information provided for self-test devices sold in high street shops in the UK and to assess their suitability for informed decision making based on use, interpretation, and post-test actions. Design Cross sectional review of information on self-test boxes and instructions for use leaflets. Setting Supermarkets, pharmacies, and health and wellbeing shops within a 10 mile radius of the University of Birmingham’s campus at Edgbaston in 2023. Main outcome measures Information on intended use of test, biomarker and clinical condition, interpretation of test results, recommendations for post-test actions, and coherence of intended use and post-test recommendations with evidence based guidance. Results 30 self-tests assessing 20 biomarkers for 19 different conditions were included. Information to guide purchase was present on a few boxes: who should use the test and when (8/30, 27%), action after the test result (7/30, 23%), and numerical test performance (10/30, 33%). From the information provided either on the box or within the instructions for use leaflets, 21 (70%) self-tests were judged to be used for diagnosis and 15 (50%) to be used for screening, although 3/21 (14%) did not provide any information about symptoms and 10/15 (67%) did not provide any information about risk factors to guide use. 27 (90%) self-tests recommended follow-up with a healthcare professional if results were positive or abnormal, and 14 (47%) if test results were negative or normal. Use of tests for 11 of 19 (58%) conditions was judged contrary to evidence based guidance in one or more of the intended population, frequency of testing, test threshold, or investigative approach required for a condition. Conclusions The current market for self-tests does not support consumer informed decisions about their use, interpretation of test results, and subsequent actions. Clinicians working downstream of self-tests are likely to face important challenges in incorporating the results in practice. As the use of self-tests continues to increase, improved regulatory oversight is urgently needed to protect the public and healthcare systems from misuse.
Objective To describe recommendations applicable to new diagnostic and screening tests brought to market in the UK as of 01 June 2023; and extract agreements, disagreements and gaps.Design Extant regulations, recommendations and guidelines for new diagnostic and screening tests applicable to new products placed in the UK market as of 01 June 2023. Non-English and references not applicable to new tests seeking market access in the UK on 01 June 2023 were excluded.Data sources PubMed, Web of Science Core Collection and Scopus; grey literature via EuropePMC and Google, government regulations and guidelines, and relevant professional societies. References of relevant included data were scanned for includable articles. Resultant data were thematically analysed and presented as a narrative scoping review.Eligibility for selecting studies PubMed, Web of Science Core Collection and Scopus were searched from 1/1/2018 for regulations, guidelines or recommendations for in vitro diagnositic tests as applied to the UK on 1/6/23. Relevant papers also had references searched.Results 943 items were initially identified with 892 excluded. Reference searching located a further 31 papers and 82 items were analysed. Seven themes were identified: regulation, companion diagnostics and lab developed tests, safety and evidence, test specific recommendations, data, innovation and recommendations for patients/the public. Wide agreement included the need to reduce bureaucracy and duplication; to mitigate to avoid unintended consequences of In Vitro Diagnostic Regulation. Disagreement over whether high-quality evidence should precede regulatory approval, or could be gathered as part of postmarketing surveillance emerged.Conclusions Industry, regulators, academics, patients representing a variety of views, should collaborate to work through areas of disagreement.
Introduction/Background ROCkeTS investigated accuracy of risk prediction models for diagnosing ovarian cancer (OC); no previous studies investigate all tests as head-to-head comparisons. Methodology Recruitment: ROCkeTS recruited newly presenting women with non-specific symptoms and raised CA125 and/or abnormal imaging to donate blood and undergo ultrasound scan performed mainly by sonographers. Sonographers achieved certification in IOTA models pre-participation and underwent quality assurance. Index tests: IOTA ADNeX model 3% and 10% thresholds, RMI1 at 200, ROMA at >14.4%, >25.3%, >27.7%, > 29.9% (manufacturer recommended threshold) and CA125 at 35 IU/ml. Comparator: RMI1 at 250 threshold. Reference standard: Tissue biopsy/cytology or follow-up at 12 month Primary outcome: accuracy defined as primary invasive OC versus benign or normal Secondary outcomes: accuracy defined as primary invasive OC, secondary malignant, borderline and neoplasms of uncertain behaviour versus benign or normal Study sample size required 150 OCs to detect 13% sensitivity improvement from 70% to 84% with 90% power, assuming positive correlation of test errors. Sensitivity, specificity, c-index (area under Receiver operating characteristic (ROC) curve), PPV and Negative Predictive value (NPV) were reported with calibration plots. Results 1242 postmenopausal women recruited from 23 hospitals. For primary outcome, comparing RMI 1 250, sensitivity 82.9% (95% CI: 76.7 to 88.0), specificity 87.4% (95% CI: 84.9 to 89.6), IOTA ADNeX 10% was more sensitive 96.1% (95% CI: 92.2 to 98.4) and less specific 58.5% (95% CI: 54.7 TO 62.1), p< 0.001, whilst sensitivity of ROMA at 29.9% was comparable 88.0% (95% CI: 82.5 to 92.2) with lower specificity 79.9% (95% CI: 76.9 to 82.7), p=0.0001. Analysis of secondary outcome was similar. (table 1). Conclusion Across all analyses, IOTA ADNeX had highest sensitivity but lower specificity. ROMA at 29.9 had marginal improvement of sensitivity over RMI 250 but reduction in specificity. IOTA ADNEX at 10% should be standard of care diagnostic in OC for postmenopausal women. Disclosures Davenport C, Rai N, Sharma P, Deeks JJ, Berhane S, Mallett S, Saha P, Champaneria R, Bayliss SE, Snell KIE, Sundar S. Menopausal status, ultrasound and biomarker tests in combination for the diagnosis of ovarian cancer in symptomatic women. Cochrane Database of Systematic Reviews 2022, Issue 7. Art. No.: CD011964. DOI: 10.1002/14651858.CD011964.pub2. Accessed 30 November 2023.
The coronavirus disease (Covid-19) pandemic raised challenges for everyday life. Development of new diagnostic tests was necessary, but under such enormous pressure risking inadequate evaluation. Against a background of concern about standards applied to the evaluation of in vitro diagnostic tests (IVDs), clear statistical thinking was needed on the principles of diagnostic testing in general, and their application in a pandemic. Therefore, in July 2020, the Royal Statistical Society convened a Working Group of six biostatisticians to review the statistical evidence needed to ensure the performance of new tests, especially IVDs for infectious diseases—for regulators, decision-makers, and the public. The Working Group’s review was undertaken when the Covid-19 pandemic shone an unforgiving light on current processes for evaluating and regulating IVDs for infectious diseases. The report’s findings apply more broadly than to the pandemic and IVDs, to diagnostic test evaluations in general. A section of the report focussed on lessons learned during the pandemic and aimed to contribute to the UK Covid-19 Inquiry’s examination of the response to, and impact of, the Covid-19 pandemic to learn lessons for the future. The review made 22 recommendations on what matters for study design, transparency, and regulation.
INTRODUCTION:ElaTION is a large multicenter pragmatic randomized controlled trial, performed in 18 secondary/tertiary hospitals across England, comparing elastography ultrasound-guided fine needle aspiration cytology (EUS-FNAC) with ultrasound-guided FNAC (US-FNAC) alone in the diagnostic assessment of thyroid nodules. Secondary trial outcomes, reported here, assessed the accuracy of ultrasound alone (US) compared with US-FNAC to inform and update current practice guidelines. METHODS:Adults with single or multiple thyroid nodules who had not undergone previous FNAC were eligible. Radiologists assessed all thyroid nodules using US alone, thereby enabling assessment of its accuracy (sensitivity and specificity) vs US-FNAC. RESULTS:Of the 982 participants, a final definitive diagnosis was obtained in 688, who were included in the final analyses. The sensitivity of US alone was the same as US-FNAC (0.91 [95% CI, 0.85-0.97] vs 0.87 [95% CI, 0.80-0.95] P = .37). US alone had statistically significant lower specificity than US-FNAC alone (0.48 vs 0.67 respectively, P < .0001). The malignancy rate on histology in a nodule classified as benign on ultrasound (U2) was 9/263 (3.42%) and on cytology (Thy2) was 15/353 (4.25%), whereas the malignancy rate in a nodule that was benign on both (U2, Thy2) was 3/210 (1.43%). Malignancy risk for U3, U4, and U5 nodules was 68/304 (22.4%), 43/83 (51.8%), and 29/38 (76.3%), respectively (P < .0001). Yet 80/982 (8%) patients were discharged despite having U3-U5 scans with Thy1 (nondiagnostic) FNAC and no definitive diagnosis.Malignancy risk was higher in smaller nodules: < 10 mm 23/60 (38.3%), 10-20 mm 46/162 (28.4%), and >20 mm 80/466 (17.2%) (P < .0001). Nodules with indeterminate cytology with atypical features (Thy3a) carried a similar malignancy risk to those with indeterminate cytology (Thy3/3f): 27/95 (28.4%) vs 42/113 (37.2%) respectively (P = .18). CONCLUSION:Ultrasound alone appears to be an effective diagnostic modality in thyroid nodules, confirming the recommendations of recent guidelines and the British Thyroid Association classification. However, findings also suggest caution regarding existing recommendations for conservative management of nondiagnostic (Thy1/Bethesda I) and atypical (Thy3a/Bethesda III) nodules. In those cases, ultrasound (U3-U5) features may help identify high-risk subgroups for more proactive management.
BACKGROUND:Sample collection is a key driver of accuracy in the diagnosis of SARS-CoV-2 infection. Viral load may vary at different anatomical sampling sites and accuracy may be compromised by difficulties obtaining specimens and the expertise of the person taking the sample. It is important to optimise sampling accuracy within cost, safety and accessibility constraints. OBJECTIVES:To compare the sensitivity of different sampling collection sites and methods for the detection of current SARS-CoV-2 infection with any molecular or antigen-based test. SEARCH METHODS:Electronic searches of the Cochrane COVID-19 Study Register and the COVID-19 Living Evidence Database from the University of Bern (which includes daily updates from PubMed and Embase and preprints from medRxiv and bioRxiv) were undertaken on 22 February 2022. We included independent evaluations from national reference laboratories, FIND and the Diagnostics Global Health website. We did not apply language restrictions. SELECTION CRITERIA:We included studies of symptomatic or asymptomatic people with suspected SARS-CoV-2 infection undergoing testing. We included studies of any design that compared results from different sample types (anatomical location, operator, collection device) collected from the same participant within a 24-hour period. DATA COLLECTION AND ANALYSIS:Within a sample pair, we defined a reference sample and an index sample collected from the same participant within the same clinical encounter (within 24 hours). Where the sample comparison was different anatomical sites, the reference standard was defined as a nasopharyngeal or combined naso/oropharyngeal sample collected into the same sample container and the index sample as the alternative anatomical site. Where the sample comparison was concerned with differences in the sample collection method from the same site, we defined the reference sample as that closest to standard practice for that sample type. Where the sample pair comparison was concerned with differences in personnel collecting the sample, the more skilled or experienced operator was considered the reference sample. Two review authors independently assessed the risk of bias and applicability concerns using the QUADAS-2 and QUADAS-C checklists, tailored to this review. We present estimates of the difference in the sensitivity (reference sample (%) minus index sample sensitivity (%)) in a pair and as an average across studies for each index sampling method using forest plots and tables. We examined heterogeneity between studies according to population (age, symptom status) and index sample (time post-symptom onset, operator expertise, use of transport medium) characteristics. MAIN RESULTS:This review includes 106 studies reporting 154 evaluations and 60,523 sample pair comparisons, of which 11,045 had SARS-CoV-2 infection. Ninety evaluations were of saliva samples, 37 nasal, seven oropharyngeal, six gargle, six oral and four combined nasal/oropharyngeal samples. Four evaluations were of the effect of operator expertise on the accuracy of three different sample types. The majority of included evaluations (146) used molecular tests, of which 140 used RT-PCR (reverse transcription polymerase chain reaction). Eight evaluations were of nasal samples used with Ag-RDTs (rapid antigen tests). The majority of studies were conducted in Europe (35/106, 33%) or the USA (27%) and conducted in dedicated COVID-19 testing clinics or in ambulatory hospital settings (53%). Targeted screening or contact tracing accounted for only 4% of evaluations. Where reported, the majority of evaluations were of adults (91/154, 59%), 28 (18%) were in mixed populations with only seven (4%) in children. The median prevalence of confirmed SARS-CoV-2 was 23% (interquartile (IQR) 13%-40%). Risk of bias and applicability assessment were hampered by poor reporting in 77% and 65% of included studies, respectively. Risk of bias was low across all domains in only 3% of evaluations due to inappropriate inclusion or exclusion criteria, unclear recruitment, lack of blinding, nonrandomised sampling order or differences in testing kit within a sample pair. Sixty-eight percent of evaluation cohorts were judged as being at high or unclear applicability concern either due to inflation of the prevalence of SARS-CoV-2 infection in study populations by selectively including individuals with confirmed PCR-positive samples or because there was insufficient detail to allow replication of sample collection. When used with RT-PCR • There was no evidence of a difference in sensitivity between gargle and nasopharyngeal samples (on average -1 percentage points, 95% CI -5 to +2, based on 6 evaluations, 2138 sample pairs, of which 389 had SARS-CoV-2). • There was no evidence of a difference in sensitivity between saliva collection from the deep throat and nasopharyngeal samples (on average +10 percentage points, 95% CI -1 to +21, based on 2192 sample pairs, of which 730 had SARS-CoV-2). • There was evidence that saliva collection using spitting, drooling or salivating was on average -12 percentage points less sensitive (95% CI -16 to -8, based on 27,253 sample pairs, of which 4636 had SARS-CoV-2) compared to nasopharyngeal samples. We did not find any evidence of a difference in the sensitivity of saliva collected using spitting, drooling or salivating (sensitivity difference: range from -13 percentage points (spit) to -21 percentage points (salivate)). • Nasal samples (anterior and mid-turbinate collection combined) were, on average, 12 percentage points less sensitive compared to nasopharyngeal samples (95% CI -17 to -7), based on 9291 sample pairs, of which 1485 had SARS-CoV-2. We did not find any evidence of a difference in sensitivity between nasal samples collected from the mid-turbinates (3942 sample pairs) or from the anterior nares (8272 sample pairs). • There was evidence that oropharyngeal samples were, on average, 17 percentage points less sensitive than nasopharyngeal samples (95% CI -29 to -5), based on seven evaluations, 2522 sample pairs, of which 511 had SARS-CoV-2. A much smaller volume of evidence was available for combined nasal/oropharyngeal samples and oral samples. Age, symptom status and use of transport media do not appear to affect the sensitivity of saliva samples and nasal samples. When used with Ag-RDTs • There was no evidence of a difference in sensitivity between nasal samples compared to nasopharyngeal samples (sensitivity, on average, 0 percentage points -0.2 to +0.2, based on 3688 sample pairs, of which 535 had SARS-CoV-2). AUTHORS' CONCLUSIONS:When used with RT-PCR, there is no evidence for a difference in sensitivity of self-collected gargle or deep-throat saliva samples compared to nasopharyngeal samples collected by healthcare workers when used with RT-PCR. Use of these alternative, self-collected sample types has the potential to reduce cost and discomfort and improve the safety of sampling by reducing risk of transmission from aerosol spread which occurs as a result of coughing and gagging during the nasopharyngeal or oropharyngeal sample collection procedure. This may, in turn, improve access to and uptake of testing. Other types of saliva, nasal, oral and oropharyngeal samples are, on average, less sensitive compared to healthcare worker-collected nasopharyngeal samples, and it is unlikely that sensitivities of this magnitude would be acceptable for confirmation of SARS-CoV-2 infection with RT-PCR. When used with Ag-RDTs, there is no evidence of a difference in sensitivity between nasal samples and healthcare worker-collected nasopharyngeal samples for detecting SARS-CoV-2. The implications of this for self-testing are unclear as evaluations did not report whether nasal samples were self-collected or collected by healthcare workers. Further research is needed in asymptomatic individuals, children and in Ag-RDTs, and to investigate the effect of operator expertise on accuracy. Quality assessment of the evidence base underpinning these conclusions was restricted by poor reporting. There is a need for further high-quality studies, adhering to reporting standards for test accuracy studies.
BACKGROUND:Burn damage to skin often results in scarring; however in some individuals the failure of normal wound-healing processes results in excessive scar tissue formation, termed 'hypertrophic scarring'. The most commonly used method for the prevention and treatment of hypertrophic scarring is pressure-garment therapy (PGT). PGT is considered standard care globally; however, there is continued uncertainty around its effectiveness. OBJECTIVES:To evaluate the benefits and harms of pressure-garment therapy for the prevention of hypertrophic scarring after burn injury. SEARCH METHODS:We used standard, extensive Cochrane search methods. We searched CENTRAL, MEDLINE, Embase, two other databases, and two trials registers on 8 June 2023 with reference checking, citation searching, and contact with study authors to identify additional studies. SELECTION CRITERIA:We included randomised controlled trials (RCTs) comparing PGT (alone or in combination with other scar-management therapies) with scar management therapies not including PGT, or comparing different PGT pressures or different types of PGT. DATA COLLECTION AND ANALYSIS:At least two review authors independently selected trials for inclusion using predetermined inclusion criteria, extracted data, and assessed risk of bias using the Cochrane RoB 1 tool. We assessed the certainty of evidence using GRADE. MAIN RESULTS:We included 15 studies in this review (1179 participants), 14 of which (1057 participants) presented useable data. The sample size of included studies ranged from 17 to 159 participants. Most studies included both adults and children. Eight studies compared a pressure garment (with or without another scar management therapy) with scar management therapy alone, five studies compared the same pressure garment at a higher pressure versus a lower pressure, and two studies compared two different types of pressure garments. Studies used a variety of pressure garments (e.g. in-house manufactured or a commercial brand). Types of scar management therapies included were lanolin massage, topical silicone gel, silicone sheet/dressing, and heparin sodium ointment. Meta-analysis was not possible as there was significant clinical and methodological heterogeneity between studies. Main outcome measures were scar improvement assessed using the Vancouver Scar Scale (VSS) or the Patient and Observer Scar Assessment Scale (POSAS) (or both), pain, pruritus, quality of life, adverse events, and adherence to therapy. Studies additionally reported a further 14 outcomes, mostly individual scar parameters, some of which contributed to global scores on the VSS or POSAS. The amount of evidence for each individual outcome was limited. Most studies had a short follow-up, which may have affected results as the full effect of any therapy on scar healing may not be seen until around 18 months. PGT versus no treatment/lanolin We included five studies (378 participants). The evidence is very uncertain on whether PGT improves scars as assessed by the VSS compared with no treatment/lanolin. The evidence is also very uncertain for pain, pruritus, adverse events, and adherence. No study used the POSAS or assessed quality of life. One additional study (122 participants) did not report useable data. PGT versus silicone We included three studies (359 participants). The evidence is very uncertain on the effect of PGT compared with silicone, as assessed by the VSS and POSAS. The evidence is also very uncertain for pain, pruritus, quality of life, adverse events, adherence, and other scar parameters. It is possible that silicone may result in fewer adverse events or better adherence compared with PGT but this was also based on very low-certainty evidence. PGT plus silicone versus no treatment/lanolin We included two studies (200 participants). The evidence is very uncertain on whether PGT plus silicone improves scars as assessed by the VSS compared with no treatment/lanolin. The evidence is also very uncertain for pain, pruritus, and adverse events. No study used the POSAS or assessed quality of life or adherence. PGT plus silicone versus silicone We included three studies (359 participants). The evidence is very uncertain on the effect of PGT plus silicone compared with silicone, as assessed by the VSS and POSAS. The evidence is also very uncertain for pain, pruritus, quality of life, adverse events, and adherence. PGT plus scar management therapy including silicone versus scar management therapy including silicone We included one study (88 participants). The evidence is very uncertain on the effect of PGT plus scar management therapy including silicone versus scar management therapy including silicone, as assessed by the VSS and POSAS. The evidence is also very uncertain for pain, pruritus, quality of life, adverse events, and adherence. High-pressure versus low-pressure garments We included five studies (262 participants). The evidence is very uncertain on the effect of high pressure versus low pressure PGT on adverse events and adherence. No study used the VSS or the POSAS or assessed pain, pruritus, or quality of life. Different types of PGT (Caroskin Tricot + an adhesive silicone gel sheet versus Gecko Nanoplast (silicone gel bandage)) We included one study (60 participants). The evidence is very uncertain on the effect of Caroskin Tricot versus Gecko Nanoplast on the POSAS, pain, pruritus, and adverse events. The study did not use the VSS or assess quality of life or adherence. Different types of pressure garments (Jobst versus Tubigrip) We included one study (110 participants). The evidence is very uncertain on the adherence to either Jobst or Tubigrip. This study did not report any other outcomes. AUTHORS' CONCLUSIONS:There is insufficient evidence to recommend using either PGT or an alternative for preventing hypertrophic scarring after burn injury. PGT is already commonly used in practice and it is possible that continuing to do so may provide some benefit to some people. However, until more evidence becomes available, it may be appropriate to allow patient preference to guide therapy.
Background:Strain and shear wave elastography which is commonly used with concurrent real-time imaging known as real-time ultrasound shear/strain wave elastography is a new diagnostic technique that has been reported to be useful in the diagnosis of nodules in several organs. There is conflicting evidence regarding its benefit over ultrasound-guided fine-needle aspiration cytology alone in thyroid nodules. Objectives:To determine if ultrasound strain and shear wave elastography in conjunction with fine-needle aspiration cytology will reduce the number of patients who have a non-diagnostic first fine-needle aspiration cytology results as compared to conventional ultrasound-only guided fine-needle aspiration cytology. Design:A pragmatic, unblinded, multicentre randomised controlled trial. Setting:Eighteen centres with a radiology department across England. Participants:Adults who had not undergone previous fine-needle aspiration cytology with single or multiple nodules undergoing investigation. Interventions:Ultrasound shear/strain wave elastography-ultrasound guided fine-needle aspiration cytology (intervention arm) - strain or shear wave elastography-guided fine-needle aspiration cytology. Ultrasound-only guided fine-needle aspiration cytology (control arm) - routine ultrasound-only guided fine-needle aspiration cytology (the current standard recommended by the British Thyroid Association guidelines). Main outcome measure:The proportion of patients who have a non-diagnostic cytology (Thy 1) result following the first fine-needle aspiration cytology. Randomisation:Patients were randomised at a 1 : 1 ratio to the interventional or control arms. Results:A total of 982 participants (80% female) were randomised: 493 were randomised to ultrasound shear/strain wave elastography-ultrasound guided fine-needle aspiration cytology and 489 were randomised to ultrasound-only guided fine-needle aspiration cytology. There was no evidence of a difference between ultrasound shear/strain wave elastography and ultrasound in non-diagnostic cytology (Thy 1) rate following the first fine-needle aspiration cytology (19% vs. 16% respectively; risk difference: 0.030; 95% confidence interval -0.007 to 0.066; p = 0.11), the number of fine-needle aspiration cytologies needed (odds ratio: 1.10; 95% confidence interval 0.82 to 1.49; p = 0.53) or in the time to reach a definitive diagnosis (hazard ratio: 0.94; 95% confidence interval 0.81 to 1.10; p = 0.45). There was a small, non-significant reduction in the number of thyroid operations undertaken when ultrasound shear/strain wave elastography was used (37% vs. 40% respectively; risk difference: -0.02; 95% confidence interval -0.06 to 0.009; p = 0.15), but no difference in the number of operations yielding benign histology - 23% versus 24% respectively, p = 0.70 (i.e. no increase in identification of malignant cases) - or in the number of serious adverse events (2% vs. 1%). There was no difference in anxiety and depression, pain or quality of life between the two arms. Limitations:The study was not powered to detect differences in malignancy. Conclusions:Ultrasound shear/strain wave elastography does not appear to have additional benefit over ultrasound-guided fine-needle aspiration cytology in the diagnosis of thyroid nodules. Future work:The findings of the ElaTION trial suggest that further research into the use of shear wave elastography in the diagnostic setting of thyroid nodules is unlikely to be warranted unless there are improvements in the technology. The diagnostic difficulty in distinguishing between benign and malignant lesions still persists. Future studies might examine the role of genomic testing on fine-needle aspiration samples. There is growing use of targeted panels of molecular markers, particularly aimed at improving the diagnostic accuracy of indeterminate (i.e. Thy3) cytology results. The application of these tests is not uniform, and their cost effectiveness has not been assessed in large-scale trials. Study registration:This study is registered as ISRCTN (ISRCTN18261857). Funding:This award was funded by the National Institute for Health and Care Research (NIHR) Health Technology Assessment programme (NIHR award ref: 12/19/04) and is published in full in Health Technology Assessment; Vol. 28, No. 46. See the NIHR Funding and Awards website for further award information.
Background There is variable evidence and no randomized trials on the benefit of US elastography-guided fine-needle aspiration cytology (FNAC) over conventional US-guided FNAC alone for thyroid nodules. Purpose To compare the efficacy of US elastography-guided FNAC versus US-guided FNAC in reducing nondiagnostic rates for thyroid nodules. Materials and Methods A pragmatic, multicenter randomized controlled trial was performed at 18 secondary and tertiary hospitals across England between February 2015 and September 2018. Eligible adults with single or multiple thyroid nodules who had not previously undergone FNAC were randomized (1:1 ratio) to US elastography FNAC (intervention) or conventional US FNAC (control). The primary outcome was the proportion of patients who have a nondiagnostic cytologic Thy1 (British Thyroid Association system) result following the first FNAC. Results A total of 982 participants (mean age, 51.3 years ± 15 [SD] [IQR, 39-63]; male-to-female ratio, 1:4) were randomized. Of the 493 participants who underwent US elastography, 467 (94.7%) were examined with strain US elastography. There was no difference between the two arms in the nondiagnostic (Thy1) rate following the first FNAC (19% vs 16%; risk difference [RD], 0.03 [95% CI: -0.01, 0.07]; P = .11) or in the median time to reach the final definitive diagnosis (3.3 months [IQR, 1.5-6.4] for US elastography FNAC vs 3.4 months [IQR, 1.5-6.2] for US FNAC). All sensitivity analyses supported the primary analysis. Fewer participants in the US elastography FNAC arm underwent diagnostic hemithyroidectomy than in the US FNAC arm (183 of 493 [37%] vs 196 of 489 [40%]), but this was not statistically significant (adjusted RD, 0.02 [95% CI: -0.06, 0.01]; P = 0.15). There was no evidence of a difference in malignancy rates between the two arms: 70 of 493 (14%) in US elastography FNAC arm versus 79 of 489 (16%) in US FNAC arm (P = .39). There was also no difference in the rate of benign histologic findings between the groups (RD, -0.01 [95% CI: -0.04, 0.03]; P = .7). Conclusion Strain US elastography does not appear to have additional benefit over conventional US FNAC in the diagnosis of malignancy in thyroid nodules. Clinical trial registration no. ISRCTN18261857 Published under a CC BY 4.0 license. Supplemental material is available for this article. See also the editorial by Isikbay and Harwin in this issue.
Objectives: To investigate psychological correlates in women referred with suspected ovarian cancer via the fast-track pathway, explore how anxiety and distress levels change 12 months post-testing and report cancer conversion rates by age and referral pathway. Design: Single arm prospective cohort study Setting: Multicentre. Secondary care including outpatient clinics and emergency admissions. Participants: 2596 newly presenting symptomatic women with a raised CA125 level, abnormal imaging or both. Methods: Women completed anxiety and distress questionnaires at recruitment and at 12 months for those who had not undergone surgery or a biopsy within 3 months of recruitment. Main outcome measures: Anxiety and distress levels measured using STAI-6 and IES-r questionnaires. OC conversion rates by age, menopausal status and referral pathway. Results: 1355/2596 (52.1%) and 1781/2596 (68.6%) experienced moderate-to-severe distress and anxiety at recruitment. Younger age and emergency presentations had higher distress levels. Clinical category for anxiety and distress remained unchanged/worsened in 76% at 12 months despite a non-cancer diagnosis. OC rates by age were 1.6% (95% CI 0.5 to 5.9) under 40 and 10.9 % (95% CI 8.7 to 13.6) over 40 years. In women referred through fast-track pathways, 3.3% (95% CI 1.9 to 5.7) of pre- and 18.5% (95% CI 16.1 to 21.0) of postmenopausal women were diagnosed with OC. Conclusions: Women undergoing diagnostic testing display severe anxiety and distress. Younger women are especially vulnerable and should be targeted for support. Women under 40 have low conversion rates and we advocate reducing testing in this group to reduce harms of testing.
Introduction The diagnosis of neovascular age-related macular degeneration (nAMD), the leading cause of visual impairment in the developed world, relies on the interpretation of various imaging tests of the retina. These include invasive angiographic methods, such as Fundus Fluorescein Angiography (FFA) and, on occasion, Indocyanine-Green Angiography (ICGA). Newer, non-invasive imaging modalities, predominately Optical Coherence Tomography (OCT) and Optical Coherence Tomography Angiography (OCTA), have drastically transformed the diagnostic approach to nAMD. The aim of this study is to undertake a comprehensive diagnostic accuracy assessment of the various imaging modalities used in clinical practice for the diagnosis of nAMD (OCT, OCTA, FFA and, when a variant of nAMD called Polypoidal Choroidal Vasculopathy is suspected, ICGA) both alone and in various combinations.Methods and analysis This is a non-inferiority, prospective, randomised diagnostic accuracy study of 1067 participants. Participants are patients with clinical features consistent with nAMD who present to a National Health Service secondary care ophthalmology unit in the UK. Patients will undergo OCT as per standard practice and those with suspicious features of nAMD on OCT will be approached for participation in the study. Patients who agree to take part will also undergo both OCTA and FFA (and ICGA if indicated). Interpretation of the imaging tests will be undertaken by clinicians at recruitment sites. A randomised design was selected to avoid bias from consecutive review of all imaging tests by the same clinician. The primary outcome of the study will be the difference in sensitivity and specificity between OCT+OCTA and OCT+FFA (±ICGA) for nAMD detection as interpreted by clinicians at recruitment sites.Ethics and dissemination The study has been approved by the South Central—Oxford B Research Ethics Committee with reference number 21/SC/0412.Dissemination of study results will involve peer-review publications, presentations at major national and international scientific conferences.Trial registration number ISRCTN18313457.