OBJECTIVE:To search for and critically appraise the psychometric quality of patient-reported outcome measures (PROMs) developed or validated in optic neuritis, in order to support high-quality research and care.METHODS:We systematically searched MEDLINE(Ovid), Embase(Ovid), PsycINFO(Ovid) and CINAHLPlus(EBSCO), and additional grey literature to November 2021, to identify PROM development or validation studies applicable to optic neuritis associated with any systemic or neurologic disease in adults. We included instruments developed using classic test theory or Rasch analysis approaches. We used established quality criteria to assess content development, validity, reliability, and responsiveness, grading multiple domains from A (high quality) to C (low quality).RESULTS:From 3142 screened abstracts we identified five PROM instruments potentially applicable to optic neuritis: three differing versions of the National Eye Institute (NEI)-Visual Function Questionnaire (VFQ): the 51-item VFQ; the 25-item VFQ and a 10-item neuro-ophthalmology supplement; and the Impact of Visual Impairment Scale (IVIS), a constituent of the Multiple Sclerosis Quality of Life Inventory (MSQLI) handbook, derived from the Functional Assessment of Multiple Sclerosis (FAMS). Psychometric appraisal revealed the NEI-VFQ-51 and 10-item neuro module had some relevant content development but weak psychometric development, and the FAMS had stronger psychometric development using Rasch Analysis, but was only somewhat relevant to optic neuritis. We identified no content or psychometric development for IVIS.CONCLUSION:There is unmet need for a PROM with strong content and psychometric development applicable to optic neuritis for use in virtual care pathways and clinical trials to support drug marketing authorisation.
Patient reported outcome measures (PROMs) capture impact of disease and treatment on quality of life, and have an emerging role in clinical trial outcome measurement. This study included a systematic review and quality appraisal of PROMs developed or validated for use in adults with uveitis or scleritis. We searched MEDLINE, EMBASE, PsycINFO, CINAHL and grey literature sources, to 5 November 2021. We used established quality criteria to grade each PROM instrument in multiple domains from A (high quality) to C (low quality), and assessed content development, validity, reliability and responsiveness. For instruments developed using classic test theory-based psychometric approaches, we assessed acceptability, item targeting and internal consistency. For instruments developed using Item Response Theory (IRT) (e.g. Rasch analysis), we assessed response categories, dimensionality, measurement precision, item fit statistics, differential item functioning and targeting. We identified and appraised four instruments applicable to certain uveitis types, but none for scleritis. Specifically, the National Eye Institute Visual Function Questionnaire-25 (NEI-VFQ), a 3-part PROM for Birdshot retinochoroiditis (Birdshot Disease & Medication Symptoms Questionnaire [BD&MSQ], the quality of life (QoL) impact of Birdshot Chorioretinopathy [QoL BCR], and the QoL impact of BCR medication [QoL Meds], the Kings Sarcoidosis Questionnaire (KSQ), and a PROM for cytomegalovirus retinitis. These instruments had limited coverage for these heterogeneous conditions, with a focus on very rare subtypes. Psychometric appraisal revealed considerable variability between instruments, limited content development, and only one developed using Item Response Theory. In conclusion, there are few validated PROMs for patients with uveitis and none for scleritis, and existing instruments have suboptimal psychometric performance. We articulate why we do not recommend their inclusion as clinical trial outcome measures for drug licensing purposes, and highlight an unmet need for PROMs applicable to uveitis and scleritis.
Background: Ovarian cancer (OC) is a diagnostic challenge, with the majority diagnosed at late stages. Existing systematic reviews of diagnostic models either use inappropriate meta-analytic methods or do not conduct statistical comparisons of models or stratify test performance by menopausal status. Methods: We searched CENTRAL, MEDLINE, EMBASE, CINAHL, CDSR, DARE, Health Technology Assessment Database and SCI Science Citation Index, trials registers, conference proceedings from 1991 to June 2019. Cochrane collaboration review methods included QUADAS-2 quality assessment and meta-analysis using hierarchical modelling. RMI, ROMA or ADNEX at any test positivity threshold were investigated. Histology or clinical follow-up was the reference standard. We excluded screening studies, studies restricted to pregnancy, recurrent or metastatic OC. 2 × 2 diagnostic tables were extracted separately for pre- and post-menopausal women. Results: We included 58 studies (30,121 patients, 9061 cases of ovarian cancer). Prevalence of OC ranged from 16 to 55% in studies. For premenopausal women, ROMA at a threshold of 13.1 (+/−2) and ADNEX at a threshold of 10% demonstrated significantly higher sensitivity compared to RMI I at 200 (p < 0.0001) 77.8 (72.5, 82.4), 94.9 (92.5, 96.6), and 57.1% (50.6 to 63.4) but lower specificity (p < 0.002), 92.5 (90.0, 94.4), 84.3 (81.3, 86.8), and 78.2 (75.8, 80.4). For postmenopausal women, ROMA at a threshold of 27.7 (+/−2) and AdNEX at a threshold of 10% demonstrated significantly higher sensitivity compared to RMI I at a threshold of 200 (p < 0.001) 90.4 (87.4, 92.7), 97.6 (96.2, 98.5), and 78.7 (74.3, 82.5), specificity of ROMA was comparable, whilst ADneX was lower, 85.5 (81.3, 88.9), 81.3 (76.9, 85.0) (p = 0.155), compared to RMI 55.2 (51.2, 59.1) (p < 0.001). Conclusions: In pre-menopausal women, ROMA and ADNEX offer significantly higher sensitivity but significantly decreased specificity. In post-menopausal women, ROMA demonstrates significantly higher sensitivity and comparable specificity to RMI I, ADNEX has the highest sensitivity of all models, but with significantly reduced specificity. RMI I has poor sensitivity compared to ROMA or ADNEX. Choice between ROMA and ADNEX as a replacement test will depend on cost effectiveness and resource implications.
BACKGROUND Ovarian cancer (OC) has the highest case fatality rate of all gynaecological cancers. Diagnostic delays are caused by non-specific symptoms. Existing systematic reviews have not comprehensively covered tests in current practice, not estimated accuracy separately in pre- and postmenopausal women, or used inappropriate meta-analytic methods. OBJECTIVES To establish the accuracy of combinations of menopausal status, ultrasound scan (USS) and biomarkers for the diagnosis of ovarian cancer in pre- and postmenopausal women and compare the accuracy of different test combinations. SEARCH METHODS We searched CENTRAL, MEDLINE (Ovid), Embase (Ovid), five other databases and three trial registries from 1991 to 2015 and MEDLINE (Ovid) and Embase (Ovid) form June 2015 to June 2019. We also searched conference proceedings from the European Society of Gynaecological Oncology, International Gynecologic Cancer Society, American Society of Clinical Oncology and Society of Gynecologic Oncology, ZETOC and Conference Proceedings Citation Index (Web of Knowledge). We searched reference lists of included studies and published systematic reviews. SELECTION CRITERIA We included cross-sectional diagnostic test accuracy studies evaluating single tests or comparing two or more tests, randomised trials comparing two or more tests, and studies validating multivariable models for the diagnosis of OC investigating test combinations, compared with a reference standard of histological confirmation or clinical follow-up in women with a pelvic mass (detected clinically or through USS) suspicious for OC. DATA COLLECTION AND ANALYSIS Two review authors independently extracted data and assessed quality using QUADAS-2. We used the bivariate hierarchical model to indirectly compare tests at commonly reported thresholds in pre- and postmenopausal women separately. We indirectly compared tests across all thresholds and estimated sensitivity at fixed specificities of 80% and 90% by fitting hierarchical summary receiver operating characteristic (HSROC) models in pre- and postmenopausal women separately. MAIN RESULTS We included 59 studies (32,059 women, 9545 cases of OC). Two tests evaluated the accuracy of a combination of menopausal status and USS findings (IOTA Logistic Regression Model 2 (LR2) and the Assessment of Different NEoplasias in the adneXa model (ADNEX)); one test evaluated the accuracy of a combination of menopausal status, USS findings and serum biomarker CA125 (Risk of Malignancy Index (RMI)); and one test evaluated the accuracy of a combination of menopausal status and two serum biomarkers (CA125 and HE4) (Risk of Ovarian Malignancy Algorithm (ROMA)). Most studies were at high or unclear risk of bias in participant, reference standard, and flow and timing domains. All studies were in hospital settings. Prevalence was 16% (RMI, ROMA), 22% (LR2) and 27% (ADNEX) in premenopausal women and 38% (RMI), 45% (ROMA), 52% (LR2) and 55% (ADNEX) in postmenopausal women. The prevalence of OC in the studies was considerably higher than would be expected in symptomatic women presenting in community-based settings, or in women referred from the community to hospital with a suspicion of OC. Studies were at high or unclear applicability because presenting features were not reported, or USS was performed by experienced ultrasonographers for RMI, LR2 and ADNEX. The higher sensitivity and lower specificity observed in postmenopausal compared to premenopausal women across all index tests and at all thresholds may reflect highly selected patient cohorts in the included studies. In premenopausal women, ROMA at a threshold of 13.1 (± 2), LR2 at a threshold to achieve a post-test probability of OC of 10% and ADNEX (post-test probability 10%) demonstrated a higher sensitivity (ROMA: 77.4%, 95% CI 72.7% to 81.5%; LR2: 83.3%, 95% CI 74.7% to 89.5%; ADNEX: 95.5%, 95% CI 91.0% to 97.8%) compared to RMI (57.2%, 95% CI 50.3% to 63.8%). The specificity of ROMA and ADNEX were lower in premenopausal women (ROMA: 84.3%, 95% CI 81.2% to 87.0%; ADNEX: 77.8%, 95% CI 67.4% to 85.5%) compared to RMI 92.5% (95% CI 90.3% to 94.2%). The specificity of LR2 was comparable to RMI (90.4%, 95% CI 84.6% to 94.1%). In postmenopausal women, ROMA at a threshold of 27.7 (± 2), LR2 (post-test probability 10%) and ADNEX (post-test probability 10%) demonstrated a higher sensitivity (ROMA: 90.3%, 95% CI 87.5% to 92.6%; LR2: 94.8%, 95% CI 92.3% to 96.6%; ADNEX: 97.6%, 95% CI 95.6% to 98.7%) compared to RMI (78.4%, 95% CI 74.6% to 81.7%). Specificity of ROMA at a threshold of 27.7 (± 2) (81.5, 95% CI 76.5% to 85.5%) was comparable to RMI (85.4%, 95% CI 82.0% to 88.2%), whereas for LR2 (post-test probability 10%) and ADNEX (post-test probability 10%) specificity was lower (LR2: 60.6%, 95% CI 50.5% to 69.9%; ADNEX: 55.0%, 95% CI 42.8% to 66.6%). AUTHORS' CONCLUSIONS In specialist healthcare settings in both premenopausal and postmenopausal women, RMI has poor sensitivity. In premenopausal women, ROMA, LR2 and ADNEX offer better sensitivity (fewer missed cancers), but for ROMA and ADNEX this is off-set by a decrease in specificity and increase in false positives. In postmenopausal women, ROMA demonstrates a higher sensitivity and comparable specificity to RMI. ADNEX has the highest sensitivity in postmenopausal women, but reduced specificity. The prevalence of OC in included studies is representative of a highly selected referred population, rather than a population in whom referral is being considered. The comparative accuracy of tests observed here may not be transferable to non-specialist settings. Ultimately health systems need to balance accuracy and resource implications to identify the most suitable test.
Aims Antenatal pelvic floor muscle training (PFMT) may be effective for the prevention and treatment of urinary and fecal incontinence both in pregnancy and postnatally, but it is not routinely implemented in practice despite guideline recommendations. This review synthesizes evidence that exposes challenges, opportunities, and concerns regarding the implementation of PFMT during the childbearing years, from the perspective of individuals, healthcare professionals (HCPs), and organizations. Methods Critical interpretive synthesis of systematically identified primary quantitative or qualitative studies or research syntheses of women's and HCPs attitudes, beliefs, or experiences of implementing PFMT. Results Fifty sources were included. These focused on experiences of postnatal urinary incontinence (UI) and perspectives of individual postnatal women, with limited evidence exploring the views of antenatal women and HCP or wider organizational and environmental issues. The concept of agency (people's ability to effect change through their interaction with other people, processes, and systems) provides an over-arching explanation of how PFMT can be implemented during childbearing years. This requires both individual and collective action of women, HCPs, maternity services and organizations, funders and policymakers. Conclusion Numerous factors constrain women's and HCPs capacity to implement PFMT. It is unrealistic to expect women and HCPs to implement PFMT without reforming policy and service delivery. The implementation of PFMT during pregnancy, as recommended by antenatal care and UI management guidelines, requires policymakers, organizations, HCPs, and women to value the prevention of incontinence throughout women's lives by using low-risk, low-cost, and proven strategies as part of women's reproductive health.
Objective To examine the validity and findings of studies that examine the accuracy of algorithm based smartphone applications (“apps”) to assess risk of skin cancer in suspicious skin lesions. Design Systematic review of diagnostic accuracy studies. Studies of any design that evaluated algorithm based smartphone apps to assess images of skin lesions suspicious for skin cancer. Reference standards included histological diagnosis or follow-up, and expert recommendation for further investigation or intervention. Two authors independently extracted data and assessed validity using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2 tool). Estimates of sensitivity and specificity were reported for each app. Nine studies that evaluated six different identifiable smartphone apps were included. Six verified results by using histology or follow-up (n=725 lesions), and three verified results by using expert recommendations (n=407 lesions). Studies were small and of poor methodological quality, with selective recruitment, high rates of unevaluable images, and differential verification. Lesion selection and image acquisition were performed by clinicians rather than smartphone users. Two CE (Conformit Europenne) marked apps are available for download. SkinScan was evaluated in a single study (n=15, five melanomas) with 0% sensitivity and 100% specificity for the detection of melanoma. SkinVision was evaluated in two studies (n=252, 61 malignant or premalignant lesions) and achieved a sensitivity of 80%
Aims We assessed the performance of modelsf (risk scores) for predicting recurrence of atrial fibrillation (AF) in patients who have undergone catheter ablation. Methods and results Systematic searches of bibliographic databases were conducted (November 2018). Studies were eligible for inclusion if they reported the development, validation, or impact assessment of a model for predicting AF recurrence after ablation. Model performance (discrimination and calibration) measures were extracted. The Prediction Study Risk of Bias Assessment Tool (PROBAST) was used to assess risk of bias. Meta-analysis was not feasible due to clinical and methodological differences between studies, but c-statistics were presented in forest plots. Thirty-three studies developing or validating 13 models were included; eight studies compared two or more models. Common model variables were left atrial parameters, type of AF, and age. Model discriminatory ability was highly variable and no model had consistently poor or good performance. Most studies did not assess model calibration. The main risk of bias concern was the lack of internal validation which may have resulted in overly optimistic and/or biased model performance estimates. No model impact studies were identified. Conclusion Our systematic review suggests that clinical risk prediction of AF after ablation has potential, but there remains a need for robust evaluation of risk factors and development of risk scores.
Abstract Objective To examine the validity and findings of studies that examine the accuracy of algorithm based smartphone applications (“apps”) to assess risk of skin cancer in suspicious skin lesions. Design Systematic review of diagnostic accuracy studies. Data sources Cochrane Central Register of Controlled Trials, MEDLINE, Embase, CINAHL, CPCI, Zetoc, Science Citation Index, and online trial registers (from database inception to 10 April 2019). Eligibility criteria for selecting studies Studies of any design that evaluated algorithm based smartphone apps to assess images of skin lesions suspicious for skin cancer. Reference standards included histological diagnosis or follow-up, and expert recommendation for further investigation or intervention. Two authors independently extracted data and assessed validity using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2 tool). Estimates of sensitivity and specificity were reported for each app. Results Nine studies that evaluated six different identifiable smartphone apps were included. Six verified results by using histology or follow-up (n=725 lesions), and three verified results by using expert recommendations (n=407 lesions). Studies were small and of poor methodological quality, with selective recruitment, high rates of unevaluable images, and differential verification. Lesion selection and image acquisition were performed by clinicians rather than smartphone users. Two CE (Conformit Europenne) marked apps are available for download. No published peer reviewed study was found evaluating the TeleSkin skinScan app. SkinVision was evaluated in three studies (n=267, 66 malignant or premalignant lesions) and achieved a sensitivity of 80% (95% confidence interval 63% to 92%) and a specificity of 78% (67% to 87%) for the detection of malignant or premalignant lesions. Accuracy of the SkinVision app verified against expert recommendations was poor (three studies). Conclusions Current algorithm based smartphone apps cannot be relied on to detect all cases of melanoma or other skin cancers. Test performance is likely to be poorer than reported here when used in clinically relevant populations and by the intended users of the apps. The current regulatory process for awarding the CE marking for algorithm based apps does not provide adequate protection to the public. Systematic review registration PROSPERO CRD42016033595.
BACKGROUND Melanoma is one of the most aggressive forms of skin cancer, with the potential to metastasise to other parts of the body via the lymphatic system and the bloodstream. Melanoma accounts for a small percentage of skin cancer cases but is responsible for the majority of skin cancer deaths. Various imaging tests can be used with the aim of detecting metastatic spread of disease following a primary diagnosis of melanoma (primary staging) or on clinical suspicion of disease recurrence (re-staging). Accurate staging is crucial to ensuring that patients are directed to the most appropriate and effective treatment at different points on the clinical pathway. Establishing the comparative accuracy of ultrasound, computed tomography (CT), magnetic resonance imaging (MRI), and positron emission tomography (PET)-CT imaging for detection of nodal or distant metastases, or both, is critical to understanding if, how, and where on the pathway these tests might be used. OBJECTIVES Primary objectivesWe estimated accuracy separately according to the point in the clinical pathway at which imaging tests were used. Our objectives were:• to determine the diagnostic accuracy of ultrasound or PET-CT for detection of nodal metastases before sentinel lymph node biopsy in adults with confirmed cutaneous invasive melanoma; and• to determine the diagnostic accuracy of ultrasound, CT, MRI, or PET-CT for whole body imaging in adults with cutaneous invasive melanoma:○ for detection of any metastasis in adults with a primary diagnosis of melanoma (i.e. primary staging at presentation); and○ for detection of any metastasis in adults undergoing staging of recurrence of melanoma (i.e. re-staging prompted by findings on routine follow-up).We undertook separate analyses according to whether accuracy data were reported per patient or per lesion.Secondary objectivesWe sought to determine the diagnostic accuracy of ultrasound, CT, MRI, or PET-CT for whole body imaging (detection of any metastasis) in mixed or not clearly described populations of adults with cutaneous invasive melanoma.For study participants undergoing primary staging or re-staging (for possible recurrence), and for mixed or unclear populations, our objectives were:• to determine the diagnostic accuracy of ultrasound, CT, MRI, or PET-CT for detection of nodal metastases;• to determine the diagnostic accuracy of ultrasound, CT, MRI, or PET-CT for detection of distant metastases; and• to determine the diagnostic accuracy of ultrasound, CT, MRI, or PET-CT for detection of distant metastases according to metastatic site. SEARCH METHODS We undertook a comprehensive search of the following databases from inception up to August 2016: Cochrane Central Register of Controlled Trials; MEDLINE; Embase; CINAHL; CPCI; Zetoc; Science Citation Index; US National Institutes of Health Ongoing Trials Register; NIHR Clinical Research Network Portfolio Database; and the World Health Organization International Clinical Trials Registry Platform. We studied reference lists as well as published systematic review articles. SELECTION CRITERIA We included studies of any design that evaluated ultrasound (with or without the use of fine needle aspiration cytology (FNAC)), CT, MRI, or PET-CT for staging of cutaneous melanoma in adults, compared with a reference standard of histological confirmation or imaging with clinical follow-up of at least three months' duration. We excluded studies reporting multiple applications of the same test in more than 10% of study participants. DATA COLLECTION AND ANALYSIS Two review authors independently extracted all data using a standardised data extraction and quality assessment form (based on the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2)). We estimated accuracy using the bivariate hierarchical method to produce summary sensitivities and specificities with 95% confidence and prediction regions. We undertook analysis of studies allowing direct and indirect comparison between tests. We examined heterogeneity between studies by visually inspecting the forest plots of sensitivity and specificity and summary receiver operating characteristic (ROC) plots. Numbers of identified studies were insufficient to allow formal investigation of potential sources of heterogeneity. MAIN RESULTS We included a total of 39 publications reporting on 5204 study participants; 34 studies reporting data per patient included 4980 study participants with 1265 cases of metastatic disease, and seven studies reporting data per lesion included 417 study participants with 1846 potentially metastatic lesions, 1061 of which were confirmed metastases. The risk of bias was low or unclear for all domains apart from participant flow. Concerns regarding applicability of the evidence were high or unclear for almost all domains. Participant selection from mixed or not clearly defined populations and poorly described application and interpretation of index tests were particularly problematic.The accuracy of imaging for detection of regional nodal metastases before sentinel lymph node biopsy (SLNB) was evaluated in 18 studies. In 11 studies (2614 participants; 542 cases), the summary sensitivity of ultrasound alone was 35.4% (95% confidence interval (CI) 17.0% to 59.4%) and specificity was 93.9% (95% CI 86.1% to 97.5%). Combining pre-SLNB ultrasound with FNAC revealed summary sensitivity of 18.0% (95% CI 3.58% to 56.5%) and specificity of 99.8% (95% CI 99.1% to 99.9%) (1164 participants; 259 cases). Four studies demonstrated lower sensitivity (10.2%, 95% CI 4.31% to 22.3%) and specificity (96.5%,95% CI 87.1% to 99.1%) for PET-CT before SLNB (170 participants, 49 cases). When these data are translated to a hypothetical cohort of 1000 people eligible for SLNB, 237 of whom have nodal metastases (median prevalence), the combination of ultrasound with FNAC potentially allows 43 people with nodal metastases to be triaged directly to adjuvant therapy rather than having SLNB first, at a cost of two people with false positive results (who are incorrectly managed). Those with a false negative ultrasound will be identified on subsequent SLNB.Limited test accuracy data were available for whole body imaging via PET-CT for primary staging or re-staging for disease recurrence, and none evaluated MRI. Twenty-four studies evaluated whole body imaging. Six of these studies explored primary staging following a confirmed diagnosis of melanoma (492 participants), three evaluated re-staging of disease following some clinical indication of recurrence (589 participants), and 15 included mixed or not clearly described population groups comprising participants at a number of different points on the clinical pathway and at varying stages of disease (1265 participants). Results for whole body imaging could not be translated to a hypothetical cohort of people due to paucity of data.Most of the studies (6/9) of primary disease or re-staging of disease considered PET-CT, two in comparison to CT alone, and three studies examined the use of ultrasound. No eligible evaluations of MRI in these groups were identified. All studies used histological reference standards combined with follow-up, and two included FNAC for some participants. Observed accuracy for detection of any metastases for PET-CT was higher for re-staging of disease (summary sensitivity from two studies: 92.6%, 95% CI 85.3% to 96.4%; specificity: 89.7%, 95% CI 78.8% to 95.3%; 153 participants; 95 cases) compared to primary staging (sensitivities from individual studies ranged from 30% to 47% and specificities from 73% to 88%), and was more sensitive than CT alone in both population groups, but participant numbers were very small.No conclusions can be drawn regarding routine imaging of the brain via MRI or CT. AUTHORS' CONCLUSIONS Review authors found a disappointing lack of evidence on the accuracy of imaging in people with a diagnosis of melanoma at different points on the clinical pathway. Studies were small and often reported data according to the number of lesions rather than the number of study participants. Imaging with ultrasound combined with FNAC before SLNB may identify around one-fifth of those with nodal disease, but confidence intervals are wide and further work is needed to establish cost-effectiveness. Much of the evidence for whole body imaging for primary staging or re-staging of disease is focused on PET-CT, and comparative data with CT or MRI are lacking. Future studies should go beyond diagnostic accuracy and consider the effects of different imaging tests on disease management. The increasing availability of adjuvant therapies for people with melanoma at high risk of disease spread at presentation will have a considerable impact on imaging services, yet evidence for the relative diagnostic accuracy of available tests is limited.
Atrial fibrillation (AF) is the arrhythmia most commonly diagnosed in clinical practice. It is associated with significant morbidity and mortality. Prevalence of AF and complications of AF, estimated by hospitalisations, have increased dramatically in the last decade. Being able to predict AF would allow tailoring of management strategies and a focus on primary or secondary prevention. Models predicting recurrent AF would have particular clinical use for the selection of rhythm control therapy. There are existing prognostic models which combine several predictors or risk factors to generate an individualised estimate of risk of AF. The aim of this systematic review is to summarise and compare model performance measures and predictive accuracy across different models and populations at risk of developing incident or recurrent AF. Methods tailored to systematic reviews of prognostic models will be used for study identification, risk of bias assessment and synthesis. Studies will be eligible for inclusion where they report an internally or externally validated model. The quality of studies reporting a prognostic model will be assessed using the Prediction Study Risk Of Bias Assessment Tool (PROBAST). Studies will be narratively described and included variables and predictive accuracy compared across different models and populations. Meta-analysis of model performance measures for models validated in similar populations will be considered where possible. To the best of our knowledge, this will be the first systematic review to collate evidence from all studies reporting on validated prognostic models, or on the impact of such models, in any population at risk of incident or recurrent AF. The review may identify models which are suitable for impact assessment in clinical practice. Should gaps in the evidence be identified, research recommendations relating to model development, validation or impact assessment will be made. Findings will be considered in the context of any models already used in clinical practice, and the extent to which these have been validated. PROSPERO ( CRD42018111649 ).
College of Medicine and Health, University of Exeter, Exeter, UK Department of Women's and Children's Health, Dunedin School of Medicine, University of Otago, Dunedin, New Zealand Department of Physiotherapy, School of Primary and Allied Health Care, Monash University, Melbourne, Australia Warwick Business School, University of Warwick, Coventry, UK Institute of Applied Health Research, University of Birmingham, Birmingham, UK Warwick Clinical Trials Unit, University of Warwick, Coventry, UK Wolfson Palliative Care Research Centre, Institute for Clinical & Applied Health Research, Hull York Medical School, University of Hull, Heslington, UK
DS01 Cochrane systematic review of the diagnostic accuracy of reflectance confocal microscopy for detection of melanoma J. Dinnes, J. Deeks, D. Saleh, N. Chuchu, S. Bayliss, L. Patel, C. Davenport, Y. Takwoingi, K. Godfrey, R. Matin, R. Patalay and H. Williams Institute of Applied Health Research, University of Birmingham, Birmingham, U.K., Newcastle Hospitals NHS Trust, Royal Victoria Infirmary, Newcastle, U.K., Royal Stoke Hospital, Stoke on Trent, U.K., The University of Nottingham, Nottingham, U.K., Churchill Hospital, Oxford, U.K. and Guy’s and St Thomas’ NHS Foundation Trust, London, U.K. Early detection of melanoma is key to improve survival rates. Reflectance confocal microscopy (RCM) in conjunction with clinical and/or dermoscopic examination may reduce unnecessary excisions without missing melanomas. The aim of this study was to conduct a Cochrane systematic review of the diagnostic accuracy of RCM for detection of melanoma in adults with (i) any lesion suspicious for melanoma and (ii) lesions that are difficult to diagnose, and to compare its accuracy with that of dermoscopy. A comprehensive search of 10 databases up to August 2016 identified studies of any design evaluating RCM in adults with lesions suspicious for melanoma, compared with histology or clinical follow-up. Two reviewers independently extracted data and quality assessment. The accuracy was estimated using hierarchical summary ROC methods; sensitivities and specificities were estimated for selected points on the summary receiver operating characteristic (SROC) curve. A total of 18 studies were included. For any lesion, suspicious for melanoma (9 datasets; 1452 lesions, 370 melanomas), at a fixed sensitivity for both tests of 90%, specificities were 82% for RCM and 42% for dermoscopy for the detection of melanoma or intraepidermal melanocytic variants. In a population of 1000 lesions at median observed melanoma prevalence of 30%, unnecessary excisions would be reduced by 280. In equivocal lesions (7 datasets; 1177 lesions, 180 melanomas), specificities were 86% and 49% for RCM and for dermoscopy. At a median melanoma prevalence of 20%, unnecessary excisions would be reduced by 296. The sensitivity and specificity of the Pellacani RCM score at a threshold of ≥3 were estimated at 92% (95% confidence interval (CI) 87–95) and 72% (95% CI 62–81). RCM has a potential role for assessing lesions that are difficult to diagnose using visual inspection and dermoscopy alone. There is a paucity of data comparing RCM with dermoscopy in a real world setting in a representative population. DS02 Cochrane systematic review of diagnostic accuracy of dermoscopy in comparison to visual inspection for the diagnosis of melanoma R.N. Matin, N. Chuchu, J. Dinnes, J.J. Deeks, L. Ferrante di Ruffano, D.R. Thomson, K.Y. Wong, R.B. Aldridge, R. Abbott, M. Fawzy, S.E. Bayliss, M.J. Grainge, Y. Takwoingi, C. Davenport, K. Godfrey, F.M. Walter and H. Williams Oxford University Hospitals NHS Foundation Trust, Oxford, U.K., University of Birmingham, Birmingham, U.K., St George’s Hospital, London, U.K., Salisbury NHS Foundation Trust, Salisbury, U.K., NHS Lothian, Edinburgh, U.K., University Hospital of Wales, Cardiff, U.K., Norfolk and Norwich University Hospitals NHS Trust, Norwich, U.K., University of Nottingham School of Medicine, Nottingham, U.K., University of Cambridge, Cambridge, U.K. and University of Nottingham, Nottingham, U.K. Early detection of melanoma is essential to improve survival. The additional value of dermoscopy over and above visual inspection (VI) of a suspicious skin lesion is critical to understand its contribution to the diagnosis of melanoma. A Cochrane systematic review of the diagnostic accuracy of dermoscopy for detection of melanoma in adults was undertaken for (i) in-person diagnosis and (ii) diagnosis based on dermoscopic images, and to compare its accuracy with VI alone. A comprehensive search of 10 databases up to August 2016 identified studies of any design evaluating dermoscopy in adults with lesions suspicious for melanoma, compared with histology or clinical follow-up. Two reviewers independently extracted data and quality assessment (using QUADAS-2). The accuracy was estimated using hierarchical summary ROC methods; sensitivities and specificities were estimated for selected points on the summary receiver operating characteristic (SROC) curve. Overall, 106 publications were included. The detection of melanoma or intraepidermal melanocytic variants was analysed for 27 in-person (23 487 lesions; 1737 melanomas) and 60 image-based (13 475 lesions; 2851 melanomas) datasets. In-person dermoscopy was more accurate than image-based interpretation [relative diagnostic odds ratio (RDOR) 4.5; 95% CI 2.3–8.5, P < 0.0001]. Dermoscopy was more accurate than VI alone; RDORs (i) 4.8 (95% CI: 3.1– 7.4; P < 0.0001) for in-person and (ii) 5.6 (95% CI: 3.7–8.5; P < 0.0001) for image-based evaluations. Predicted increases in sensitivity were (i) 16% (92% vs. 76%) and (ii) 34% (81% vs. 47%) at a fixed specificity of 80%. Use of a published algorithm to assist dermoscopy had no significant impact on accuracy. The accuracy was significantly higher for experienced observers compared with less experienced. Dermoscopy is a valuable tool to support VI of suspicious skin lesions to detect melanoma, particularly in referred populations and for
BACKGROUND Stillbirth affects 2.6 million pregnancies worldwide each year. Whilst the majority of cases occur in low- and middle-income countries, stillbirth remains an important clinical issue for high-income countries (HICs) - with both the UK and the USA reporting rates above the mean for HICs. In HICs, the most frequently reported association with stillbirth is placental dysfunction. Placental dysfunction may be evident clinically as fetal growth restriction (FGR) and small-for-dates infants. It can be caused by placental abruption or hypertensive disorders of pregnancy and many other disorders and factorsPlacental abnormalities are noted in 11% to 65% of stillbirths. Identification of FGA is difficult in utero. Small-for-gestational age (SGA), as assessed after birth, is the most commonly used surrogate measure for this outcome. The degree of SGA is associated with the likelihood of FGR; 30% of infants with a birthweight < 10th centile are thought to be FGR, while 70% of infants with a birthweight < 3rd centile are thought to be FGR. Critically, SGA is the most significant antenatal risk factor for a stillborn infant. Correct identification of SGA infants is associated with a reduction in the perinatal mortality rate. However, currently used tests, such as measurement of symphysis-fundal height, have a low reported sensitivity and specificity for the identification of SGA infants. OBJECTIVES The primary objective was to assess and compare the diagnostic accuracy of ultrasound assessment of fetal growth by estimated fetal weight (EFW) and placental biomarkers alone and in any combination used after 24 weeks of pregnancy in the identification of placental dysfunction as evidenced by either stillbirth, or birth of a SGA infant. Secondary objectives were to investigate the effect of clinical and methodological factors on test performance. SEARCH METHODS We developed full search strategies with no language or date restrictions. The following sources were searched: MEDLINE, MEDLINE In Process and Embase via Ovid, Cochrane (Wiley) CENTRAL, Science Citation Index (Web of Science), CINAHL (EBSCO) with search strategies adapted for each database as required; ISRCTN Registry, UK Clinical Trials Gateway, WHO International Clinical Trials Portal and ClinicalTrials.gov for ongoing studies; specialist abstract and conference proceeding resources (British Library's ZETOC and Web of Science Conference Proceedings Citation Index). Search last conducted in Ocober 2016. SELECTION CRITERIA We included studies of pregnant women of any age with a gestation of at least 24 weeks if relevant outcomes of pregnancy (live birth/stillbirth; SGA infant) were assessed. Studies were included irrespective of whether pregnant women were deemed to be low or high risk for complications or were of mixed populations (low and high risk). Pregnancies complicated by fetal abnormalities and multi-fetal pregnancies were excluded as they have a higher risk of stillbirth from non-placental causes. With regard to biochemical tests, we included assays performed using any technique and at any threshold used to determine test positivity. DATA COLLECTION AND ANALYSIS We extracted the numbers of true positive, false positive, false negative, and true negative test results from each study. We assessed risk of bias and applicability using the QUADAS-2 tool. Meta-analyses were performed using the hierarchical summary ROC model to estimate and compare test accuracy. MAIN RESULTS We included 91 studies that evaluated seven tests - blood tests for human placental lactogen (hPL), oestriol, placental growth factor (PlGF) and uric acid, ultrasound EFW and placental grading and urinary oestriol - in a total of 175,426 pregnant women, in which 15,471 pregnancies ended in the birth of a small baby and 740 pregnancies which ended in stillbirth. The quality of included studies was variable with most domains at low risk of bias although 59% of studies were deemed to be of unclear risk of bias for the reference standard domain. Fifty-three per cent of studies were of high concern for applicability due to inclusion of only high- or low-risk women.Using all available data for SGA (86 studies; 159,490 pregnancies involving 15,471 SGA infants), there was evidence of a difference in accuracy (P < 0.0001) between the seven tests for detecting pregnancies that are SGA at birth. Ultrasound EFW was the most accurate test for detecting SGA at birth with a diagnostic odds ratio (DOR) of 21.3 (95% CI 13.1 to 34.6); hPL was the most accurate biochemical test with a DOR of 4.78 (95% CI 3.21 to 7.13). In a hypothetical cohort of 1000 pregnant women, at the median specificity of 0.88 and median prevalence of 19%, EFW, hPL, oestriol, urinary oestriol, uric acid, PlGF and placental grading will miss 50 (95% CI 32 to 68), 116 (97 to 133), 124 (108 to 137), 127 (95 to 152), 139 (118 to 154), 144 (118 to 161), and 144 (122 to 161) SGA infants, respectively. For the detection of pregnancies ending in stillbirth (21 studies; 100,687 pregnancies involving 740 stillbirths), in an indirect comparison of the four biochemical tests, PlGF was the most accurate test with a DOR of 49.2 (95% CI 12.7 to 191). In a hypothetical cohort of 1000 pregnant women, at the median specificity of 0.78 and median prevalence of 1.7%, PlGF, hPL, urinary oestriol and uric acid will miss 2 (95% CI 0 to 4), 4 (2 to 8), 6 (6 to 7) and 8 (3 to 13) stillbirths, respectively. No studies assessed the accuracy of ultrasound EFW for detection of pregnancy ending in stillbirth. AUTHORS' CONCLUSIONS Biochemical markers of placental dysfunction used alone have insufficient accuracy to identify pregnancies ending in SGA or stillbirth. Studies combining U and placental biomarkers are needed to determine whether this approach improves diagnostic accuracy over the use of ultrasound estimation of fetal size or biochemical markers of placental dysfunction used alone. Many of the studies included in this review were carried out between 1974 and 2016. Studies of placental substances were mostly carried out before 1991 and after 2013; earlier studies may not reflect developments in test technology.
COPD self-management reduces hospital admissions and improves health-related quality of life (HRQoL). However, whilst most patients are managed in primary care, the majority of self-management trials have recruited participants with more severe disease from secondary care. We report the findings of a systematic review of the effectiveness of community-based self-management interventions in primary care patients with COPD. We systematically searched eleven electronic databases and identified 12 eligible randomised controlled trials with seven included in meta-analyses for HRQoL, anxiety and depression. We report no difference in HRQoL at final follow-up (St George’s Respiratory Questionnaire total score −0.29; 95%CI −2.09, 1.51; I 2 0%), nor any difference in anxiety or depression. In conclusion, supported self-management interventions delivered in the community to patients from primary care do not appear to be effective. Further research is recommended to identify effective self-management interventions suitable for primary care populations, particularly those with milder disease.
Background Early accurate detection of all skin cancer types is essential to guide appropriate management, reduce morbidity and improve survival. Basal cell carcinoma (BCC) is usually localised to the skin but has potential to infiltrate and damage surrounding tissue, while cutaneous squamous cell carcinoma (cSCC) and melanoma have a much higher potential to metastasise and ultimately lead to death. Exfoliative cytology is a non-invasive test that uses the Tzanck smear technique to identify disease by examining the structure of cells obtained from scraped samples. This simple procedure is a less invasive diagnostic test than a skin biopsy, and for BCC it has the potential to provide an immediate diagnosis that avoids an additional clinic visit to receive skin biopsy results. This may benefit patients scheduled for either Mohs micrographic surgery or non-surgical treatments such as radiotherapy. A cytology scrape can never give the same information as a skin biopsy, however, so it is important to better understand in which skin cancer situations it may be helpful. Objectives To determine the diagnostic accuracy of exfoliative cytology for detecting basal cell carcinoma (BCC) in adults, and to compare its accuracy with that of standard diagnostic practice (visual inspection with or without dermoscopy). Secondary objectives were: to determine the diagnostic accuracy of exfoliative cytology for detecting cSCC, invasive melanoma and atypical intraepidermal melanocytic variants, and any other skin cancer; and for each of these secondary conditions to compare the accuracy of exfoliative cytology with visual inspection with or without dermoscopy in direct test comparisons; and to determine the effect of observer experience. Search methods We undertook a comprehensive search of the following databases from inception up to August 2016: Cochrane Central Register of Controlled Trials; MEDLINE; Embase; CINAHL; CPCI; Zetoc; Science Citation Index; US National Institutes of Health Ongoing Trials Register; NIHR Clinical Research Network Portfolio Database; and the World Health Organization International Clinical Trials Registry Platform. We also studied the reference lists of published systematic review articles. Selection criteria Studies evaluating exfoliative cytology in adults with lesions suspicious for BCC, cSCC or melanoma, compared with a reference standard of histological confirmation. Data collection and analysis Two review authors independently extracted all data using a standardised data extraction and quality assessment form (based on QUADAS-2). Where possible we estimated summary sensitivities and specificities using the bivariate hierarchical model. Main results We synthesised the results of nine studies contributing a total of 1655 lesions to our analysis, including 1120 BCCs (14 datasets), 41 cSCCs (amongst 401 lesions in 2 datasets), and 10 melanomas (amongst 200 lesions in 1 dataset). Three of these datasets (one each for BCC, melanoma and any malignant condition) were derived from one study that also performed a direct comparison with dermoscopy. Studies were of moderate to poor quality, providing inadequate descriptions of participant selection, thresholds used to make cytological and histological diagnoses, and blinding. Reporting of participants' prior referral pathways was particularly poor, as were descriptions of the cytodiagnostic criteria used to make diagnoses. No studies evaluated the use of exfoliative cytology as a primary diagnostic test for detecting BCC or other skin cancers in lesions suspicious for skin cancer. Pooled data from seven studies using standard cytomorphological criteria (but various stain methods) to detect BCC in participants with a high clinical suspicion of BCC estimated the sensitivity and specificity of exfoliative cytology as 97.5% (95% CI 94.5% to 98.9%) and 90.1% (95% CI 81.1% to 95.1%). respectively. When applied to a hypothetical population of 1000 clinically suspected BCC lesions with a median observed BCC prevalence of 86%, exfoliative cytology would miss 21 BCCs and would lead to 14 false positive diagnoses of BCC. No false positive cases were histologically confirmed to be melanoma. Insufficient data are available to make summary statements regarding the accuracy of exfoliative cytology to detect melanoma or cSCC, or its accuracy compared to dermoscopy. Authors' conclusions The utility of exfoliative cytology for the primary diagnosis of skin cancer is unknown, as all included studies focused on the use of this technique for confirming strongly suspected clinical diagnoses. For the confirmation of BCC in lesions with a high clinical suspicion, there is evidence of high sensitivity and specificity. Since decisions to treat low-risk BCCs are unlikely in practice to require diagnostic confirmation given that clinical suspicion is already high, exfoliative cytology might be most useful for cases of BCC where the treatments being contemplated require a tissue diagnosis (e.g. radiotherapy). The small number of included studies, poor reporting and varying methodological quality prevent us from drawing strong conclusions to guide clinical practice. Despite insufficient data on the use of cytology for cSCC or melanoma, it is unlikely that cytology would be useful in these scenarios since preservation of the architecture of the whole lesion that would be available from a biopsy provides crucial diagnostic information. Given the paucity of good quality data, appropriately designed prospective comparative studies may be required to evaluate both the diagnostic value of exfoliative cytology by comparison to dermoscopy, and its confirmatory value in adequately reported populations with a high probability of BCC scheduled for further treatment requiring a tissue diagnosis.
Background Melanoma accounts for a small proportion of all skin cancer cases but is responsible for most skin cancer-related deaths. Early detection and treatment can improve survival. Smartphone applications are readily accessible and potentially offer an instant risk assessment of the likelihood of malignancy so that the right people seek further medical attention from a clinician for more detailed assessment of the lesion. There is, however, a risk that melanomas will be missed and treatment delayed if the application reassures the user that their lesion is low risk. Objectives To assess the diagnostic accuracy of smartphone applications to rule out cutaneous invasive melanoma and atypical intraepidermal melanocytic variants in adults with concerns about suspicious skin lesions. Search methods We undertook a comprehensive search of the following databases from inception to August 2016: Cochrane Central Register of Controlled Trials; MEDLINE; Embase; CINAHL; CPCI; Zetoc; Science Citation Index; US National Institutes of Health Ongoing Trials Register; NIHR Clinical Research Network Portfolio Database; and the World Health Organization International Clinical Trials Registry Platform. We studied reference lists and published systematic review articles. Selection criteria Studies of any design evaluating smartphone applications intended for use by individuals in a community setting who have lesions that might be suspicious for melanoma or atypical intraepidermal melanocytic variants versus a reference standard of histological confirmation or clinical follow-up and expert opinion. Data collection and analysis Two review authors independently extracted all data using a standardised data extraction and quality assessment form (based on QUADAS-2). Due to scarcity of data and poor quality of studies, we did not perform a meta-analysis for this review. For illustrative purposes, we plotted estimates of sensitivity and specificity on coupled forest plots for each application under consideration. Main results This review reports on two cohorts of lesions published in two studies. Both studies were at high risk of bias from selective participant recruitment and high rates of non-evaluable images. Concerns about applicability of findings were high due to inclusion only of lesions already selected for excision in a dermatology clinic setting, and image acquisition by clinicians rather than by smartphone app users. We report data for five mobile phone applications and 332 suspicious skin lesions with 86 melanomas across the two studies. Across the four artificial intelligence-based applications that classified lesion images (photographs) as melanomas (one application) or as high risk or 'problematic' lesions (three applications) using a pre-programmed algorithm, sensitivities ranged from 7% (95% CI 2% to 16%) to 73% (95% CI 52% to 88%) and specificities from 37% (95% CI 29% to 46%) to 94% (95% CI 87% to 97%). The single application using store-and-forward review of lesion images by a dermatologist had a sensitivity of 98% (95% CI 90% to 100%) and specificity of 30% (95% CI 22% to 40%). The number of test failures (lesion images analysed by the applications but classed as 'unevaluable' and excluded by the study authors) ranged from 3 to 31 (or 2% to 18% of lesions analysed). The store-and-forward application had one of the highest rates of test failure (15%). At least one melanoma was classed as unevaluable in three of the four application evaluations. Authors' conclusions Smartphone applications using artificial intelligence-based analysis have not yet demonstrated sufficient promise in terms of accuracy, and they are associated with a high likelihood of missing melanomas. Applications based on store-and-forward images could have a potential role in the timely presentation of people with potentially malignant lesions by facilitating active self-management health practices and early engagement of those with suspicious skin lesions; however, they may incur a significant increase in resource and workload. Given the paucity of evidence and low methodological quality of existing studies, it is not possible to draw any implications for practice. Nevertheless, this is a rapidly advancing field, and new and better applications with robust reporting of studies could change these conclusions substantially.
BACKGROUNDMelanoma has one of the fastest rising incidence rates of any cancer. It accounts for a small percentage of skin cancer cases but is responsible for the majority of skin cancer deaths. Early detection and treatment is key to improving survival; however, anxiety around missing early cases needs to be balanced against appropriate levels of referral and excision of benign lesions. Used in conjunction with clinical or dermoscopic suspicion of malignancy, or both, reflectance confocal microscopy (RCM) may reduce unnecessary excisions without missing melanoma cases.OBJECTIVESTo determine the diagnostic accuracy of reflectance confocal microscopy for the detection of cutaneous invasive melanoma and atypical intraepidermal melanocytic variants in adults with any lesion suspicious for melanoma and lesions that are difficult to diagnose, and to compare its accuracy with that of dermoscopy.SEARCH METHODSWe undertook a comprehensive search of the following databases from inception up to August 2016: Cochrane Central Register of Controlled Trials; MEDLINE; Embase; and seven other databases. We studied reference lists and published systematic review articles.SELECTION CRITERIAStudies of any design that evaluated RCM alone, or RCM in comparison to dermoscopy, in adults with lesions suspicious for melanoma or atypical intraepidermal melanocytic variants, compared with a reference standard of either histological confirmation or clinical follow-up.DATA COLLECTION AND ANALYSISTwo review authors independently extracted all data using a standardised data extraction and quality assessment form (based on QUADAS-2). We contacted authors of included studies where information related to the target condition or diagnostic threshold were missing. We estimated summary sensitivities and specificities per algorithm and threshold using the bivariate hierarchical model. To compare RCM with dermoscopy, we grouped studies by population (defined by difficulty of lesion diagnosis) and combined data using hierarchical summary receiver operating characteristic (SROC) methods. Analysis of studies allowing direct comparison between tests was undertaken. To facilitate interpretation of results, we computed values of specificity at the point on the SROC curve with 90% sensitivity as this value lies within the estimates for the majority of analyses. We investigated the impact of using a purposely developed RCM algorithm and in-person test interpretation.MAIN RESULTSThe search identified 18 publications reporting on 19 study cohorts with 2838 lesions (including 658 with melanoma), which provided 67 datasets for RCM and seven for dermoscopy. Studies were generally at high or unclear risk of bias across almost all domains and of high or unclear concern regarding applicability of the evidence. Selective participant recruitment, lack of blinding of the reference test to the RCM result, and differential verification were particularly problematic. Studies may not be representative of populations eligible for RCM, and test interpretation was often undertaken remotely from the patient and blinded to clinical information.Meta-analysis found RCM to be more accurate than dermoscopy in studies of participants with any lesion suspicious for melanoma and in participants with lesions that were more difficult to diagnose (equivocal lesion populations). Assuming a fixed sensitivity of 90% for both tests, specificities were 82% for RCM and 42% for dermoscopy for any lesion suspicious for melanoma (9 RCM datasets; 1452 lesions and 370 melanomas). For a hypothetical population of 1000 lesions at the median observed melanoma prevalence of 30%, this equated to a reduction in unnecessary excisions with RCM of 280 compared to dermoscopy, with 30 melanomas missed by both tests. For studies in equivocal lesions, specificities of 86% would be observed for RCM and 49% for dermoscopy (7 RCM datasets; 1177 lesions and 180 melanomas). At the median observed melanoma prevalence of 20%, this reduced unnecessary excisions by 296 with RCM compared with dermoscopy, with 20 melanomas missed by both tests. Across all populations, algorithms and thresholds assessed, the sensitivity and specificity of the Pellacani RCM score at a threshold of three or greater were estimated at 92% (95% confidence interval (CI) 87 to 95) for RCM and 72% (95% CI 62 to 81) for dermoscopy.AUTHORS' CONCLUSIONSRCM may have a potential role in clinical practice, particularly for the assessment of lesions that are difficult to diagnose using visual inspection and dermoscopy alone, where the evidence suggests that RCM may be both more sensitive and specific in comparison to dermoscopy. Given the paucity of data to allow comparison with dermoscopy, the results presented require further confirmation in prospective studies comparing RCM with dermoscopy in a real-world setting in a representative population.
Background Early accurate detection of all skin cancer types is essential to guide appropriate management and to improve morbidity and survival. Melanoma and cutaneous squamous cell carcinoma (cSCC) are high-risk skin cancers which have the potential to metastasise and ultimately lead to death, whereas basal cell carcinoma (BCC) is usually localised with potential to infiltrate and damage surrounding tissue. Anxiety around missing early curable cases needs to be balanced against inappropriate referral and unnecessary excision of benign lesions. Computer-assisted diagnosis (CAD) systems use artificial intelligence to analyse lesion data and arrive at a diagnosis of skin cancer. When used in unreferred settings ('primary care'), CAD may assist general practitioners (GPs) or other clinicians to more appropriately triage high-risk lesions to secondary care. Used alongside clinical and dermoscopic suspicion of malignancy, CAD may reduce unnecessary excisions without missing melanoma cases. Objectives To determine the accuracy of CAD systems for diagnosing cutaneous invasive melanoma and atypical intraepidermal melanocytic variants, BCC or cSCC in adults, and to compare its accuracy with that of dermoscopy. Search methods We undertook a comprehensive search of the following databases from inception up to August 2016: Cochrane Central Register of Controlled Trials (CENTRAL); MEDLINE; Embase; CINAHL; CPCI; Zetoc; Science Citation Index; US National Institutes of Health Ongoing Trials Register; NIHR Clinical Research Network Portfolio Database; and the World Health Organization International Clinical Trials Registry Platform. We studied reference lists and published systematic review articles. Selection criteria Studies of any design that evaluated CAD alone, or in comparison with dermoscopy, in adults with lesions suspicious for melanoma or BCC or cSCC, and compared with a reference standard of either histological confirmation or clinical follow-up. Data collection and analysis Two review authors independently extracted all data using a standardised data extraction and quality assessment form (based on QUADAS-2). We contacted authors of included studies where information related to the target condition or diagnostic threshold were missing. We estimated summary sensitivities and specificities separately by type of CAD system, using the bivariate hierarchical model. We compared CAD with dermoscopy using (a) all available CAD data (indirect comparisons), and (b) studies providing paired data for both tests (direct comparisons). We tested the contribution of human decision-making to the accuracy of CAD diagnoses in a sensitivity analysis by removing studies that gave CAD results to clinicians to guide diagnostic decision-making. Main results We included 42 studies, 24 evaluating digital dermoscopy-based CAD systems (Derm-CAD) in 23 study cohorts with 9602 lesions (1220 melanomas, at least 83 BCCs, 9 cSCCs), providing 32 datasets for Derm-CAD and seven for dermoscopy. Eighteen studies evaluated spectroscopy-based CAD (Spectro-CAD) in 16 study cohorts with 6336 lesions (934 melanomas, 163 BCC, 49 cSCCs), providing 32 datasets for Spectro-CAD and six for dermoscopy. These consisted of 15 studies using multispectral imaging (MSI), two studies using electrical impedance spectroscopy (EIS) and one study using diffuse-reflectance spectroscopy. Studies were incompletely reported and at unclear to high risk of bias across all domains. Included studies inadequately address the review question, due to an abundance of low-quality studies, poor reporting, and recruitment of highly selected groups of participants. Across all CAD systems, we found considerable variation in the hardware and software technologies used, the types of classification algorithm employed, methods used to train the algorithms, and which lesion morphological features were extracted and analysed across all CAD systems, and even between studies evaluating CAD systems. Meta-analysis found CAD systems had high sensitivity for correct identification of cutaneous invasive melanoma and atypical intraepidermal melanocytic variants in highly selected populations, but with low and very variable specificity, particularly for Spectro-CAD systems. Pooled data from 22 studies estimated the sensitivity of Derm-CAD for the detection of melanoma as 90.1% (95% confidence interval (CI) 84.0% to 94.0%) and specificity as 74.3% (95% CI 63.6% to 82.7%). Pooled data from eight studies estimated the sensitivity of multispectral imaging CAD (MSI-CAD) as 92.9% (95% CI 83.7% to 97.1%) and specificity as 43.6% (95% CI 24.8% to 64.5%). When applied to a hypothetical population of 1000 lesions at the mean observed melanoma prevalence of 20%, Derm-CAD would miss 20 melanomas and would lead to 206 false-positive results for melanoma. MSI-CAD would miss 14 melanomas and would lead to 451 false diagnoses for melanoma. Preliminary findings suggest CAD systems are at least as sensitive as assessment of dermoscopic images for the diagnosis of invasive melanoma and atypical intraepidermal melanocytic variants. We are unable to make summary statements about the use of CAD in unreferred populations, or its accuracy in detecting keratinocyte cancers, or its use in any setting as a diagnostic aid, because of the paucity of studies. Authors' conclusions In highly selected patient populations all CAD types demonstrate high sensitivity, and could prove useful as a back-up for specialist diagnosis to assist in minimising the risk of missing melanomas. However, the evidence base is currently too poor to understand whether CAD system outputs translate to different clinical decision-making in practice. Insufficient data are available on the use of CAD in community settings, or for the detection of keratinocyte cancers. The evidence base for individual systems is too limited to draw conclusions on which might be preferred for practice. Prospective comparative studies are required that evaluate the use of already evaluated CAD systems as diagnostic aids, by comparison to face-to-face dermoscopy, and in participant populations that are representative of those in which the test would be used in practice.