AimsCompare the prevalence of age-related cataract and the cataract surgical coverage rate between Indigenous and non-Indigenous Australians and explore differences in these estimates across location and time.MethodsThe Joanna Briggs Institute guidance for systematic reviews of prevalence studies was followed. A systematic search of Medline, Embase, Web of Science and grey literature from database inception to June 2022 was performed. All studies reporting cataract prevalence in Australian populations were included. Pooled prevalence estimates were derived using meta-analyses with a random-effects model. Nine studies enrolling 36 302 participants were included. Most studies only reported the prevalence of cataract causing vision loss (visual acuity<6/12) or blindness (visual acuity<6/60), restricting our meta-analysis to these definitions.ResultsCataract causing unilateral vision loss was common in both Indigenous and non-Indigenous adults (3.5% and 3.6%, p=0.891). Indigenous adults had a higher prevalence of bilateral vision loss (3.6% vs 1.1%, p=0.011) and bilateral blindness (0.385% vs 0.001%, p=0.002) than non-Indigenous adults. Cataract surgical coverage was lower in Indigenous (68.0%; 95% CI, 55.9 to 79.0) than non-Indigenous (88.4%; 95% CI, 79.9 to 94.8) adults (p=0.004). No differences in bilateral vision loss, blindness or surgical coverage were found between rural and urban subgroups or between studies conducted before and after the year 2000.ConclusionsCataract causes vision loss in a substantial number of adults living in urban and rural Australia. Policies to improve diagnosis and surgery rates should be prioritised, particularly for Indigenous Australians who experience a disproportionate burden of advanced cataract and reduced access to surgery.PROSPERO registration numberCRD42022340197.
Few metrics exist to describe phenotypic diversity within ophthalmic imaging datasets, with researchers often using ethnicity as a surrogate marker for biological variability. We derived a continuous, measured metric, the retinal pigment score (RPS), that quantifies the degree of pigmentation from a colour fundus photograph of the eye. RPS was validated using two large epidemiological studies with demographic and genetic data (UK Biobank and EPIC-Norfolk Study) and reproduced in a Tanzanian, an Australian, and a Chinese dataset. A genome-wide association study (GWAS) of RPS from UK Biobank identified 20 loci with known associations with skin, iris and hair pigmentation, of which eight were replicated in the EPIC-Norfolk cohort. There was a strong association between RPS and ethnicity, however, there was substantial overlap between each ethnicity and the respective distributions of RPS scores. RPS decouples traditional demographic variables from clinical imaging characteristics. RPS may serve as a useful metric to quantify the diversity of the training, validation, and testing datasets used in the development of AI algorithms to ensure adequate inclusion and explainability of the model performance, critical in evaluating all currently deployed AI models. The code to derive RPS is publicly available at: https://github.com/uw-biomedical-ml/retinal-pigmentation-score.
The advent of foundation models (FMs) is transforming medical domain. In ophthalmology, RETFound, a retina-specific FM pre-trained sequentially on 1.4 million natural images and 1.6 million retinal images, has demonstrated high adaptability across clinical applications. Conversely, DINOv2, a general-purpose vision FM pre-trained on 142 million natural images, has shown promise in non-medical domains. However, its applicability to clinical tasks remains underexplored. To address this, we conducted head-to-head evaluations by fine-tuning RETFound and three DINOv2 models (large, base, small) for ocular disease detection and systemic disease prediction tasks, across eight standardized open-source ocular datasets, as well as the Moorfields AlzEye and the UK Biobank datasets. DINOv2-large model outperformed RETFound in detecting diabetic retinopathy (AUROC=0.850-0.952 vs 0.823-0.944, across three datasets, all P<=0.007) and multi-class eye diseases (AUROC=0.892 vs. 0.846, P<0.001). In glaucoma, DINOv2-base model outperformed RETFound (AUROC=0.958 vs 0.940, P<0.001). Conversely, RETFound achieved superior performance over all DINOv2 models in predicting heart failure, myocardial infarction, and ischaemic stroke (AUROC=0.732-0.796 vs 0.663-0.771, all P<0.001). These trends persisted even with 10
Background/aims Deep learning systems (DLSs) for diabetic retinopathy (DR) detection show promising results but can underperform in racial and ethnic minority groups, therefore external validation within these populations is critical for health equity. This study evaluates the performance of a DLS for DR detection among Indigenous Australians, an understudied ethnic group who suffer disproportionately from DR-related blindness. Methods We performed a retrospective external validation study comparing the performance of a DLS against a retinal specialist for the detection of more-than-mild DR (mtmDR), vision-threatening DR (vtDR) and all-cause referable DR. The validation set consisted of 1682 consecutive, single-field, macula-centred retinal photographs from 864 patients with diabetes (mean age 54.9 years, 52.4% women) at an Indigenous primary care service in Perth, Australia. Three-person adjudication by a panel of specialists served as the reference standard. Results For mtmDR detection, sensitivity of the DLS was superior to the retina specialist (98.0% (95% CI, 96.5 to 99.4) vs 87.1% (95% CI, 83.6 to 90.6), McNemar’s test p<0.001) with a small reduction in specificity (95.1% (95% CI, 93.6 to 96.4) vs 97.0% (95% CI, 95.9 to 98.0), p=0.006). For vtDR, the DLS’s sensitivity was again superior to the human grader (96.2% (95% CI, 93.4 to 98.6) vs 84.4% (95% CI, 79.7 to 89.2), p<0.001) with a slight drop in specificity (95.8% (95% CI, 94.6 to 96.9) vs 97.8% (95% CI, 96.9 to 98.6), p=0.002). For all-cause referable DR, there was a substantial increase in sensitivity (93.7% (95% CI, 91.8 to 95.5) vs 74.4% (95% CI, 71.1 to 77.5), p<0.001) and a smaller reduction in specificity (91.7% (95% CI, 90.0 to 93.3) vs 96.3% (95% CI, 95.2 to 97.4), p<0.001). Conclusion The DLS showed improved sensitivity and similar specificity compared with a retina specialist for DR detection. This demonstrates its potential to support DR screening among Indigenous Australians, an underserved population with a high burden of diabetic eye disease.
Background Evidence on the performance of Generative Pre-trained Transformer 4 (GPT-4), a large language model (LLM), in the ophthalmology question-answering domain is needed. Methods We tested GPT-4 on two 260-question multiple choice question sets from the Basic and Clinical Science Course (BCSC) Self-Assessment Program and the OphthoQuestions question banks. We compared the accuracy of GPT-4 models with varying temperatures (creativity setting) and evaluated their responses in a subset of questions. We also compared the best-performing GPT-4 model to GPT-3.5 and to historical human performance. Results GPT-4- 0.3 (GPT-4 with a temperature of 0.3) achieved the highest accuracy among GPT-4 models, with 75.8% on the BCSC set and 70.0% on the OphthoQuestions set. The combined accuracy was 72.9%, which represents an 18.3% raw improvement in accuracy compared with GPT-3.5 (p<0.001). Human graders preferred responses from models with a temperature higher than 0 (more creative). Exam section, question difficulty and cognitive level were all predictive of GPT-4- 0.3 answer accuracy. GPT-4- 0.3' s performance was numerically superior to human performance on the BCSC (75.8% vs 73.3%) and OphthoQuestions (70.0% vs 63.0%), but the difference was not statistically significant (p=0.55 and p=0.09). Conclusion GPT-4, an LLM trained on non-ophthalmology-specific data, performs significantly better than its predecessor on simulated ophthalmology board-style exams. Remarkably, its performance tended to be superior to historical human performance, but that difference was not statistically significant in our study.
Objective: Recent developments in artificial intelligence (AI) have positioned it to transform several stages of the clinical trial process. In this study, we explore the role of AI in clinical trial recruitment of individuals with geographic atrophy (GA), an advanced stage of age-related macular degeneration, amidst numerous ongoing clinical trials for this condition. Design: Cross-sectional study. Subjects: Retrospective dataset from the INSIGHT Health Data Research Hub at Moorfields Eye Hospital in London, United Kingdom, including 306 651 patients (602 826 eyes) with suspected retinal disease who underwent OCT imaging between January 1, 2008 and April 10, 2023. Methods: A deep learning model was trained on OCT scans to identify patients potentially eligible for GA trials, using AI-generated segmentations of retinal tissue. This method's efficacy was compared against a traditional keyword-based electronic health record (EHR) search. A clinical validation with fundus auto- fluorescence (FAF) images was performed to calculate the positive predictive value of this approach, by comparing AI predictions with expert assessments. Main Outcome Measures: The primary outcomes included the positive predictive value of AI in identifying trial-eligible patients, and the secondary outcome was the intraclass correlation between GA areas segmented on FAF by experts and AI-segmented OCT scans. Results: The AI system shortlisted a larger number of eligible patients with greater precision (1139, positive predictive value: 63%; 95% confidence interval [CI]: 54%-71%) compared with the EHR search (693, positive predictive value: 40%; 95% CI: 39%-42%). A combined AI-EHR approach identified 604 eligible patients with a positive predictive value of 86% (95% CI: 79%-92%). Intraclass correlation of GA area segmented on FAF versus AI-segmented area on OCT was 0.77 (95% CI: 0.68-0.84) for cases meeting trial criteria. The AI also adjusts to the distinct imaging criteria from several clinical trials, generating tailored shortlists ranging from 438 to 1817 patients. Conclusions: This study demonstrates the potential for AI in facilitating automated prescreening for clinical trials in GA, enabling site feasibility assessments, data-driven protocol design, and cost reduction. Once treatments are available, similar AI systems could also be used to identify individuals who may benefit from treatment. Financial Disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article. Ophthalmology Science 2024;4:100566 (c) 2024 by the American Academy of Ophthalmology. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/ 4.0/).
Purpose of review Last year marked the development of the first foundation model in ophthalmology, RETFound, setting the stage for generalizable medical artificial intelligence (GMAI) that can adapt to novel tasks. Additionally, rapid advancements in large language model (LLM) technology, including models such as GPT-4 and Gemini, have been tailored for medical specialization and evaluated on clinical scenarios with promising results. This review explores the opportunities and challenges for further advancements in these technologies. Recent findings RETFound outperforms traditional deep learning models in specific tasks, even when only fine-tuned on small datasets. Additionally, LMMs like Med-Gemini and Medprompt GPT-4 perform better than out-of-the-box models for ophthalmology tasks. However, there is still a significant deficiency in ophthalmology-specific multimodal models. This gap is primarily due to the substantial computational resources required to train these models and the limitations of high-quality ophthalmology datasets. Summary Overall, foundation models in ophthalmology present promising opportunities but face challenges, particularly the need for high-quality, standardized datasets for training and specialization. Although development has primarily focused on large language and vision models, the greatest opportunities lie in advancing large multimodal models, which can more closely mimic the capabilities of clinicians.
Foundation models represent a paradigm shift in artificial intelligence (AI), evolving from narrow models designed for specific tasks to versatile, generalisable models adaptable to a myriad of diverse applications. Ophthalmology as a specialty has the potential to act as an exemplar for other medical specialties, offering a blueprint for integrating foundation models broadly into clinical practice. This review hopes to serve as a roadmap for eyecare professionals seeking to better understand foundation models, while equipping readers with the tools to explore the use of foundation models in their own research and practice. We begin by outlining the key concepts and technological advances which have enabled the development of these models, providing an overview of novel training approaches and modern AI architectures. Next, we summarise existing literature on the topic of foundation models in ophthalmology, encompassing progress in vision foundation models, large language models and large multimodal models. Finally, we outline major challenges relating to privacy, bias and clinical validation, and propose key steps forward to maximise the benefit of this powerful technology.
Purpose: To examine the association of physical activity (PA) with glaucoma and related traits, to assess whether genetic predisposition to glaucoma modified these associations, and to probe causal relationships using Mendelian randomization (MR). Design: Cross-sectional observational and gene-environment interaction analyses in the UK Biobank. Two-sample MR experiments using summary statistics from large genetic consortia. Participants: UK Biobank participants with data on self-reported or accelerometer-derived PA and intraocular pressure (IOP; n = 94 206 and n = 27 777, respectively), macular inner retinal OCT measurements (n = 36 274 and n = 9991, respectively), and glaucoma status (n = 86 803 and n = 23 556, respectively). Methods: We evaluated multivariable-adjusted associations of self-reported (International Physical Activity Questionnaire) and accelerometer-derived PA with IOP and macular inner retinal OCT parameters using linear regression and with glaucoma status using logistic regression. For all outcomes, we examined gene-PA interactions using a polygenic risk score (PRS) that combined the effects of 2673 genetic variants associated with glaucoma. Main Outcome Measures: Intraocular pressure, macular retinal nerve fiber layer (mRNFL) thickness, macular ganglion cell-inner plexiform layer (mGCIPL) thickness, and glaucoma status. Results: In multivariable-adjusted regression models, we found no association of PA level or time spent in PA with glaucoma status. Higher overall levels and greater time spent in higher levels of both self-reported and accelerometer-derived PA were associated positively with thicker mGCIPL (P < 0.001 for trend for each). Compared with the lowest quartile of PA, participants in the highest quartiles of accelerometer-derived moderate -and vigorous-intensity PA showed a thicker mGCIPL by +0.57 mm (P < 0.001) and +0.42 mm (P = 0.005). No association was found with mRNFL thickness. High overall level of self-reported PA was associated with a modestly higher IOP of +0.08 mmHg (P = 0.01), but this was not replicated in the accelerometry data. No associations were modified by a glaucoma PRS, and MR analyses did not support a causal relationship between PA and any glaucoma-related outcome. Conclusions: Higher overall PA level and greater time spent in moderate and vigorous PA were not associated with glaucoma status but were associated with thicker mGCIPL. Associations with IOP were modest and inconsistent. Despite the well-documented acute reduction in IOP after PA, we found no evidence that high levels of habitual PA are associated with glaucoma status or IOP in the general population. Financial Disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article. Ophthalmology 2023;130:1024-1036 (c) 2023 by the American Academy of Ophthalmology. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
Our website uses cookies to enhance your experience. By continuing to use our site, or clicking "Continue," you are agreeing to our Cookie Policy | Continue JAMA Ophthalmology HomeNew OnlineCurrent IssueFor Authors Podcast Journals JAMA JAMA Network Open JAMA Cardiology JAMA Dermatology JAMA Health Forum JAMA Internal Medicine JAMA Neurology JAMA Oncology JAMA Ophthalmology JAMA Otolaryngology–Head & Neck Surgery JAMA Pediatrics JAMA Psychiatry JAMA Surgery Archives of Neurology & Psychiatry (1919-1959) JN Learning / CMESubscribeJobsInstitutions / LibrariansReprints & Permissions Terms of Use | Privacy Policy | Accessibility Statement 2024 American Medical Association. All Rights Reserved Search All JAMA JAMA Network Open JAMA Cardiology JAMA Dermatology JAMA Forum Archive JAMA Health Forum JAMA Internal Medicine JAMA Neurology JAMA Oncology JAMA Ophthalmology JAMA Otolaryngology–Head & Neck Surgery JAMA Pediatrics JAMA Psychiatry JAMA Surgery Archives of Neurology & Psychiatry Input Search Term Sign In Individual Sign In Sign inCreate an Account Access through your institution Sign In Purchase Options: Buy this article Rent this article Subscribe to the JAMA Ophthalmology journal
Purpose: To examine the associations of alcohol consumption with glaucoma and related traits, to assess whether a genetic predisposition to glaucoma modified these associations, and to perform Mendelian random-ization (MR) experiments to probe causal effects.Design: Cross-sectional observational and gene-environment interaction analyses in the UK Biobank. Two -sample MR experiments using summary statistics from large genetic consortia. Participants: UK Biobank participants with data on intraocular pressure (IOP) (n = 109 097), OCT-derived macular inner retinal layer thickness measures (n = 46 236) and glaucoma status (n = 173 407).Methods: Participants were categorized according to self-reported drinking behaviors. Quantitative esti-mates of alcohol intake were derived from touchscreen questionnaires and food composition tables. We per-formed a 2-step analysis, first comparing categories of alcohol consumption (never, infrequent, regular, and former drinkers) before assessing for a dose-response effect in regular drinkers only. Multivariable linear, logistic, and restricted cubic spline regression, adjusted for key sociodemographic, medical, anthropometric, and lifestyle factors, were used to examine associations. We assessed whether any association was modified by a multitrait glaucoma polygenic risk score. The inverse-variance weighted method was used for the main MR analyses.Main Outcome Measures: Intraocular pressure, macular retinal nerve fiber layer (mRNFL) thickness, mac-ular ganglion cell-inner plexiform layer (mGCIPL) thickness, and prevalent glaucoma.Results: Compared with infrequent drinkers, regular drinkers had higher IOP (+0.17 mmHg; P < 0.001) and thinner mGCIPL (-0.17 mm; P = 0.049), whereas former drinkers had a higher prevalence of glaucoma (odds ratio, 1.53; P = 0.002). In regular drinkers, alcohol intake was adversely associated with all outcomes in a dose -dependent manner (all P < 0.001). Restricted cubic spline regression analyses suggested nonlinear associa-tions, with apparent threshold effects at approximately 50 g (w6 UK or 4 US alcoholic units)/week for mRNFL and mGCIPL thickness. Significantly stronger alcohol-IOP associations were observed in participants at higher ge-netic susceptibility to glaucoma (Pinteraction < 0.001). Mendelian randomization analyses provided evidence for a causal association with mGCIPL thickness.Conclusions: Alcohol intake was consistently and adversely associated with glaucoma and related traits, and at levels below current United Kingdom (< 112 g/week) and United States (women, < 98 g/week; men, < 196 g/week) guidelines. Although we cannot infer causality definitively, these results will be of interest to people with or at risk of glaucoma and their advising physicians.Financial Disclosure(s): Proprietary or commercial disclosure may be found after the references. Ophthalmology Glaucoma 2023;6:366-379 & COPY; 2022 by the American Academy of Ophthalmology. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
Importance Democratizing artificial intelligence (AI) enables model development by clinicians with a lack of coding expertise, powerful computing resources, and large, well-labeled data sets.Objective To determine whether resource-constrained clinicians can use self-training via automated machine learning (ML) and public data sets to design high-performing diabetic retinopathy classification models.Design, Setting, and Participants This diagnostic quality improvement study was conducted from January 1, 2021, to December 31, 2021. A self-training method without coding was used on 2 public data sets with retinal images from patients in France (Messidor-2 [n = 1748]) and the UK and US (EyePACS [n = 58 689]) and externally validated on 1 data set with retinal images from patients of a private Egyptian medical retina clinic (Egypt [n = 210]). An AI model was trained to classify referable diabetic retinopathy as an exemplar use case. Messidor-2 images were assigned adjudicated labels available on Kaggle; 4 images were deemed ungradable and excluded, leaving 1744 images. A total of 300 images randomly selected from the EyePACS data set were independently relabeled by 3 blinded retina specialists using the International Classification of Diabetic Retinopathy protocol for diabetic retinopathy grade and diabetic macular edema presence; 19 images were deemed ungradable, leaving 281 images. Data analysis was performed from February 1 to February 28, 2021.Exposures Using public data sets, a teacher model was trained with labeled images using supervised learning. Next, the resulting predictions, termed pseudolabels, were used on an unlabeled public data set. Finally, a student model was trained with the existing labeled images and the additional pseudolabeled images.Main Outcomes and Measures The analyzed metrics for the models included the area under the receiver operating characteristic curve (AUROC), accuracy, sensitivity, specificity, and F1 score. The Fisher exact test was performed, and 2-tailed P values were calculated for failure case analysis.Results For the internal validation data sets, AUROC values for performance ranged from 0.886 to 0.939 for the teacher model and from 0.916 to 0.951 for the student model. For external validation of automated ML model performance, AUROC values and accuracy were 0.964 and 93.3% for the teacher model, 0.950 and 96.7% for the student model, and 0.890 and 94.3% for the manually coded bespoke model, respectively.Conclusions and Relevance These findings suggest that self-training using automated ML is an effective method to increase both model performance and generalizability while decreasing the need for costly expert labeling. This approach advances the democratization of AI by enabling clinicians without coding expertise or access to large, well-labeled private data sets to develop their own AI models.
Medical artificial intelligence (AI) offers great potential for recognizing signs of health conditions in retinal images and expediting the diagnosis of eye diseases and systemic disorders(1). However, the development of AI models requires substantial annotation and models are usually task-specific with limited generalizability to different clinical applications(2). Here, we present RETFound, a foundation model for retinal images that learns generalizable representations from unlabelled retinal images and provides a basis for label-efficient model adaptation in several applications. Specifically, RETFound is trained on 1.6million unlabelled retinal images by means of self-supervised learning and then adapted to disease detection tasks with explicit labels. We show that adapted RETFound consistently outperforms several comparison models in the diagnosis and prognosis of sight-threatening eye diseases, as well as incident prediction of complex systemic disorders such as heart failure and myocardial infarction with fewer labelled data. RETFound provides a generalizable solution to improve model performance and alleviate the annotation workload of experts to enable broad clinical AI applications from retinal imaging. RETFound, a foundation model for retinal images that learns generalizable representations from unlabelled images, is trained on 1.6million unlabelled images by self-supervised learning and then adapted to disease detection tasks with explicit labels.
KEYWORDS: Collaborative careindigenous healthoutreach eye carerural eye caretelemedicineteleophthalmology