PRECIS:Artificial intelligence-derived macular thinning patterns were associated with central visual field progression in glaucoma and outperformed global thickness metrics in predicting progression across disease severities. PURPOSE:To provide spatial patterns of ganglion cell complex thickness and assess their associations with central visual field progression in glaucoma. METHODS:Macular patterns from the ganglion cell complex were determined using an artificial intelligence algorithm termed archetypal analysis (AA). The diagnostic accuracy of spatial patterns for detecting 10-2 central visual field progression in eyes with at least five 10-2 visual field tests was calculated and compared with the mean global ganglion cell complex thickness. Eyes with progression on either of 2 trend-based methods (significant MD slope less than -0.5 dB/y or clustered pointwise linear regression) were classified as "progressors." RESULTS:A total of 4031 macular scans of 1093 eyes (611 patients) were included, with a mean (SD) age of 67.8 (12.7) years. Eleven distinct spatial patterns were identified. While the macular vulnerable zone was preferentially affected in 4 patterns, most of the less vulnerable zones were preserved. The AA models at baseline achieved AUROC [0.73 (95% CI, 0.62-0.84)] and outperformed global ganglion cell complex thickness [0.55 (95% CI, 0.46-0.61), P =0.01] for predicting central VF progression in eyes with early disease at baseline. The AA models AUROC [0.70 (95% CI, 0.59-0.80)] also outperformed ganglion cell complex thickness [0.55 (95% CI, 0.48-0.60), P =0.02] for predicting central VF progression across all severities. CONCLUSIONS:Using unsupervised artificial intelligence, characteristic patterns of macular thinning were identified and associated with progression in the central visual field. Spatial macular pattern analysis may enhance individualized care and improve risk stratification for those at risk of central VF damage.
PURPOSE:To evaluate associations between parapapillary choriocapillaris microvascular dropout (MvD) and optical coherence tomography (OCT)-detected deep optic nerve head (ONH) structures in glaucomatous eyes with and without myopia. DESIGN:Cross-sectional study from clinical trial data. METHODS:394 eyes from 262 patients with primary open-angle glaucoma (POAG) and glaucoma suspects were stratified into three groups of no myopia (axial length (AL)<24 mm; n = 144), mild myopia (24 mm ≤ AL < 26 mm; n = 174), and high myopia (AL ≥ 26 mm; n = 76). Spectralis ONH OCT radial B-scans were acquired relative to the Foveal-Bruch's Membrane Opening (FoBMO) axis. Bruch's Membrane Opening (BMO) and anterior scleral canal opening (ASCO) were manually segmented, and their size and shape were calculated. BMO/ASCO offset magnitude, neural canal obliqueness, and neural canal minimum cross-sectional area (NCMCA) were measured. The presence, area, and angular circumference of juxtapapillary MvD were evaluated using OCT-angiography en face choroidal images and B-scans. RESULTS:The MvD area (95% CI) was significantly greater in highly myopic eyes (0.38 [0.30, 0.47] mm²), compared with mild myopia (0.33 [0.27, 0.39] mm²) and no myopia (0.21 [0.14, 0.27] mm²) (P = .002). The MvD angular circumference was also significantly larger in mild myopia (75.4 [64.0, 86.9]°), followed by high myopia (74.5 [58.0, 90.9]°) and no myopia (52.6[39.9, 65.3]°) (P = .017). The highly myopic group showed a significantly larger BMO area, NCMCA ovality index, BMO/ASCO offset magnitude, and neural canal obliqueness, along with smaller NCMCA, compared to the other groups (all P < .01). In multivariable analysis, NCMCA, NCMCA ovality index, BMO/ASCO offset magnitude, and neural canal obliqueness were significantly associated with both MvD presence (all P < .05) and MvD area (all P < .05). Additionally, NCMCA ovality index and neural canal obliqueness were associated with MvD angular circumference (P = .01 and P = .004, respectively). CONCLUSIONS:In myopic POAG eyes, the presence and area of MvD were associated with NCMCA, NCMCA ovality index, BMO/ASCO offset magnitude, and neural canal obliqueness, whereas MvD angular circumference was associated only with NCMCA ovality index and neural canal obliqueness. Evaluating choriocapillaris MvD alongside deep ONH structural alterations may provide clinical insights into the pathogenesis of glaucoma in myopia.
Objective:To develop an explainable multimodal large language model (MM-LLM) that (1) screens optic nerve head (ONH) OCT circle scans for quality and (2) generates structured clinical reports that include glaucoma diagnosis and sector-wise retinal nerve fiber layer (RNFL) thinning assessments. Design:A retrospective cohort study using longitudinal data from the Diagnostic Innovations in Glaucoma Study and the African Descent and Glaucoma Evaluation Study. Participants:A total of 43 849 Spectralis circumpapillary B-scans centered on the ONH from 1310 subjects, including 1331 glaucomatous and 867 healthy eyes. Methods:An MM-LLM (Llama 3.2 Vision-Instruct model) was fine-tuned to generate clinical descriptions of OCT imaging data. Training data included paired OCT images and automatically generated, structured clinical reports that described global and sectoral RNFL thinning. Poor-quality scans were labeled as unusable and paired with a fixed refusal statement. The model was evaluated on a held-out test set for 3 tasks: quality assessment, glaucoma detection, and RNFL thinning classification across 7 anatomical sectors. Evaluation metrics included accuracy, sensitivity, specificity, precision, and F1-score. Model description quality was also evaluated using standard text evaluation metrics (BLEU, ROUGE, METEOR, and BERTScore). Main Outcome Measures:Diagnostic accuracy metrics for each task; text evaluation metrics for description quality. Results:The model achieved 0.90 accuracy and 0.98 specificity for quality triage. For glaucoma detection, accuracy was 0.86 (sensitivity 0.93, specificity 0.65, and F1-score 0.91). Retinal nerve fiber layer thinning prediction accuracy ranged from 0.83 to 0.94, with the highest performance in global, temporal, temporal superior, and temporal inferior sectors. Text generation scores (mean ± standard deviation) showed strong alignment with reference reports (BLEU: 0.82 ± 0.19; ROUGE-1: 0.94 ± 0.08; ROUGE-2: 0.87 ± 0.17; ROUGE-L: 0.92 ± 0.11; BERTScore-F1: 0.99 ± 0.02). Stratified analysis revealed better RNFL thinning detection in moderate-to-advanced glaucoma cases, especially in temporal sectors, while performance in nasal regions was better for mild cases. Conclusions:The fine-tuned MM-LLM generated accurate clinical descriptions based on OCT imaging. The model achieved high accuracy in identifying image quality issues and detecting glaucoma. The model provided sectoral descriptions of RNFL thinning to support clinical OCT evaluation. This approach shows potential as a scalable tool for clinical decision support, but further validation across additional datasets is needed. Financial Disclosures:Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow had three steps: (1) LLM initial assessment; (2) function calling to invoke specialized tools for image quality (QAModel, FundaQ-8), glaucoma classification (SwinV2-Tiny), and optic disc/cup segmentation (SegFormer-B0); and (3) LLM reflection integrating the initial impression with tool outputs. Two LLMs (Gemini 2.5 Flash, GPT-5.4 mini) were evaluated on two public datasets (ORIGA, n=100; RIM-ONE-v3, n=100) under uncropped and cropped fields of view; all images were independently graded by a masked fellowship-trained glaucoma specialist. The agentic workflow improved classification accuracy by 16 to 47 percentage points across all conditions, reaching within 6 points of the specialist; on RIM-ONE-v3 the best configurations matched the specialist accuracy of 88
Background: Interpretation of optical coherence tomography (OCT) imaging requires clinicians to translate complex imaging findings into clinical narrative reports, contributing substantially to documentation burden and physician burnout. Our goal is to develop and evaluate a fine-tuned multimodal large language model (LLM) for automated generation of clinician-style OCT reports to document retinal disease. Methods: We curated 80,102 macula OCT B-scan images paired with corresponding clinician-authored reports from 2,208 patients at a tertiary retina center (2011–2024). A multimodal LLM (Llama 3.2 Vision-Instruct) was fine-tuned to generate clinician-style reports from OCT B-scans. Performance was evaluated based on assessments by three masked expert graders. A set of 250 AI-generated and 250 clinician-authored reports were assessed for predicted source (AI vs. clinician) and three quality metrics (correctness, completeness, and conciseness) using a 5-point Likert scale. Intra- and inter-grader agreements as well as quantitative text similarity metrics (BLEU, ROUGE, METEOR, BERTScore) were also computed. Findings: In masked expert evaluations, graders demonstrated near-chance accuracy (49–55%) in distinguishing AI-generated from clinician-authored reports. Experts graded AI-generated reports as less accurate than clinician-authored reports (3.75 vs. 3.99, p=0.0002), but there were no significant differences between AI and clinician reports for completeness (4.13 vs. 4.23, p=0.08) or conciseness (4.21 vs. 4.13, p=0.14). Experts also graded reports that they perceived as clinician-authored significantly higher regardless of true source (p<0.0001). Inter-grader rater agreement was highest for conciseness (weighted 𝝹 0.52–0.58) and followed by correctness (weighted 𝝹 0.47–0.51) and completeness (weighted 𝝹 0.33–0.49).The AI-generated reports achieved high semantic similarity with clinician-authored reports (BERTScore F1 0.89) despite lower lexical overlap (BLEU 0.09, ROUGE-1 F1 0.36). Interpretation: A fine‑tuned multimodal LLM can generate retinal OCT reports that are indistinguishable from clinician-authored reports and comparable in correctness, completeness and conciseness to those written by clinicians. These findings suggest that LLMs can assist with clinical documentation, reduce reporting burden, and improve workflow efficiency in high-volume ophthalmic imaging clinics.
Purpose:To evaluate the accuracy of a three-dimensional (3D) deep learning (3D DL) and 3D cross domain deep learning (3D CD-DL) classifiers compared to standard macular ganglion cell-inner plexiform layer (GCIPL) thickness measurements for classifying eyes with glaucoma using optical coherence tomography (OCT). Methods:A total of 502 primary open-angle glaucoma eyes from 295 patients and 119 healthy eyes from 63 individuals were included. Two classifiers were compared: (1) a 3D DL model trained on Spectralis macular OCT and applied to Spectralis macular OCT images and (2) 3D CD-DL model trained on synthetic Spectralis images generated from 3D Cirrus macular OCT using Cycle-consistent adversarial networks (CycleGAN) and applied to real Spectralis macula OCT images. An additional 100 different eyes (50 Cirrus, 50 Spectralis) were used to train the CycleGAN. Age, axial length, and disc area adjusted area under the receiver operating curves (AUROC) were used to compare model accuracy. Results:Adjusted AUROC for 3D DL model was 0.92 (95% confidence interval [CI], 0.85-0.95). This was significantly higher than global GCIPL thickness (0.83 [0.78-0.85], p ≤ 0.001) but similar to 3D CD-DL (0.91 [0.84-0.95], P = 0.45). Using only early glaucoma eyes (mean deviation ≥ -3.0 dB), the 3D DL model showed significantly higher diagnostic accuracy (0.90 [0.84-0.94]) compared to global GCIPL thickness (0.80 [0.76-0.82], P ≤ 0.001) but similar to the 3D CD-DL model (0.90 [0.83-0.93], P = 0.51). Conclusions:The 3D DL classifier showed significantly higher diagnostic accuracy than global GCIPL thickness but was similar in performance to the 3D CD-DL classifier. By using synthetic data and diverse training sets, cross-domain learning produces robust, generalizable models across different imaging devices as demonstrated by the comparable accuracy of the 3D CD-DL and device-specific 3D DL models. More data from other OCT devices are needed to further validate these findings. Translational Relevance:The 3D Deep learning models significantly surpass traditional GCIPL thickness measurements for accurately detecting glaucoma. The cross-domain model closely matches the performance of the device-specific model in glaucoma classification potentially reducing the need for device-specific models in clinical practice.
Mentor: Robert N. Weinreb Program: Ophthalmology Type: Original Research Background: Preserving central retinal ganglion cells is crucial in glaucoma management, as macular damage in the 10-2 visual field (VF) impacts quality of life. The purpose of this study is to characterize spatial patterns of macular ganglion cell complex (GCC) thickness and assess their associations with central VF progression in glaucoma. Methods: A total of 4219 macula scans of 1116 eyes (613 participants) were included in this retrospective cohort study (NCT00221897, NCT00221923). Macular patterns from GCC were determined by an artificial intelligence (AI) unsupervised algorithm (archetypal analysis (AA)). Diagnostic accuracy of spatial patterns to detect 10-2 central VF progression in 262 eyes (175 patients) with a minimum of five 10-2 VF tests was calculated and compared with mean naive GCC thickness. VF progression was defined based on pointwise linear regression and trend-based methods. Results: AA identified 11 distinct spatial patterns across different glaucoma stages ( Figure 1). The AA model at baseline achieved Area Under the Receiver Operating Characteristic curve (AUROC) of 0.73 (95% CI [0.60 – 0.84]) and outperformed mean GCC thickness (0.55, 95% CI [0.46 – 0.61], P = 0.006) for predicting central VF progression tin eyes with early stages of glaucoma. The AA model (AUROC: 0.70 95% CI [0.59 – 0.80]) outperformed mean GCC (0.55 [95% CI 0.48 – 0.60], P = 0.012) thickness for predicting central VF progression across all severities. Figure 1 Eleven macula ocular coherence tomography (OCT) patterns for varying glaucoma severity, as determined by unsupervised archetypal analysis. The percentage above each archetype indicates the respective average decomposition weight in our dataset for each pattern. Conclusion: AI-driven spatial GCC patterns exhibit distinct characteristics associated with central VF progression in glaucoma. Macular patterns may enhance patient care and risk stratification.
Objective:To develop an explainable multimodal large language model (MM-LLM) that (1) screens optic nerve head (ONH) OCT circle scans for quality and (2) generates structured clinical reports that include glaucoma diagnosis and sector-wise retinal nerve fiber layer (RNFL) thinning assessments. Design:Retrospective cohort study using longitudinal data from the Diagnostic Innovations in Glaucoma Study (DIGS) and the African Descent and Glaucoma Evaluation Study (ADAGES). Participants:43,849 Spectralis ONH OCT circle scans from 1,310 subjects, including 1,331 glaucomatous and 867 healthy eyes. Methods:A MM-LLM (Llama 3.2 Vision-Instruct model) was fine-tuned to generate clinical descriptions of OCT imaging data. Training data included paired OCT images and automatically generated, structured clinical reports that described global and sectoral RNFL thinning. Poor-quality scans were labeled as unusable and paired with a fixed refusal statement. The model was evaluated on a held-out test set for three tasks: quality assessment, glaucoma detection, and RNFL thinning classification across seven anatomical sectors. Evaluation metrics included accuracy, sensitivity, specificity, precision, and F1-score. Model description quality was also evaluated using standard text evaluation metrics (BLEU, ROUGE, METEOR, BERTScore). Results:The model achieved 0.90 accuracy and 0.98 specificity for quality triage. For glaucoma detection, accuracy was 0.86 (sensitivity 0.91, specificity 0.73, F1-score 0.91). RNFL thinning prediction accuracy ranged from 0.83 to 0.94, with highest performance in global and temporal sectors. Text generation scores (mean ± SD) showed strong alignment with reference reports (BLEU: 0.82 ± 0.19; ROUGE-1: 0.94 ± 0.08; ROUGE-2: 0.87 ± 0.17; ROUGE-L: 0.92 ± 0.11; BERTScore-F1: 0.99 ± 0.02). Stratified analysis revealed better RNFL thinning detection in moderate-to-advanced glaucoma cases, especially in temporal sectors, while performance in nasal regions was better for mild cases. Conclusions:The fine-tuned MM-LLM generated accurate clinical descriptions based on OCT imaging. The model achieved high accuracy in identifying image quality issues and detecting glaucoma. The model also provided sectoral descriptions of RNFL thinning to help support clinical OCT evaluation. This approach shows potential as a scalable tool for clinical decision support, but further validation across additional datasets is needed.
Purpose: The aim is to assess GPT-4V's (OpenAI) diagnostic accuracy and its capability to identify glaucoma-related features compared to expert evaluations. Design: Evaluation of multimodal large language models for reviewing fundus images in glaucoma. Subjects: A total of 300 fundus images from 3 public datasets (ACRIMA, ORIGA, and RIM-One v3) that included 139 glaucomatous and 161 nonglaucomatous cases were analyzed. Methods: Preprocessing ensured each image was centered on the optic disc. GPT-4's vision-preview model (GPT-4V) assessed each image for various glaucoma-related criteria: image quality, image gradability, cup-to-disc ratio, peripapillary atrophy, disc hemorrhages, rim thinning (by quadrant and clock hour), glaucoma status, and estimated probability of glaucoma. Each image was analyzed twice by GPT-4V to evaluate consistency in its predictions. Two expert graders independently evaluated the same images using identical criteria. Comparisons between GPT-4V's assessments, expert evaluations, and dataset labels were made to determine accuracy, sensitivity, specificity, and Cohen kappa. Main Outcome Measures: The main parameters measured were the accuracy, sensitivity, specificity, and Cohen kappa of GPT-4V in detecting glaucoma compared with expert evaluations. Results: GPT-4V successfully provided glaucoma assessments for all 300 fundus images across the datasets, although approximately 35% required multiple prompt submissions. GPT-4V's overall accuracy in glaucoma detection was slightly lower (0.68, 0.70, and 0.81, respectively) than that of expert graders (0.78, 0.80, and 0.88, for expert grader 1 and 0.72, 0.78, and 0.87, for expert grader 2, respectively), across the ACRIMA, ORIGA, and RIM-ONE datasets. In Glaucoma detection, GPT-4V showed variable agreement by dataset and expert graders, with Cohen kappa values ranging from 0.08 to 0.72. In terms of feature detection, GPT-4V demonstrated high consistency (repeatability) in image gradability, with an agreement accuracy of ≥89% and substantial agreement in rim thinning and cup-to-disc ratio assessments, although kappas were generally lower than expert-to-expert agreement. Conclusions: GPT-4V shows promise as a tool in glaucoma screening and detection through fundus image analysis, demonstrating generally high agreement with expert evaluations of key diagnostic features, although agreement did vary substantially across datasets. Financial Disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
PurposeTo evaluate the diagnostic accuracy of a deep learning autoencoder-based model utilizing regions of interest (ROI) from optical coherence tomography (OCT) texture enface images for detecting glaucoma in myopic eyes.MethodsThis cross-sectional study included a total of 453 eyes from 315 participants from the multi-center "Swept-Source OCT (SS-OCT) Myopia and Glaucoma Study", composed of 268 eyes from 168 healthy individuals and 185 eyes from 147 glaucomatous individuals. All participants underwent swept-source optical coherence tomography (SS-OCT) imaging, from which texture enface images were constructed and analyzed. The study compared four methods: (1) global RNFL thickness, (2) texture enface image, (3) a single autoencoder model trained only on healthy eyes, and (4) a dual autoencoder model trained on both healthy and glaucomatous eyes. Diagnostic accuracy was assessed using the area under the receiver operating curves (AUROC) and precision recall curves (AUPRC).ResultsThe dual autoencoder model achieved the highest AUROC (95% CI) (0.92 [0.88, 0.95]), significantly outperforming the single autoencoder model trained only on healthy eyes (0.86 [0.83, 0.88], p = 0.01), the global RNFL thickness model (0.84 [0.80, 0.86], p = 0.003), and the texture enface model (0.83 [0.79, 0.85], p = 0.005). Using AUPRC (95% CI), the dual autoencoder model (0.86 [0.83, 0.89]) also outperformed the single autoencoder model trained only on healthy eyes (0.80 [0.78, 0.82], p = 0.02), the global RNFL thickness model (0.74 [0.70, 0.76], p = 0.001), and the texture enface model (0.71 [0.68, 0.73], p<0.001). No significant difference was observed between the global RNFL thickness measurement and the texture enface measurement (p = 0.47).DiscussionThe dual autoencoder model, which integrates reconstruction errors from both healthy and glaucomatous training data, demonstrated superior diagnostic accuracy compared to the single autoencoder model, global RNFL thickness and texture enface-based approaches. These findings suggest that deep learning models leveraging ROI-based reconstruction error from texture enface images may enhance glaucoma classification in myopic eyes, providing a robust alternative to conventional structural thickness metrics.
Purpose:To compare the performance of unimodal and multimodal implementation of the self-supervised learning model RETFound in detecting glaucoma using color fundus photographs (CFPs) and OCT images, and to assess its generalizability across different ethnicities, age groups, and disease severities. Design:Evaluation of a diagnostic technology. Subjects Participants and Controls:Fourteen thousand five hundred ten CFPs and 32 640 OCTs from 1948 eyes of 1098 participants (60.8% glaucoma, 39.2% healthy) from the Diagnostic Innovations in Glaucoma Study and the African Descent and Glaucoma Evaluation Study were included. Glaucoma was defined as photograph-based glaucomatous optic neuropathy with or without repeatable glaucoma visual field damage. Methods:A multimodal RETFound model was developed using paired CFPs and OCT images. The model was compared to unimodal RETFound models using solely CFP or OCT images. Performance was also stratified by race (Black vs. White), age (<60 vs. ≥60 years), and disease severity (mild vs. moderate-to-severe glaucoma). Main Outcome Measures:Diagnostic accuracy of unimodal and multimodal RETFound models using CFP and OCT for detecting glaucoma was assessed using the area under the receiver operating characteristic curve (AUC), precision, and recall. Results:The multimodal model for glaucoma detection achieved an AUC of 0.94 (95% confidence interval: 0.91-0.97), significantly outperforming the CFP unimodal model (AUC 0.86 [95% confidence interval: 0.81-0.89], P < 0.001) but not the OCT unimodal model (AUC 0.93 [95% confidence interval: 0.90-0.96], P = 0.47). Precision and recall were higher (0.96 and 0.87, respectively) for the multimodal model compared with the CFP model (0.92 and 0.69) across all subgroups. No significant differences based on race or age were found in either unimodal or multimodal glaucoma detection models. All models exhibited better performance in detecting moderate-to-severe glaucoma than mild glaucoma, with significant differences in the unimodal CFP (P = 0.002) and OCT (P = 0.005) models. Conclusions:The multimodal RETFound model demonstrated improved diagnostic ability compared with the CFP unimodal model but did not significantly outperform the OCT unimodal model in glaucoma detection. As clinical implementation of a unimodal artificial intelligence (AI) model is easier than a multimodal counterpart, our results suggest unimodal OCT AI models may be sufficient for detecting glaucoma. Financial Disclosures:Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
This study aims to develop deep learning (DL) models to predict the retinal nerve fiber layer (RNFL) thickness changes in glaucoma, facilitating the early diagnosis and monitoring of disease progression. Using the longitudinal data from two glaucoma studies (Diagnostic Innovations in Glaucoma Study (DIGS) and African Descent and Glaucoma Evaluation Study (ADAGES)), we constructed models using optical coherence tomography (OCT) scans from 251 participants (437 eyes). The models were trained to predict the RNFL thickness at a future visit based on previous scans. We evaluated four models: linear regression (LR), support vector regression (SVR), gradient boosting regression (GBR), and a custom 1D convolutional neural network (CNN). The GBR model achieved the best performance in predicting pointwise RNFL thickness changes (MAE = 5.2 μm, R2 = 0.91), while the custom 1D CNN excelled in predicting changes to average global and sectoral RNFL thickness, providing greater resolution and outperforming the traditional models (MAEs from 2.0–4.2 μm, R2 from 0.94–0.98). Our custom models used a novel approach that incorporated longitudinal OCT imaging to achieve consistent performance across different demographics and disease severities, offering potential clinical decision support for glaucoma diagnosis. Patient-level data splitting enhances the evaluation robustness, while predicting detailed RNFL thickness provides a comprehensive understanding of the structural changes over time.
Purpose To compare the performance of unimodal and multimodal implementation of the self-supervised learning model RETFound in detecting glaucoma using color fundus photographs (CFP) and optical coherence tomography (OCT) images, and to assess its generalizability across different ethnicities, age groups, and disease severities. Design Evaluation of a diagnostic technology Subjects, Participants, and Controls 14,510 CFPs and 32,640 OCTs from 1,948 eyes of 1,098 participants (60.8% glaucoma, 39.2% healthy) from the Diagnostic Innovations in Glaucoma Study (DIGS) and the African Descent and Glaucoma Evaluation Study (ADAGES) were included. Glaucoma was defined as photograph-based glaucomatous optic neuropathy (GON) with or without repeatable glaucoma visual field damage (GVFD). Methods A multimodal RETFound model was developed using paired CFPs and OCT images. The model was compared to unimodal RETFound models using solely CFP or OCT images. Performance was also stratified by race (Black vs. White), age (<60 vs. ≥60 years), and disease severity (mild vs. moderate-to-severe glaucoma). Main Outcome Measures Diagnostic accuracy of unimodal and multimodal RETFound models using CFP and OCT for detecting glaucoma was assessed using the area under the receiver operating characteristic curve (AUC), precision, and recall. Results The multimodal model for glaucoma detection achieved an AUC of 0.94 (95% CI: 0.91–0.97), significantly outperforming the CFP unimodal model (AUC 0.86 [95% CI: 0.81–0.89], p < 0.001) but not the OCT unimodal model (AUC 0.93 [95% CI: 0.90–0.96], p = 0.47). Precision and recall were higher (0.96 and 0.87, respectively) for the multimodal model compared to the CFP model (0.92 and 0.69) across all subgroups. No significant differences based on race or age were found in either unimodal or multimodal glaucoma detection models. All models exhibited better performance in detecting moderate-to-severe glaucoma than mild glaucoma, with significant differences in the unimodal CFP (p = 0.002) and OCT (p = 0.005) models. Conclusions The multimodal RETFound model demonstrated improved diagnostic ability compared to the CFP unimodal model but did not significantly outperform the OCT unimodal model in glaucoma detection. As real-world clinical implementation of a unimodal AI model is easier than a multimodal counterpart, our results suggest unimodal OCT AI models may be sufficient for detecting glaucoma.
PRÉCIS:Larger choriocapillaris microvasculature dropout area and wider angular circumference are significantly associated with 24-2C central visual field damage in primary open angle glaucoma eyes with and without axial myopia. PURPOSE:To evaluate the relationship between a juxtapapillary choriocapillaris microvasculature dropout (MvD) and central visual field (VF) damage in primary open angle glaucoma (POAG) patients with or without axial myopia. METHODS:This cross-sectional study included 125 patients with POAG or glaucoma suspects stratified into no axial myopia (axial length (AL) ≤24 mm; 46 eyes), mild axial myopia (24 mm < AL ≤26 mm; 81 eyes), and high axial myopia (AL >26 mm; 59 eyes). Presence, area, and angular circumference of juxtapapillary MvD were evaluated on OCT-A en-face choroidal images and B-scans. Perimetry was conducted using the 24-2C and 10-2 Humphrey program. RESULTS:Mean 24-2C VF mean deviation was significantly worse in eyes with MvD compared with eyes without MvD across all groups (all P <0.042). Central VF defects detected in the 24-2C and 10-2 VF tests were significantly more prevalent among eyes with MvD (68.3% and 81.7%, respectively) compared with eyes without MvD (19.0% and 38.1%, respectively) ( P <0.001) in the mild axial myopia group. In multivariable analysis, larger MvD area ( P =0.014) and wider MvD angular circumference ( P =0.006) were significantly associated with higher likelihood of the presence of 24-2C central VF damage in overall cohort. CONCLUSIONS:MvD area and angular circumference are significantly associated with central VF damage detected by VF 24-2C in POAG eyes with and without axial myopia. Choriocapillaris MvD assessment shows promise for identifying POAG patients with a higher risk of having central VF defects and may provide clinical insights into the pathogenesis of glaucoma in myopia.
PRÉCIS:Artificial intelligence applied to OCTA images demonstrated high accuracy in estimating 24-2 visual field maps by leveraging information from the parapapillary area. PURPOSE:To develop deep learning (DL) models estimating 24-2 visual field (VF) maps from optical coherence tomography angiography (OCTA) optic nerve head (ONH) en face images. METHODS:A total of 3148 VF OCTA pairs were collected from 994 participants (1684 eyes). DL models were trained using radial peripapillary capillary (RPC), superficial, and choroidal, as well as combined ONH VD layers, to estimate 24-2 mean deviation (MD), pattern standard deviation (PSD), 52 total deviation (TD), and pattern deviation (PD) values and compared with a linear regression (LR) model. Model accuracy was assessed by calculating mean absolute error (MAE) and R (Pearson correlation coefficient) between estimated and actual VF values. RESULTS:DL models outperformed LR estimates for the estimation of VF values using individual and combined layers ( P <0.001). For example, in the estimation of MD using RPC, DL achieved an R of 0.79 and MAEs of 1.77 dB. Average estimated TDs using RPC had R of 0.63 and MAEs of 3.08 dB. DL estimation using combined layers slightly improved the choroid in the estimation of MD ( P <0.01) and had comparable performance with RPC and superficial layers. It also slightly improved RPC, superficial and choroidal layer in the estimation of TDs ( P <0.01). CONCLUSIONS:DL models from OCTA images demonstrated high accuracy in estimating 24-2 VF maps by leveraging information from ONH layers. By extending the application of DL to OCTA images using RPC or superficial layers, it may be possible to reduce the frequency of VF testing to individual patients.