Purpose:The purpose of this study was to evaluate whether explicitly modeling diabetes mellitus (DM) without diabetic retinopathy (DR) as its own stage enables deep learning (DL) to detect early retinal changes for early risk identification of DR severity spectrum. Methods:We developed 3 DL classification models that explicitly incorporated DM without DR as a distinct stage using 3-class, 4-class, and 6-class staging granularity using 6069 color fundus images from the University of Illinois Chicago Hospital, including 1996 no-DM cases, 1852 DM without DR cases, and 2221 DR cases (516 mild, 220 moderate, 103 severe, and 1382 proliferative DR [PDR]). We developed segmentation models for the optic nerve head (ONH) and retinal vessels to quantify the impact of these regions on classification performance through targeted perturbations. We also examined spatial changes in retinal features across DR stages by measuring the alignment between DL saliency maps and ONH location. Results:For the 3-class model, areas under the curve (AUCs) were 92.2% (no-DM), 80.3% (DM without DR), and 74.1% (mild DR). For the 4-class model, AUCs were 94.0% (no-DM), 71.9% (DM without DR), 61.5% (mild DR), and 80.3% (referable DR). For the 6-class model, AUCs were 94.0% (no-DM), 65.7% (DM without DR), 65.6% (mild DR), 58.9% (moderate DR), 58.6% (severe DR), and 76.1% (PDR). Vessel perturbations reduced performance by 16% to 31% across models, and greater DR severity was associated with increased saliency-to-ONH distances (Pearson r = 0.69-0.72, P < 0.001). Conclusions:Explicitly modeling diabetes without retinopathy improved early-stage discrimination and revealed feature-reliance shifts with DR severity. Vessel- and saliency-based analyses identified subtle retinal changes preceding clinical DR. Translational Relevance:Treating diabetes without retinopathy as its own stage may enhance early DR risk identification and aid development of clinically useful artificial intelligence (AI) tools.
Purpose:The purpose of this study was to evaluate a semi-supervised Mean Teacher (MT) framework for semantic segmentation of cataract surgical images, addressing the challenge of limited labeled data in real-world clinical applications. Methods:We adapted the MT framework for four-class segmentation of the iris, pupil, intraocular lens, and surgical instruments using a small labeled set and 40,000 unlabeled Cataract-1K frames. Performance was assessed using Dice Similarity Coefficient (DSC) and the 95th percentile Hausdorff Distance (HD95). MT was compared with a fully supervised UNet under varying labeled and unlabeled conditions, and additional baselines provided broader context. Ablation studies evaluated noise types and key hyperparameters, including consistency weight (λ) and exponential moving average (EMA) decay (α). Internal validation used Cataract-1K, and external testing was done on CaDIS and CatInstSeg. Results:Our MT model consistently outperformed the supervised UNet baseline. With 100 labeled images and 40,000 unlabeled frames, MT achieved a DSC of 0.76 ± 0.12 compared to UNet with 0.59 ± 0.17 (t = 21.23, P < 0.05). On CaDIS and CatInstSeg, MT reached DSCs of 0.69 ± 0.17 and 0.71 ± 0.20, outperforming UNet at 0.62 ± 0.08 and 0.64 ± 0.23. Optimal performance was observed with λ = 0.1, α = 0.995, and Gaussian noise σ = 15. Conclusions:The MT framework provides an effective semi-supervised solution for surgical image segmentation with limited annotations. Full source code and utility scripts will be released upon acceptance at: https://github.com/mahtabfaraji1/Semi-supervised-segmentation-of-cataract-surgical-images. Translational Relevance:Accurate segmentation of ocular anatomy and instruments supports surgical guidance, intraoperative decision making, and training tools in data-limited clinical environments.
Objectives:Uveal melanoma (UM) is the most common primary intraocular malignancy in adults and carries significant metastatic risk. Early and accurate diagnosis is essential, but challenging due to overlapping clinical features with benign choroidal nevi. Deep learning (DL) offers potential to support early and accessible detection, but model performance is limited by dataset size and quality. This study evaluates data-centric optimisation strategies for DL classification of UM using ultra-widefield (UWF) fundus photography. Methods:This retrospective study analysed UWF fundus photographs from 784 patients (864 images) seen at the University of Illinois Chicago eye clinic. A baseline binary classification model (UM vs. choroidal nevus) was compared with seven models incorporating optimisation strategies across three categories: class addition, dataset augmentation, and enhanced feature selection. Performance was assessed using AUC, F1 score, precision, and recall, with calibration evaluated via expected calibration error. Results:The baseline model achieved an AUC of 0.906 ± 0.032. The top-performing model incorporated healthy retinal controls as an additional class, achieving an AUC of 0.987 ± 0.011 and F1 scores of 0.966 (nevus) and 0.941 (UM). Dataset augmentation approaches yielded minimal performance gain, and Multi-strategy models showed no additive benefit. Conclusions:Data-centric optimisation significantly influences DL performance for UM detection. Three principles emerge: healthy class addition improves specificity and anatomical feature learning; data quality may outweigh quantity; and contextual input tuning is a key model parameter. These findings offer a practical framework for developing clinically robust, physician-supervised AI tools to support early UM triage and reduce diagnostic variability.
Purpose:To investigate the diagnostic utility of the ganglion cell layer (GCL) to inner plexiform layer (IPL) thickness ratio, in differentiating branch retinal vein occlusion (BRVO) from primary open-angle glaucoma (POAG) exhibiting hemifield defects. Given the impact of POAG on the ganglion cell complex (GCC), we hypothesized that POAG wound show disproportionately greater thinning in the GCL and tested if the GCL:IPL ratio could differentiate between these two conditions. Methods:We conducted a retrospective case series of macular OCT images (Spectralis [Heidelberg Engineering, Heidelberg, Germany]) from patients with old BRVO/HRVO and POAG. Inclusion criteria were 1) BRVO/HRVO Diagnosis at least 6 months, without macular edema at the time of imaging, and 2) POAG with an arcuate or altitudinal hemifield defect on Humphrey Visual Field (Carl Zeiss Meditec, Inc., Dublin, CA). Exclusion criteria were the presence of poor-quality OCT scans, history of pan-retinal photocoagulation (PRP), corneal, retinal or neuroophthalmological conditions. Using the Heidelberg automated segmentation analysis of the macular cube, a 20-degree PMB grid was centered over the foveal pit and utilized to generate a 6×10 grid to provide a comprehensive assessment of the retinal layers in the macular region. The GCL:IPL ratio was calculated by dividing the GCL by the corresponding IPL thickness. Calculation of the inner retina:total retina ratio was done in an identical manner, and the average thicknesses and ratios were then compared using the Wilcoxon rank-sum test (MedCalc Statistical Software, Ostend Belgium). Results:Final analysis included 60 eyes of 60 patients (mean age 67 ± 13 years; 57% female; 42% African American, 28% Hispanic, 12% White, 18% as others) diagnosed with old BRVO/HRVO (n = 30) or POAG (n = 30). Patients with POAG had an average visual field mean deviation of -16.22dB ± 5.02. The GCL:IPL ratio was significantly lower in patients with POAG was 0.91 (95% CI [0.85, 0.94]) compared to RVO with 1.14 (95% CI [1.09, 1.17], (P < 0.0001). By adopting the GCL:IPL ratio of less than 1 as a diagnostic marker for POAG, the area under the curve (AUC) was 0.83, with a sensitivity of 90.0% and a specificity of 76.7%. Conclusions:There was disproportionately greater thinning in the GCL compared to the IPL in patients with POAG compared to those with RVO as evidenced by the observed differences in GCL:IPL ratios. Our findings characterized by high AUC and sensitivity suggest that the GCL:IPL ratio has potential to be a marker for distinguishing between old BRVO and POAG with hemifield defects.
Purpose:To develop and evaluate an unsupervised domain adaptation (UDA) framework for glaucoma classification from fundus images that improves the generalizability of deep learning (DL) models across heterogeneous imaging characteristics and clinical settings. Methods:We developed an adversarial UDA framework that adapts a labeled source domain to an unlabeled target domain by jointly optimizing glaucoma classification and domain discrimination. A total of 6906 fundus images were included: 1422 images (full-view and optic nerve head-cropped) derived from 711 fundus photographs of 520 patients from the University of Illinois Chicago and 5484 images from three public datasets (RIMONE-DL, n = 371; REFUGE, n = 259; LAG, n = 4854). Performance was evaluated across multiple source-target domain pairs. Saliency map analysis compared feature utilization between UDA and standard DL models. Results:Across diverse source-target dataset pairs, the proposed UDA framework improved standard DL classification accuracy by up to 40.0% and area under the curve (AUC) by up to 59.8% in extreme domain shift experiments. UDA also reduced the performance gap to the ideal target-trained model by up to 82.6% for accuracy and 38.1% for AUC. When applied to unseen domains, UDA improved classification accuracy by up to 14.7% relative to standard DL. Conclusions:The proposed UDA framework improves cross-domain generalizability of DL-based glaucoma classification from fundus images. Compared with standard DL models, UDA exhibits feature utilization patterns more aligned with an ideal baseline, reducing overreliance on optic nerve head-specific cues and supporting more robust decision-making under domain shift. Translational Relevance:By enabling glaucoma classification without requiring labeled data from new clinical sites, this UDA framework addresses a key barrier to deploying artificial intelligence-based screening tools across diverse real-world ophthalmic settings.
9560 Background: Early differentiation of uveal melanoma (UM) from benign intraocular lesions remains clinically challenging, particularly in regions with limited access to specialized ocular oncologists. Diagnostic uncertainty can delay referral and treatment, highlighting a gap between current practice and the desired goal of timely, accurate early UM detection. Machine learning (ML) offers a promising approach to improving diagnostic support with early detection of UM. This study developed, trained, and optimized eight ML classification models for early UM detection using ultrawidefield fundus (UWF) imaging to evaluate methods to increase clinical utility and readiness of current models. Methods: 1,592 UWF images from 784 patients were retrospectively collected and labeled into four classes: UM (n = 364), choroidal nevus (n = 824), congenital hypertrophy of the retinal pigment epithelium (CHRPE, n = 102), and healthy controls (n = 302). Images with multiple lesions, significant media opacity, or poor tumor visualization were excluded. All models used a ResNet-50 architecture, and a binary UM/nevus classifier was developed as the control. Optimization strategies included 1) class expansion with healthy control and CHRPE images, 2) training dataset augmentation with multiple images per patient and post-treatment images, and 3) a region of interest (ROI) dilation preprocessing pipeline. A 70/15/15 train/validation/test split was used with patient-level stratification to prevent data leakage. The test set was held constant across all models to ensure fair comparison. Performance was assessed using per-class F1 score and area under the curve (AUC) with 5-fold cross-validation. Grad-CAM images were generated to aid in model interpretability. Results: The control model achieved an AUC of 0.94±0.01 and F1 scores of 0.90±0.01 for nevus and 0.80±0.03 for UM. ROI dilation achieved the greatest improvement (AUC 0.99±0.001; nevus F1 0.98±0.01; UM F1 0.95±0.01). Other optimization strategies showed variable impact, with multi-optos yielding the second-best performance (AUC 0.95±0.02; nevus F1 0.91±0.02; UM F1 0.83±0.04). A final model integrating the most effective strategies demonstrated robust performance (AUC 0.95±0.01; nevus F1 0.91±0.03; UM F1 0.82±0.06). Grad-CAM confirmed an appropriate lesion-centered focus. Conclusions: Our findings show that targeted ML optimizations can meaningfully improve performance for early UM detection. While ROI dilation preprocessing had the most performance gain, results varied across optimization pathways, indicating room for refinement. Further evaluation on external datasets is needed to assess generalizability, paving the way for end-stage clinical rollout.
BackgroundPeriorbital measurements such as margin to reflex distances, palpebral fissure height, and scleral show are critical in diagnosing and managing conditions like ptosis and disorders of the eyelid. However, deployment of automated periorbital measurement algorithms in structured research workflows remains limited by the lack of integrated capture and data management infrastructure. ObjectiveWe developed and evaluated Glorbit, a lightweight, browser-based application for automated periorbital distance measurement using artificial intelligence (AI). The objective was to evaluate end-to-end workflow feasibility of the platform under simulated, operator-run conditions. MethodsThe application integrates a DeepLabV3 segmentation model into a modular image processing pipeline with secure, site-specific Google Cloud storage, supporting local preprocessing and cloud upload through Firebase-authenticated logins. The full workflow—metadata entry, facial image capture, segmentation, and upload—was tested. After the session, the participants completed a Likert-style survey. ResultsGlorbit successfully ran on all tested platforms, including laptops, tablets, and mobile phones across major browsers. A total of 15 volunteers were enrolled in this study in which the app completed predefined workflow steps in all simulated, operator-run sessions. The segmentation model produced outputs on all images, and the average session duration was 101.7 (SD 17.5) seconds. Simulated experience scores on a 5-point Likert scale were uniformly high. ConclusionsGlorbit is a cross-platform application that supports structured periorbital image capture and automated inference within a unified workflow. In simulated, operator-run testing, the platform demonstrated successful execution of predefined workflow steps across devices. These findings support the technical feasibility of the system as a research-oriented data collection framework and may inform future evaluations in broader research settings.
Purpose:Uveal melanoma (UM) is the most common intraocular malignancy in adults, with high metastatic risk and poor prognosis. Current screening and triaging methods for melanocytic choroidal tumors face inherent limitations, particularly in regions with limited access to specialized ocular oncologists. This study proposes a reference framework for future research in UM detection. It highlights the trade-offs between different computer vision (CV) approaches based on data availability, computer resources, annotation effort, and clinical applicability. Methods:In total, 864 Optos images of UM, choroidal nevi, and congenital hypertrophy of the retinal pigment epithelium were included in the study. Three CV models-classification, detection, and segmentation-were implemented using a shared ResNet-50 backbone. Performance metrics included area under the curve (AUC) for classification, F1 score for detection, and Dice score for segmentation. An ablation study evaluated robustness to data scarcity. Model interpretability was enhanced through Grad-CAM visualizations and confusion matrices. Results:Classification, detection, and segmentation achieved AUC scores of 94%, 93%, and 95%, respectively. Segmentation showed the best F1 score (89%) and a Dice score of 75%. Classification performed the best with limited data. Conclusions:In high-resource scenarios (100+ images), all models performed similarly, while in low-resource scenarios (<70 images), the classification model outperformed the others. This suggests that simpler models may offer better value for classification tasks with limited resources. Translational Relevance:Given different CV algorithms' different clinical use cases and development costs, this study takes a comparative look at these algorithms for the task of UM analysis on ultra-widefield fundus images.
Purpose: To validate a custom FIJI (ImageJ) program for more reproducible, faster curvilinear periorbital measurements, as compared with 2 custom artificial intelligence–based tools. Design: Combined technical validation and method comparison study. Subjects: Front-facing photographs of 45 cleft palate syndromic patients. Methods: A FIJI (ImageJ)-based macro script tool, OrbitJ (semiautomated), was developed, requiring 15 user input steps to generate 38 measurements per photo. The user outlines the irises to set image scale and then selects points along the lid margins and brow lines. Linear interpolation and fourth-degree polynomial fit lines were used to generate periorbital measurements. This tool was compared against our previously developed deep learning algorithm for periorbital measurements, OrbitMap (automated), another open-source algorithm PeriOrbitAI (automated), and against manual measurements. Four human graders measured 45 photos once both manually and with OrbitJ. Intrarater and interrater measurements were performed with 10 photos manually, 3 in triplicate manually and 5 in triplicate with OrbitJ. Fourteen manual measurements were performed: time per image, iris diameter, margin reflex distance (MRD) 1 and 2, inferior scleral show (ISS), medial, central, and lateral brow height, canthal tilt, vertical dystopia, interpupillary distance, and inner and outer canthal distance (OCD). Main Outcome Measures: The mean absolute error, reliability, bias, and Pearson correlation of periorbital measurements. Results: Analysis was successful in all 45 images for all methods except PeriOrbitAI, which failed on 6 images. For manual and semiautomated intra- and interrater measurements, reliability was considered moderate or better (intraclass correlation coefficient [ICC] >0.5) for all measurements excluding elapsed time. Manual interrater mean absolute error was <1 mm all measures except OCD. Reliability and correlation were high (ICC, Pearson correlation coefficient >0.8) between all OrbitJ and manual measurements. Comparing OrbitMap to manual measurements, ICC and Pearson correlation coefficient were >0.5 except for ISS, borderline for OCD (ICC = 0.51). Reliability between PeriOrbitAI and manual measurements was low (ICC <0.5) except for MRD2 and OCD, and correlation was moderate (Pearson correlation coefficient = 0.49–0.75). The mean analysis time per image was 13.4 ± 4 minutes (manual measurements), 5.4 ± 1.9 minutes (semiautomated OrbitJ), 10.71 ± 1.65 seconds (automated PeriOrbitAI), and 1.45 ± 0.15 seconds (automated OrbitMap) (P < 0.001). Conclusions: Compared with manual measurements, all semiautomated OrbitJ measurements and most automated OrbitMap measurements were reliable. Notably, only 2 PeriOrbitAI measurements were reliable. All methods demonstrated significant time savings over manual measurements. Financial Disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Machine unlearning aims to remove the influence of specific training samples from a trained model without full retraining. While prior work has largely focused on privacy-motivated settings, we recast unlearning as a general-purpose tool for post-deployment model revision. Specifically, we focus on utilizing unlearning in clinical contexts where data shifts, device deprecation, and policy changes are common. To this end, we propose a bilevel optimization formulation of boundary-based unlearning that can be solved using iterative algorithms. We provide convergence guarantees when first-order algorithms are used to unlearn. Our method introduces tunable loss design for controlling the forgetting-retention tradeoff and supports novel model composition strategies that merge the strengths of distinct unlearning runs. Across benchmark and real-world clinical imaging datasets, our approach outperforms baselines on both forgetting and retention metrics, including scenarios involving imaging devices and anatomical outliers. This work establishes machine unlearning as a modular, practical alternative to retraining for real-world model maintenance in clinical applications.
Purpose:Standard deep learning (DL) models often suffer significant performance degradation on out-of-distribution (OOD) data, where test data differs from training data, a common challenge in medical imaging due to real-world variations. Methods:We propose a unified self-censorship framework as an alternative to the standard DL models for glaucoma classification using deep evidential uncertainty quantification. Our approach detects OOD samples at both the dataset and image levels. Dataset-level self-censorship enables users to accept or reject predictions for an entire new dataset based on model uncertainty, whereas image-level self-censorship refrains from making predictions on individual OOD images rather than risking incorrect classifications. We validated our approach across diverse datasets. Results:Our dataset-level self-censorship method outperforms the standard DL model in OOD detection, achieving an average 11.93% higher area under the curve (AUC) across 14 OOD datasets. Similarly, our image-level self-censorship model improves glaucoma classification accuracy by an average of 17.22% across 4 external glaucoma datasets against baselines while censoring 28.25% more data. Conclusions:Our approach addresses the challenge of generalization in standard DL models for glaucoma classification across diverse datasets by selectively withholding predictions when the model is uncertain. This method reduces misclassification errors compared to state-of-the-art baselines, particularly for OOD cases. Translational Relevance:This study introduces a tunable framework that explores the trade-off between prediction accuracy and data retention in glaucoma prediction. By managing uncertainty in model outputs, the approach lays a foundation for future decision support tools aimed at improving the reliability of automated glaucoma diagnosis.
Periorbital measurements such as margin reflex distances (MRD1/2), palpebral fissure height, and scleral show are essential in diagnosing and managing conditions like ptosis and eyelid disorders. We developed Glorbit, a lightweight, browser-based application for automated periorbital distance measurement using artificial intelligence, designed for use in low-resource clinical settings. The app integrates a DeepLabV3 segmentation model into a modular pipeline with secure, site-specific Google Cloud storage. Glorbit supports offline mode, local preprocessing, and cloud upload via Firebase-authenticated logins. We evaluated usability, cross-platform compatibility, and deployment readiness through a simulated enrollment study of 15 volunteers. The app completed the full workflow – metadata entry, image capture, segmentation, and upload – on all tested sessions without error. Glorbit successfully ran on laptops, tablets, and mobile phones across major browsers. The segmentation model succeeded on all images. Average session time was 101.7 seconds (standard deviation: 17.5). Usability survey scores (1-5 scale) were uniformly high: intuitiveness and efficiency (5.0), workflow clarity (4.8), output confidence (4.9), and clinical utility (4.9). Glorbit provides a functional, scalable solution for standardized periorbital measurement in diverse environments. It supports secure data collection and may enable future development of real-time triage tools and multimodal AI-driven oculoplastics. Tool available at: https://glorbit.app
As large language model (LLM) agents increasingly undertake digital work, reliable frameworks are needed to evaluate their real-world competence, adaptability, and capacity for human collaboration. Existing benchmarks remain largely static, synthetic, or domain-limited, providing limited insight into how agents perform in dynamic, economically meaningful environments. We introduce UpBench, a dynamically evolving benchmark grounded in real jobs drawn from the global Upwork labor marketplace. Each task corresponds to a verified client transaction, anchoring evaluation in genuine work activity and financial outcomes. UpBench employs a rubric-based evaluation framework, in which expert freelancers decompose each job into detailed, verifiable acceptance criteria and assess AI submissions with per-criterion feedback. This structure enables fine-grained analysis of model strengths, weaknesses, and instruction-following fidelity beyond binary pass/fail metrics. Human expertise is integrated throughout the data pipeline (from job curation and rubric construction to evaluation) ensuring fidelity to real professional standards and supporting research on human-AI collaboration. By regularly refreshing tasks to reflect the evolving nature of online work, UpBench provides a scalable, human-centered foundation for evaluating agentic systems in authentic labor-market contexts, offering a path toward a collaborative framework, where AI amplifies human capability through partnership rather than replacement.
Objective:We aimed to create and validate a dataset for oculoplastic segmentation and periorbital distance prediction. Design:This was an experimental study. Subjects:Images of faces from 2 open-source datasets were included in this study. Methods:The images were sourced from 2 open-source datasets and cropped to include only the eyes. All images had the iris, sclera, lid, caruncle, and brow segmented by 5 trained annotators. Intergrader reliability analysis was done by having 5 annotators annotate the same 100 images randomly selected after at least a 2-week forgetting period. Intragrader analysis was done by having 5 annotators annotate the same 20 images after a 2-week forgetting period. Three DeepLabV3 segmentation models were trained for segmentation using the datasets following standard procedures. Main Outcome Measures:The quality of the annotations was evaluated by Dice score through intragrader and intergrader experiments. Segmentation models were trained to demonstrate the dataset's utility for deep learning. The Dice score was used to evaluate deep learning models. Results:We annotated 2842 images. Agreement between annotators (intergrader) on a randomly selected subset of 100 images was very high, with an average Dice score of 0.82 ± 0.01. Intragrader analysis also demonstrates that the same grader accurately reproduces annotations with an average Dice score, across all classes, of 0.81 ± 0.08. The average Dice score across all classes of a segmentation network trained on the Chicago Facial dataset, the CelebAMask-HQ dataset, and both combined was 0.90 ± 0.11, 0.81 ± 0.20, and 0.84 ± 0.18, respectively. Conclusions:We have developed a first-of-its-kind dataset for use in oculoplastic and craniofacial segmentation tasks. All the annotations are publicly available for free download. Having access to segmentation datasets designed specifically for oculoplastic surgery will permit more rapid development of clinically useful segmentation networks that can be leveraged for periorbital distance prediction and other downstream tasks. In addition to the annotations, we also provide an open-source toolkit for periorbital distance prediction from segmentation masks, which are available via an application programming interface. The weights of all models have also been open-sourced and are publicly available for use by the community. Financial Disclosures:Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Periorbital distances are critical markers for diagnosing and monitoring a range of oculoplastic and craniofacial conditions. Manual measurement, however, is subjective and prone to intergrader variability. Automated methods have been developed but remain limited by standardized imaging requirements, small datasets, and a narrow focus on individual measurements. We developed a segmentation pipeline trained on a domain-specific dataset of healthy eyes and compared its performance against the Segment Anything Model (SAM) and the prior benchmark, PeriorbitAI. Segmentation accuracy was evaluated across multiple disease classes and imaging conditions. We further investigated the use of predicted periorbital distances as features for disease classification under in-distribution (ID) and out-of-distribution (OOD) settings, comparing shallow classifiers, CNNs, and fusion models. Our segmentation model achieved state-of-the-art accuracy across all datasets, with error rates within intergrader variability and superior performance relative to SAM and PeriorbitAI. In classification tasks, models trained on periorbital distances matched CNN performance on ID data (77–78% accuracy) and substantially outperformed CNNs under OOD conditions (63–68% accuracy vs. 14%). Fusion models achieved the highest ID accuracy (80%) but were sensitive to degraded CNN features under OOD shifts. Segmentation-derived periorbital distances provide robust, explainable features for disease classification and generalize better under domain shift than CNN image classifiers. These results establish a new benchmark for periorbital distance prediction and highlight the potential of anatomy-based AI pipelines for real-world deployment in oculoplastic and craniofacial care.
Here, we examine the latest advances in glaucoma detection through Deep Learning (DL) algorithms using Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA). This study focuses on three aspects of DL-based glaucoma detection frameworks: input data modalities, processing strategies, and model architectures and applications. Moreover, we analyze trends in employing each aspect since the onset of DL in this field. Finally, we address current challenges and suggest future research directions.