Histopathology remains the gold standard for cancer diagnosis and prognosis. With the advent of transcriptome profiling, multi-modal learning combining transcriptomics with histology offers more comprehensive information. However, existing multi-modal approaches are challenged by intrinsic multi-modal heterogeneity, insufficient multi-scale integration, and reliance on paired data, restricting clinical applicability. To address these challenges, we propose a disentangled multi-modal framework with four contributions: 1) to mitigate multi-modal heterogeneity, we decompose WSIs and transcriptomes into tumor and microenvironment subspaces using a disentangled multi-modal fusion module, and introduce a confidence-guided gradient coordination strategy to balance subspace optimization; 2) to enhance multi-scale integration, we propose an inter-magnification gene-expression consistency strategy that aligns transcriptomic signals across WSI magnifications; 3) to reduce dependency on paired data, we propose a subspace knowledge distillation strategy enabling transcriptome-agnostic inference through a WSI-only student model; and 4) to improve inference efficiency, we propose an informative token aggregation module that suppresses WSI redundancy while preserving subspace semantics. Extensive experiments on cancer diagnosis, prognosis, and survival prediction demonstrate our superiority over state-of-the-art methods across multiple settings. Code is available at GitHub.
BACKGROUND:While deep learning has advanced pathological analysis in IgA nephropathy (IgAN), the lack of integrated models that combine multi-label structural identification, Oxford classification, and prognosis prediction remains a significant clinical challenge. METHODS:We developed DeepSNN, a novel deep sequential neural network that serves as a multi-task model trained on multi-center multi-modal renal datasets. The architecture integrates lesion segmentation, glomerular classification, Oxford MEST-C scoring, and prognosis prediction subnets. To ensure interpretability, we conducted visualization experiments and comparative analyses with pathologists' diagnostic patterns. Pathologist comparisons employed Cohen's Kappa with blinded re-evaluation of test and validation sets. RESULTS:DeepSNN demonstrated exceptional lesion identification capabilities across the People's Liberation Army General (PLAG) Hospital dataset (n = 245) and China-Japan Friendship (CJF) Hospital dataset (n = 32), achieving dice coefficients of 0.95 and 0.92, respectively. For Oxford classification, DeepSNN delivered outstanding outcomes with high Kappa values of 0.84, 0.79, 0.87, 0.87, and 0.82 for M, E, S, T, and C scores on the PLAG dataset. Notably, our method outperformed three junior pathologists and achieved comparable performance to senior pathologists across both datasets. During a median follow-up of 47.7 (IQR: 21.9-61.1) months, DeepSNN excelled in prognosis prediction (AUC: 0.810), demonstrating improvement over the International IgA Nephropathy Prediction Tool (IIPT) (AUC: 0.742, ΔAUC = +0.068) in PLAG Hospital dataset (n = 245). Furthermore, visualization maps showed consistent pathological region identification between pathologists and DeepSNN. CONCLUSIONS:DeepSNN successfully integrates multiple diagnostic tasks with performance comparable to senior pathologists, demonstrating substantial potential for streamlining IgAN clinical workflows. This innovation addresses critical gaps in automated renal pathology analysis while maintaining clinical interpretability.
Accurate integration of histological and molecular features is central to modern cancer diagnostics, but it is often hampered by extended turnaround times and resource-intensive parallel workflows, which can delay diagnosis and introduce variability between pathologists. We present CAMPaS (Cross-modal AI for Integrated Molecular Pathology Diagnosis and Stratification), a trustworthy AI framework that jointly predicts glioma histology, molecular markers, and WHO 2021 integrative diagnoses from hematoxylin and eosin-stained slides. CAMPaS introduces a novel architecture purposefully designed to address challenges in real-world translation. Trained on 3,367 patients (6,043 slides) across eight cohorts (six retrospective, two prospective), CAMPaS achieved high diagnostic accuracy (AUC 0.895-0.916 in training; 0.946-0.955 in prospective cohorts) and generalized robustly across diverse settings. Its interpretable cross-modal predictions aligned with expert annotations and genomic profiles, revealing biologically coherent features. Crucially, CAMPaS stratifies prognosis and treatment response, offering a scalable and biologically grounded solution to accelerate precision oncology in routine care. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Trial NCT04007185; CRUK/A19732; B2024-346-01 ### Funding Statement This study was funded by the NIHR Brain Injury MedTech Co-operative and the NIHR Cambridge Biomedical Research Centre (NIHR203312) and Addenbrooke's Charitable Trust and the National Natural Science Foundation of China (grant numbers 62271475 and 82102877). This work presents independent research funded by the National Institute for Health and Care Research (NIHR). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethics Committee of Tiantan Hospital gave ethical approval for this work; Ethics Committee of Sun Yat-sen Hospital gave ethical approval for this work; Ethics Committee of Cambridge University Hospital gave ethical approval for this work I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors after peer-reviewing period
Differentiating between diabetic nephropathy (DN) and non-diabetic renal disease (NDRD) without a kidney biopsy remains a major challenge, often leading to missed opportunities for targeted treatments that could greatly improve NDRD outcomes. To reform the traditional biopsy-all diagnostic paradigm and avoid unnecessary biopsy, we developed a transformer-based deep learning (DL) system for detecting DN and NDRD upon non-invasive multi-modal data of fundus images and clinical characteristics. Our Trans-MUF achieved an AUC of 0.980 (95% CI: 0.979 to 0.980) over the internal retrospective set and also had superior generalizability over a prospective dataset (AUC: 0.989, 95% CI: 0.987 to 0.990) and a multicenter, cross-machine and multi-operator dataset (AUC: 0.932, 95% CI: 0.931 to 0.939). Moreover, the nephrologists‘ diagnosis accuracy can be improved by 21%, through visualization assistance of the DL system. This paper lays a foundation for automatically differentiating DN and NDRD without biopsy. (Registry name: Correlation Study Between Clinical Phenotype and Pathology of Type 2 Diabetic Nephropathy. ID: NCT03865914. Date: 2017-11-30).
Background/aims To identify ocular determinants of iridolenticular contact area (ILCA), a recently introduced swept-source optical coherence tomography (SSOCT) derived parameter, and assess the association between ILCA and angle closure. Methods In this population-based cross-sectional study, right eyes of 464 subjects underwent SSOCT (SS-1000, CASIA, Tomey Corporation, Nagoya, Japan) imaging in the dark. Eight out of 128 cross-sectional images (evenly spaced 22.5° apart) were selected for analysis. Matlab (Matworks, Massachusetts, USA) was used to measure ILCA, defined as the circumferential extent of contact area between the pigmented iris epithelium and anterior lens surface. Gonioscopic angle closure (GAC) was defined as non-visibility of the posterior trabecular meshwork in two or more angle quadrants. Results The mean age of subjects was 62±6.6 years, with the majority being female (65.5%). 143/464 subjects (28.6%) had GAC. In multivariable linear regression analysis, ILCA was significantly associated with anterior chamber width (β=1.03, p=0.003), pupillary diameter (β=−1.9, p<0.001) and iris curvature (β=−17.35, p<0.001). ILCA was smaller in eyes with GAC compared with those with open angles (4.28±1.6 mm 2 vs 6.02±2.71 mm 2 , p<0.001). ILCA was independently associated with GAC (β=−0.03, p<0.001), iridotrabecular contact index (β=−6.82, p<0.001) or angle opening distance (β=0.02, p<0.001) after adjusting for covariates. The diagnostic performance of ILCA for detecting GAC was acceptable (AUC=0.69). Conclusions ILCA is a significant predictor of angle closure independent of other biometric factors and may reflect unique anatomical information associated with pupillary block. ILCA represents a novel biometric risk factor in eyes with angle closure.
The amount of face images has been witnessing an explosive increase in the last decade, where various distortions inevitably exist on transmitted or stored face images. The distortions lead to visible and undesirable degradation on face images, affecting their quality of experience (QoE). To address this issue, this paper proposes a novel Transformer-based method for quality assessment on face images (named as TransFQA). Specifically, we first establish a large-scale face image quality assessment (FIQA) database, which includes 42,125 face images with diversifying content at different distortion types. Through an extensive crowdsource study, we obtain 712,808 subjective scores, which to the best of our knowledge contribute to the largest database for assessing the quality of face images. Furthermore, by investigating the established database, we comprehensively analyze the impacts of distortion types and facial components (FCs) on the overall image quality. Accordingly, we propose the TransFQA method, in which the FC-guided Transformer network (FT-Net) is developed to integrate the global context, face region and FC detailed features via a new progressive attention mechanism. Then, a distortion-specific prediction network (DP-Net) is designed to weight different distortions and accurately predict final quality scores. Finally, the experiments comprehensively verify that our TransFQA method significantly outperforms other state-of-the-art methods for quality assessment on face images.
Multi-modal learning plays a crucial role in cancer diagnosis and prognosis. Current deep learning based multi-modal approaches are often limited by their abilities to model the complex correlations between genomics and histopathology data, addressing the intrinsic complexity of tumour ecosystem where both tumour and microenvironment contribute to malignancy. We propose a biologically interpretative and robust multimodal learning framework to efficiently integrate histopathology images and genomics by decomposing the feature subspace of histopathology images and genomics, reflecting distinct tumour and microenvironment features. To enhance cross-modal interactions, we design a knowledge-driven subspace fusion scheme, consisting of a cross-modal deformable attention module and a gene-guided consistency strategy. Additionally, in pursuit of dynamically optimizing the subspace knowledge, we further propose a novel gradient coordination learning strategy. Extensive experiments demonstrate the effectiveness of the proposed method, outperforming state-of-the-art techniques in three downstream tasks of glioma diagnosis, tumour grading, and survival analysis. Our code is available at https://github.com/helenypzhang/Subspace-Multimodal-Learning.
Multimodal learning, integrating histology images and genomics, promises to enhance precision oncology with comprehensive views at microscopic and molecular levels. However, existing methods may not sufficiently model the shared or complementary information for more effective integration. In this study, we introduce a Unified Modeling Enhanced Multimodal Learning (UMEML) framework that employs a hierarchical attention structure to effectively leverage shared and complementary features of both modalities of histology and genomics. Specifically, to mitigate unimodal bias from modality imbalance, we utilize a query-based cross-attention mechanism for prototype clustering in the pathology encoder. Our prototype assignment and modularity strategy are designed to align shared features and minimizes modality gaps. An additional registration mechanism with learnable tokens is introduced to enhance cross-modal feature integration and robustness in multimodal unified modeling. Our experiments demonstrate that our method surpasses previous state-of-the-art approaches in glioma diagnosis and prognosis tasks, underscoring its superiority in precision neuro-Oncology.
Single domain generalization aims to address the challenge of out-of-distribution generalization problem with only one source domain available. Feature distanglement is a classic solution to this purpose, where the extracted task-related feature is presumed to be resilient to domain shift. However, the absence of references from other domains in a single-domain scenario poses significant uncertainty in feature disentanglement (ill-posedness). In this paper, we propose a new framework, named \textit{Domain Game}, to perform better feature distangling for medical image segmentation, based on the observation that diagnostic relevant features are more sensitive to geometric transformations, whilist domain-specific features probably will remain invariant to such operations. In domain game, a set of randomly transformed images derived from a singular source image is strategically encoded into two separate feature sets to represent diagnostic features and domain-specific features, respectively, and we apply forces to pull or repel them in the feature space, accordingly. Results from cross-site test domain evaluation showcase approximately an ~11.8% performance boost in prostate segmentation and around ~10.5% in brain tumor segmentation compared to the second-best method.
Most recently, the pathology diagnosis of cancer is shifting to integrating molecular makers with histology features. It is a urgent need for digital pathology methods to effectively integrate molecular markers with histology, which could lead to more accurate diagnosis in the real world scenarios. This paper presents a first attempt to jointly predict molecular markers and histology features and model their interactions for classifying diffuse glioma bases on whole slide images. Specifically, we propose a hierarchical multi-task multi-instance learning framework to jointly predict histology and molecular markers. Moreover, we propose a co-occurrence probability-based label correction graph network to model the co-occurrence of molecular markers. Lastly, we design an inter-omic interaction strategy with the dynamical confidence constraint loss to model the interactions of histology and molecular markers. Our experiments show that our method outperforms other state-of-the-art methods in classifying diffuse glioma,as well as related histology and molecular markers on a multi-institutional dataset.
BACKGROUND AND OBJECTIVE:As an advanced technique, immunofluorescence (IF) is one of the most widely-used medical image for nephropathy diagnosis, due to its ease of acquisition with low cost. In practice, the clinically collected IF images are commonly corrupted by blurs at different degrees, mainly because of the inaccurate focus at the acquisition stage. Although deep neural network (DNN) methods achieve the great success in nephropathy diagnosis, their performance dramatically drops over the blurred IF images. This significantly limits the potential of leveraging the advanced DNN techniques in real-world nephropathy diagnosis scenarios.METHODS:This paper first establishes two IF databases with synthetic blurs (IFVB) and real-world blurs (Real-IF) for nephropathy diagnosis, respectively, including 1,659 patients and 6,521 IF images with various degrees of blurs. According to the analysis on these two databases, we propose a deep hierarchical multi-task learning based nephropathy diagnosis (DeepMT-ND) method to bridge the gap between the low-level vision and high-level medical tasks. Specifically, DeepMT-ND simultaneously handles the main task of automatic nephropathy diagnosis, as well as the auxiliary tasks of image quality assessment (IQA) and de-blurring.RESULTS:Extensive experiments show the superiority of our DeepMT-ND in terms of diagnosis accuracy and generalization ability. For instance, our method performs better than nephrologists with at least 15.4% and 6.5% accuracy improvements in IFVB and Real-IF, respectively. Meanwhile, our method also achieves comparable performance in two auxiliary tasks of IQA and de-blurring on blurred IF images.CONCLUSIONS:In this paper, we propose a new DeepMT-ND method for nephropathy diagnosis on blurred IF images. The proposed hierarchical multi-task learning framework provides the new scope to narrow the gap between the low-level vision and high-level medical tasks, and will contribute to nephropathy diagnosis in clinical scenarios. The diagnosis accuracy and generalization ability of DeepMT-ND are experimentally verified to be effective over both synthetic and real-world databases.
Diabetic retinopathy (DR) is a leading cause of permanent blindness among the working-age people. Automatic DR grading can help ophthalmologists make timely treatment for patients. However, the existing grading methods are usually trained with high resolution (HR) fundus images, such that the grading performance decreases a lot given low resolution (LR) images, which are common in clinic. In this paper, we mainly focus on DR grading with LR fundus images. According to our analysis on the DR task, we find that: 1) image super-resolution (ISR) can boost the performance of both DR grading and lesion segmentation; 2) the lesion segmentation regions of fundus images are highly consistent with pathological regions for DR grading. Based on our findings, we propose a convolutional neural network (CNN)-based method for joint learning of multi-level tasks for DR grading, called DeepMT-DR, which can simultaneously handle the low-level task of ISR, the mid-level task of lesion segmentation and the high-level task of disease severity classification on LR fundus images. Moreover, a novel task-aware loss is developed to encourage ISR to focus on the pathological regions for its subsequent tasks: lesion segmentation and DR grading. Extensive experimental results show that our DeepMT-DR method significantly outperforms other state-of-the-art methods for DR grading over three datasets. In addition, our method achieves comparable performance in two auxiliary tasks of ISR and lesion segmentation.
In this paper, we propose a novel task for saliency-guided image translation, with the goal of image-to-image translation conditioned on the user specified saliency map. To address this problem, we develop a novel Generative Adversarial Network (GAN)-based model, called SalG-GAN. Given the original image and target saliency map, SalG-GAN can generate a translated image that satisfies the target saliency map. In SalG-GAN, a disentangled representation framework is proposed to encourage the model to learn diverse translations for the same target saliency condition. A saliency-based attention module is introduced as a special attention mechanism for facilitating the developed structures of saliency-guided generator, saliency cue encoder and saliency-guided global and local discriminators. Furthermore, we build a synthetic dataset and a real-world dataset with labeled visual attention for training and evaluating our SalG-GAN. The experimental results over both datasets verify the effectiveness of our model for saliency-guided image translation.
Given the outbreak of COVID-19 pandemic and the shortage of medical resource, extensive deep learning models have been proposed for automatic COVID-19 diagnosis, based on 3D computed tomography (CT) scans.However, the existing models independently process the 3D lesion segmentation and disease classification, ignoring the inherent correlation between these two tasks.In this paper, we propose a joint deep learning model of 3D lesion segmentation and classification for diagnosing COVID-19, called DeepSC-COVID, as the first attempt in this direction.Specifically, we establish a large-scale CT database containing 1,805 3D CT scans with fine-grained lesion annotations, and reveal 4 findings about lesion difference between COVID-19 and community acquired pneumonia (CAP).Inspired by our findings, DeepSC-COVID is designed with 3 subnets: a cross-task feature subnet for feature extraction, a 3D lesion subnet for lesion segmentation, and a classification subnet for disease diagnosis.Besides, the task-aware loss is proposed for learning the task interaction across the 3D lesion and classification subnets.Different from all existing models for COVID-19 diagnosis, our model is interpretable with fine-grained 3D lesion distribution.Finally, extensive experimental results show that the joint learning framework in our model significantly improves the performance of 3D lesion segmentation and disease classification in both efficiency and efficacy.
Chronic kidney disease is one of the most important causes of mortality worldwide, but a shortage of nephrology pathologists has led to delays or errors in its diagnosis and treatment. Immunofluorescence (IF) images of patients with IgA nephropathy (IgAN), membranous nephropathy (MN), diabetic nephropathy (DN), and lupus nephritis (LN) were obtained from the General Hospital of Chinese PLA. The data were divided into training and test data. To simulate the inaccurate focus of the fluorescence microscope, the Gaussian method was employed to blur the IF images. We proposed a novel multi-task learning (MTL) method for image quality assessment, de-blurring, and disease classification tasks. A total of 1608 patients' IF images were included-1289 in the training set and 319 in the test set. For non-blurred IF images, the classification accuracy of the test set was 0.97, with an AUC of 1.000. For blurred IF images, the proposed MTL method had a higher accuracy (0.94 vs. 0.93, p < 0.01) and higher AUC (0.993 vs. 0.986) than the common MTL method. The novel MTL method not only diagnosed four types of kidney diseases through blurred IF images but also showed good performance in two auxiliary tasks: image quality assessment and de-blurring.
Recent years have witnessed the growing interest in disease severity grading, especially for ocular diseases based on fundus images. The existing grading methods are usually trained with high resolution (HR) images. However, the grading performance decreases a lot given low resolution (LR) images, which are common in practice. In this paper, we mainly focus on diabetic retinopathy (DR) grading with LR fundus images. According to our analysis on the DR task, we find that: 1) image super-resolution (ISR) can boost the performance of DR grading and lesion segmentation; 2) the lesion segmentation regions of fundus images are highly consistent with pathological regions for DR grading. Thus, we propose a deep multi-task learning based DR grading (DeepMT-DR) method for LR fundus images, which simultaneously handles the auxiliary tasks of ISR and lesion segmentation. Specifically, based on our findings, we propose a hierarchical deep learning structure that simultaneously processes the low-level task of ISR, the mid-level task of lesion segmentation and the high-level task of DR grading. Moreover, a novel task-aware loss is developed to encourage ISR to focus on the pathological regions for its subsequent tasks: lesion segmentation and DR grading. Extensive experimental results show that our DeepMT-DR method significantly outperforms other state-of-the-art methods for DR grading over two public datasets. In addition, our method achieves comparable performance in two auxiliary tasks of ISR and lesion segmentation.
Disease forecast is an effective solution to early treatment and prevention for some irreversible diseases, e.g., glaucoma. Different from existing disease detection methods that predict the current status of a patient, disease forecast aims to predict the future state for early treatment. This paper is a first attempt to address the glaucoma forecast task utilizing the sequential fundus images of a patient. Specifically, we establish a database of sequential fundus images for glaucoma forecast (SIGF), which includes an average of 9 images per eye, corresponding to 3,671 fundus images in total. Besides, a novel deep learning method for glaucoma forecast (DeepGF) is proposed based on our SIGF database, consisting of an attention-polar convolution neural network (AP-CNN) and a variable time interval long short-term memory (VTI-LSTM) network to learn the spatio-temporal transition at different time intervals across sequential medical images of a person. In addition, a novel active convergence (AC) training strategy is proposed to solve the imbalanced sample distribution problem of glaucoma forecast. Finally, the experimental results show the effectiveness of our DeepGF method in glaucoma forecast.
Recently, the attention mechanism has been successfully applied in convolutional neural networks (CNNs), significantly boosting the performance of many computer vision tasks. Unfortunately, few medical image recognition approaches incorporate the attention mechanism in the CNNs. In particular, there exists high redundancy in fundus images for glaucoma detection, such that the attention mechanism has potential in improving the performance of CNN-based glaucoma detection. This paper proposes an attention-based CNN for glaucoma detection (AG-CNN). Specifically, we first establish a large-scale attention based glaucoma (LAG) database, which includes 5,824 fundus images labeled with either positive glaucoma (2,392) or negative glaucoma (3,432). The attention maps of the ophthalmologists are also collected in LAG database through a simulated eye-tracking experiment. Then, a new structure of AG-CNN is designed, including an attention prediction subnet, a pathological area localization subnet and a glaucoma classification subnet. Different from other attention-based CNN methods, the features are also visualized as the localized pathological area, which can advance the performance of glaucoma detection. Finally, the experiment results show that the proposed AG-CNN approach significantly advances state-of-the-art glaucoma detection.
Current literature has not considered or provided any data on the permeability of the iris stroma. In this study, we aimed to determine the hydraulic permeability of porcine irides from the isolated stroma. Fifteen enucleated porcine eyes were acquired from the local abattoir. The iris pigment epithelium was scraped off using a pair of forceps and the dilator muscles were pinched off using a pair of colibri toothed forceps. We designed an experimental setup, based on Darcy's law, and consisting of a custom 3D-printed pressure column using acrylonitrile butadiene styrene (ABS) plastic. PBS solution was passed through the iris stroma in a 180 degrees arc shape, with a column height of approximately 204 mm (2000 Pa). Measurements of iris stromal thickness were conducted using optical coherence tomography (OCT). To measure flow rate, we measured the mass (volume) of PBS solution using a mass balance in approximately 1 min. Histology was performed using hematoxylin and eosin (H&E) and anti-smooth muscle antibody (anti-alpha-SMA) for validation. The permeability experiments demonstrated that the iris stroma is a biphasic tissue that allows fluid flow. Our image processing results determined the area of flow to be 7.55 mm(2) and the tissue thickness to be between 180 and 430 mu m. The hydraulic permeability of the porcine stroma, calculated using Darcy's law, was 5.13 +/- 2.39 x 10(-5) mm(2)/Pa.s. Histological and immunochemical studies confirmed that the tissues used for this permeability study were solely iris stroma. Additionally, anti-alpha-SMA staining revealed staining specific for stromal blood vessels, with the notable absence of dilator and sphincter muscle staining. Our study combined experimental microscopic data with the theory of biphasic materials to investigate the hydraulic permeability of the iris stroma. This work will serve as a basis on which to validate future biomechanical studies of human irides with which may ultimately aid disease diagnosis and inform the design of novel treatments.
The past few years have witnessed the great success of applying deep neural networks (DNNs) in computer-aided diagnosis. However, little attention has been paid to provide pathological evidence in the existing DNNs for medical diagnosis. In fact, feature visualization in DNNs is able to help understanding how the computer make decisions, and thus it shows promise on finding pathological evidence from computer-aided diagnosis. In this paper, we propose a novel pathologya-ware visualization approach for DNN-based glaucoma classification, which is used to locate the pathological evidence from fundus images for glaucoma. Besides, we apply the visualization framework to the glaucoma images synthesis task, through which specific pathological areas of synthesized images can be enhanced. Finally, experimental results show that the visualization heat maps can pinpoint different glaucoma pathologies with high accuracy, and that the generated glaucoma images are more pathophysiologically clear in rim loss (RL) and retinal neural fiber layer damage (RNFLD), which is verified by the ophthalmologist.