Parkinson’s disease (PD) is a prevalent neurodegenerative disorder globally. The eye’s retina is an extension of the brain, and clinical evidence has suggested the great potential of retinal pathology as surrogate biomarkers for early PD diagnosis. In particular, recent studies have shown that texture features extracted from retinal layers based on optical coherence tomography (OCT) images are strongly associated with PD-related retinal pathology. Additionally, frequency domain learning techniques can improve the representational capabilities of deep neural networks (DNNs) by decomposing frequency components that involve rich texture features, which remain underexplored for automated early PD diagnosis in OCT. To bridge this gap, we propose an Adaptive Wavelet Filter (AWF) that serves as the Practical Texture Amplifier, which fully leverages the merits of texture features from the retinal pathology view. Specifically, AWF first enhances feature map diversities and refines feature representations via channel mixer, then emphasizes informative texture feature representations with the well-designed adaptive wavelet filtering token mixer with the aid of frequency domain learning. By embedding AWFs into the DNN stem, AWFNet is constructed for automated early PD screening from OCT images. Additionally, we introduce a novel Balanced Confidence (BC) loss to boost early PD screening performance and trustworthiness of AWFNet, by mining the potential of sample-wise predicted probabilities across all classes and class frequency prior. The extensive experiments manifest the superiority of AWFNet with BC over state-of-the-art methods in terms of early PD screening performance and trustworthiness.
PURPOSE:To investigate the extent of lens opacity image features measured by AS-OCT and their association with disease severity based on Lens Opacities Classification System III (LOCS III) in eyes with nuclear cataract (NC) and to determine the diagnostic performance of relative features for grading of nuclear lens opacity. SETTING:Multicenter study at 2 sites. DESIGN:Clinical validation. METHODS:A total of 127 individuals with different severity of nuclear cataract were recruited from two different clinical centers: Thailand (Thai, n=81) and Shenzhen, China (SZRM, n=46). All patients underwent AS-OCT examination, images were graded under the LOCS III standard. Automated machine learning models were developed to extract the nuclear region annotation, feature-based quantifiers were then analyzed and evaluated through classifying cataract severity based on disease severity according to LOCS III. RESULTS:AS-OCT pixel based features such as mean, variance, Root Mean Square (RMS), interquartile range, and percentiles significantly correlate with nuclear cataract grading (p < 0.01). Features such as variance, standard deviation, and median showed high consistency, while kurtosis and skewness were negatively correlated. Prediction model achieved 0.81 accuracy on SZRM center (F1-score 0.82), and 0.87 accuracy on Thai center (F1-score 0.83). CONCLUSIONS:Automated AS-OCT image features has strong consistency in lens opacity grading. Potentials are also shown in supportive diagnosis and surgical planning in nuclear cataract.
PURPOSE:To predict multiple postoperative parameters after implantable collamer lens (ICL) surgery with generative artificial intelligence, using preoperative anterior segment optical coherence tomography (AS-OCT) images as the input. SETTING:Daikanyama Eye Clinic, Tokyo, Japan; Miyata Eye Hospital, Miyazaki, Japan; Nagoya Eye Clinic, Nagoya, Japan; Yokohama Sky Eye Clinic, Kanagawa, Japan. DESIGN:Retrospective study. METHODS:The research involved paired preoperative and postoperative AS-OCT images from 1010 patients (1585 eyes) who underwent horizontal ICL implantation and 86 patients (86 eyes) who received vertical implantation from 4 clinical centers. A implantable collamer lens-generative adversarial network (ICL-GAN) was used to predict postoperative structures based on preoperative AS-OCT slice from each eye. Postoperative parameters, including vault, AOD500, and TIA500, were measured from the predicted postoperative structures. The prediction error was evaluated using the mean absolute error (MAE) and root mean square error. The correlation and agreement between the prediction and the achieved values were also analyzed. RESULTS:The vaults measured from the predictions of postoperative anatomical structure have a strong correlation with the achieved values on horizontal data ( r = 0.659, P < .01 for ICL size of 12.1 mm; r = 0.799, P < .01 for 12.6 mm; and r = 0.737, P < .01 for 13.2 mm) and when compared with the NK-formula and KS-formula achieved the minimum prediction errors (MAE are 105 μm, 114 μm, and 111 μm). ICL-GAN also performed best on the vertical implantation data. The AOD500 and TIA500 also showed good correlation and agreement with the achieved values. CONCLUSIONS:The generative artificial intelligence demonstrated the capability to predict multiple postoperative parameters after ICL surgery.
Self-supervised monocular depth estimation serves as a key task in the development of endoscopic navigation systems. However, performance degradation persists due to uneven illumination inherent in endoscopic images, particularly in low-intensity regions. Existing low-light enhancement techniques fail to effectively guide the depth network. Furthermore, solutions from other fields, like autonomous driving, require well-lit images, making them unsuitable and increasing data collection burdens. To this end, we present DeLight-Mono - a novel self-supervised monocular depth estimation framework with illumination decoupling. Specifically, endoscopic images are represented by a designed illumination-reflectance-depth model, and are decomposed with auxiliary networks. Moreover, a self-supervised joint-optimizing framework with novel losses leveraging the decoupled components is proposed to mitigate the effects of uneven illumination on depth estimation. The effectiveness of the proposed methods was rigorously verified through extensive comparisons and an ablation study performed on two public datasets.
Quantitative analysis of retinal vascular morphology is vital for clinical decision-making and the investigation of systemic diseases. Central to this process is the accurate segmentation of retinal arteries and veins (A/V) from the background, a task challenged by substantial variations in vessel calibers and the presence of low-contrast or ambiguous structures in fundus images, especially in ultra-wide field imaging where peripheral distortions and large-scale anatomical variability are pronounced. These factors often lead to fragmented semantic representations and topological inconsistencies in automated segmentation outputs. To address these limitations, we propose Ultra, a multi-granularity topological reasoning network designed for precise A/V segmentation. Ultra adopts a cascaded two-stage architecture: PriorNet generates coarse, multi-scale vascular priors that provide structural guidance, while RefineNet performs topology-aware segmentation refinement. To further enforce topological coherence, we propose the neighboring pixel connectivity regularization (NICER) layer, which selectively integrates local connectivity information predicted by the proposed connectivity prediction union (CPU) module. This connectivity is employed as auxiliary supervision through a pixel-wise local connectivity loss, reinforcing structural reasoning and promoting anatomically consistent vascular topology inference. Extensive experiments on ultra-wide field fundus imaging (UWF) datasets demonstrate that Ultra achieves state-of-the-art performance in A/V segmentation and topological preservation. Moreover, Ultra generalizes well to conventional color fundus photography (CFP) datasets, underscoring its robustness and broad applicability. Code is publicly available at: https://github.com/iMED-Lab/Ultra.
Intra-operative Optical Coherence Tomography (iOCT) is essential for ophthalmic surgery but is heavily degraded by speckle noise, while real-time acquisition prevents multi-frame averaging. Existing self-supervised Blind-Spot Networks (BSNs) suppress noise but often over-smooth direction-sensitive anatomical boundaries. We propose GE-BSN, a Geometric-Equivariant Blind-Spot Network for edge-sensitive iOCT denoising. GE-BSN embeds geometric equivariance into the asymmetric blind-spot mechanism through Geometric Prior-Guided Conv (GPG-Conv), enabling orientation-aware structural regularization. We further introduce Adaptive Anatomy-Preserving Fusion (AAP-Fusion) to balance strong background smoothing with high-fidelity boundary preservation. Experiments on real-world iOCT datasets show competitive denoising and improved structural fidelity for downstream retinal layer segmentation.
Speckle removal in anterior segment optical coherence tomography (AS-OCT) images is a nonlinear inverse problem that improves image quality. Although untrained network priors have proven effective in inverse restoration tasks but are not directly applicable to speckle removal in clinical AS-OCT images due to spectral bias: networks focus more on low-frequency information, whereas high-frequency structural details are vital for clinical analysis. In this letter, we aim to boost the Untrained network Priors for AS-OCT image despeckling by Spectral Bias Compensation (UP-SBC), which enriches the structural information and overall regularization. Specifically, we compensate for the loss of high-frequency structural information and analyze its efficacy via analyzing the frequency band correspondence. Then, we further design a data fidelity to regularize the nonlinear nature of speckle removal and incorporate a focal frequency loss to decouple the structural details and speckle noise for network optimization. Experiments verify the efficacy of the UP-SBC, providing high-quality despeckling results while preserving structural details.
4D medical image interpolation aims to recover missing volumes from sparsely observed time points and is important for dynamic anatomical analysis in applications such as cardiac MRI and thoracic CT, where motion is often repetitive or near-periodic over clinically relevant intervals. A key challenge is that this structure is not always encoded directly in deformation representations for interpolation. In addition, physiological motion is often non-uniform, so equal temporal intervals do not necessarily correspond to equal amounts of anatomical change. To address these issues, we formulate interpolation as learning a continuous deformation process with a phase-structured prior. Given two endpoint volumes, we parameterize a phase-conditioned velocity field with a finite Fourier basis, which embeds near-periodic motion patterns directly into the deformation space and supports continuous querying at arbitrary target times. We further introduce a phase-aligned temporal reparameterization that maps normalized within-interval time to a latent motion phase according to deformation variation intensity, thereby better modeling non-uniform motion progression. Intermediate volumes are then synthesized by continuously warping both endpoints, followed by bidirectional fusion and lightweight residual refinement. Experiments on ACDC and 4D-Lung show that the proposed method achieves state-of-the-art performance over existing baselines while producing anatomically plausible and coherent intermediate volumes from sparse observations.
The robust segmentation of different targets in multiple modality images is challenging due to factors such as low contrast, variations in target size and shape, and interference from diseases, which may lead to segmentation ambiguity. In addition, the assessment of the reliability of artificial intelligence is crucial for its clinical application. This paper proposes the Online Bayesian approximation based Uncertainty-aware Network (OBU-Net) for robust ophthalmic image segmentation. Our approach introduces an efficient online Bayesian method to update a spatial uncertainty map during training continuously. Then, the Spatial Uncertainty Aware Block (SUA-B) leverages the uncertainty map to localize and prioritize attention to ambiguous regions. Additionally, we extract pixel-wise confidence from multi-scale predictions to integrate hierarchical predictions. We compare OBU-Net with state-of-the-art (SOTA) methods on six datasets. The experimental results demonstrate that our method achieves the best overall performance across different modalities and segmentation tasks, highlighting the robustness of our approach. Additionally, metamorphic testing experiments were conducted, exploring the algorithm's stability against random perturbations. Lastly, we propose an image-level uncertainty score and demonstrate its effectiveness for evaluating the model's segmentation reliability.
With the rapid advancements in computer vision technology, real-time autonomous navigation systems for blind and visually impaired individuals (BVIs) leveraging scene understanding have become increasingly feasible. Remarkably, existing navigation systems demonstrate excellent performance in obstacle detection and directional guidance for BVIs. Nevertheless, they lack the spatial perception of the entire scene for BVIs to reach the autonomous navigation. To overcome this issue, we propose a novel computer vision-based method to generate Multiple Paths with Sensory Substitution Device (MP-SSD) system, aiming to effectively and conveniently provide key autonomous navigation information for BVIs in outdoor routes. The MP-SSD system combines potential navigation target detection, path planning, and 3D Semantic Scene Completion (SSC) techniques to develop a widely applicable environmental detection and sensory substitution device (SSD), that can cater to the practical requirements of BVIs. Specifically, MP-SSD can extract semantic and spatial information from the input RGB image through the three-dimensional SSC model, and complete the information of the invisible area of the scene, thereby identifying potential target points and planning the shortest navigation path. During the interaction process, the system provides multi-path prompts through spatial audio to ensure accurate guidance. Experimental analyses on experience feedback of BVIs indicate that MP-SSD can effectively help BVIs actively acquire valuable navigation information in outdoor environments, thereby enhancing their autonomous mobility.
Efficient convolutional neural network (CNN) architecture design has attracted growing research interests. However, they typically apply single receptive field (RF), small asymmetric RFs, or pyramid RFs to learn different feature representations, still encountering two significant challenges in medical image classification tasks: i) They have limitations in capturing diverse lesion characteristics efficiently, e.g., tiny, coordination, small and salient, which have unique roles on the classification results, especially imbalanced medical image classification. ii) The predictions generated by those CNNs are often unfair/biased, bringing a high risk when employing them to real-world medical diagnosis conditions. To tackle these issues, we develop a new concept, Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields (ERoHPRF), to simultaneously boost medical image classification performance and fairness. This concept aims to mimic the multi-expert consultation mode by applying the well-designed heterogeneous pyramid RF bag to capture lesion characteristics with varying significances effectively via convolution operations with multiple heterogeneous kernel sizes. Additionally, ERoHPRF introduces an expertlike structural reparameterization technique to merge its parameters with the two-stage strategy, ensuring competitive computation cost and inference speed through comparisons to a single RF. To manifest the effectiveness and generalization ability of ERoHPRF, we incorporate it into mainstream efficient CNN architectures. The extensive experiments show that our proposed ERoHPRF maintains a better trade-off than state-of-the-art methods in terms of medical image classification, fairness, and computation overhead. The code of this paper is available at https://github.com/XiaoLing12138/Expert-Like-Reparameterization-of-Heterogeneous-Pyramid-Receptive-Fields.
Abstract Accurate semantic segmentation of lunar surface point clouds is a fundamental requirement for autonomous rover navigation during surface traversal. However, lunar terrains present unique challenges, including highly unstructured geometry, weak semantic cues, ambiguous object boundaries, and severe class imbalance, which limit the effectiveness of existing point cloud segmentation methods. To address these issues, we develop an enhanced RandLA-Net framework tailored for lunar surface obstacle perception. A geometry-guided Boundary Enhancement Module (BEM) is introduced to explicitly strengthen semantic boundary features by leveraging local curvature variation, enabling more reliable discrimination of craters and rocks with unclear edges. In addition, a lightweight Global Context Module (GCM) is designed to inject scene-level semantic priors through channel-wise attention, improving consistency in large-scale terrain understanding under random sampling. The network architecture is further optimized to balance computational efficiency and geometric detail preservation for spaceborne deployment. Experiments conducted on the LuSNAR dataset demonstrate that the proposed method achieves a mean Intersection over Union (mIoU) of 95.26%, outperforming the baseline RandLA-Net by 2.04%, with particularly notable improvements in crater segmentation accuracy. The results indicate that the proposed approach provides an effective and efficient solution for lunar surface perception, offering practical value for enhancing the autonomy and safety of future lunar rover missions.
Inadequate bearing lubrication poses a severe threat to the operational safety of rotating machinery. To tackle the diagnostic challenges arising from weak early-stage failure signals, insufficient mechanism interpretability, and the scarcity of field samples, this paper proposes a diagnostic framework integrating Multi-scale Spectral Morphological (MSM) features with adaptive ensemble learning. First, based on the elastohydrodynamic lubrication mechanism and vibration signal modulation characteristics, a multi-scale spectral morphological feature system is constructed. By mapping macro-level energy evolution, micro-level statistical properties, and spectral morphological structures, this system enhances the stability and interpretability of weak signals across evolving failure stages. Furthermore, to address the instability of decision boundaries in single classifiers under class imbalance and complex noise environments, an ADASYN-Stacking ensemble diagnostic model is developed. This model reshapes minority class decision boundaries via Adaptive Synthetic Sampling (ADASYN) and fuses heterogeneous classifiers through a Stacking strategy. Experimental results demonstrate that the proposed method exhibits superior performance on testbed data during the fault latency period. Notably, it maintains a diagnostic accuracy of 97.53% on real industrial field data containing multi-source interference, validating its robustness in early lubrication failure identification and its generalization capability under complex operating conditions.
Retinal fundus image enhancement is a crucial prerequisite for reliable ophthalmic diagnosis and downstream clinical analyses. However, state-of-the-art automated segmentation models suffer severe performance degradation when applied to clinical images due to the domain gap caused by heterogeneous, frequency-dependent artifacts. Existing enhancement networks often struggle with cross-dataset generalization, tending to over-smooth anatomical details or introduce hallucinatory artifacts. To address this, FreqMamba, a novel Frequency-Spatial Hybrid State Space Model, is proposed for generalizable and structure-preserving retinal image enhancement. The core contribution is the Frequency-Spatial Mamba Block (FSMB), which elegantly decouples degradation restoration into dual domains. The spatial branch utilizes bidirectional Vision Mamba to capture global vascular continuity with linear complexity. Concurrently, the frequency branch introduces a learnable channel-wise modulation mechanism guided by a physical Butterworth-like high-frequency prior. Optimized via task-aware spectral-spatial constraints, this mechanism implicitly maintains a mathematical residual formulation to inject high-frequency details without spectrum explosion. Comprehensive experiments demonstrate that FreqMamba exhibits exceptional zero-shot generalization across diverse clinical datasets (e.g., achieving state-of-the-art PSNR across all datasets and an SSIM of 0.854 on CHASE). Compared to recent state-of-the-art methods, it significantly boosts downstream segmentation robustness (e.g., raising the Dice score from 0.508 to 0.531 on DRIVE under severe degradations), establishing a reliable prerequisite for clinical deployment.
Accurate depth estimation is crucial for 3D reconstruction and precise navigation in ophthalmic fundus surgery. However, acquiring annotated data remains challenging due to the impracticality of depth sensors under surgical microscopes.To overcome this limitation, we introduce RetinalDepth-64K, a novel synthetic dataset comprising 64,000 stereo image pairs across 1,280 diverse scenes, developed through a Real2Sim2Real pipeline that transforms real-world fundus surgery videos into synthetic data and facilitates model deployment in real scenarios. We analyzed key characteristics such as intricate retinal textures from real-world videos to guide the Real-to-Sim phase, enabling realistic data synthesis.To improving dataset fidelity for depth estimation, we created 3D eye models using Blender with ultra-wide-field retinal textures, glass-modeled aqueous humor, and dynamic instrument trajectories, enhanced by post-processing to ensure photorealism.The dataset provides RGB images, depth maps, normal maps, and instrument segmentation masks from binocular view, supporting the training of monocular, binocular, and video-based depth estimation models to enhance robustness. In the Sim-to-Real phase, quantitative and qualitative experiments show that finetuning foundation models with RetinalDepth-64K produces accurate depth predictions for synthetic data. Comparative analysis on results of zeroshot and finetuned models further validates robust generalization to real fundus surgery scenes, offering significant potential to enhance surgical precision and support the training of novice surgeons through reliable depth cues.As the first dataset of its kind for retinal surgery, RetinalDepth-64K offers a vital resource for advancing 3D reconstruction and surgical navigation in ophthalmology.
Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential for diagnosis. However, ophthalmic VQA benchmarks primarily emphasize answer accuracy, neglecting the explicit visual evidence necessary for clinical interpretability. In this work, we introduce FundusGround, a new benchmark for clinically interpretable ophthalmic VQA with spatially-grounded lesion evidence. Specifically, we propose a three-stage pipeline that collects 10,719 fundus images with 15,595 image-level meticulously annotated lesions. To ensure anatomical consistency and clinical validity, all lesions are spatially localized using the Early Treatment Diabetic Retinopathy Study (ETDRS) grid, enabling standardized mapping to nine clinically meaningful retinal regions. Built upon this structured lesion evidence, 72,706 questions are then generated spanning four formats: open-ended, closed-ended, single-choice, and multiple-choice. We further benchmark multiple general- and medical- large vision-language models using dual metrics for answer accuracy and lesion-level reasoning. The experiments demonstrate that incorporating lesion-level visual evidence consistently improves model performance and transparency, highlighting the necessity of explicit spatial grounding for reliable and explainable ophthalmic VQA.
Cerebrovascular segmentation provides valuable cues for cerebrovascular diseases. Deep learning has achieved remarkable success in cerebrovascular segmentation, but relies on colossal computing power. To address existing challenges, we studied the intensity characteristics in cerebrovascular imaging and proposed an explicable intensity-aware cerebrovascular segmentation (EI-Seg) with 3D and tri-planar representations to promote accurate and efficient feature learning. In particular, EI-Seg has sufficient semantic interpretability, guiding the model to generate low-dimensional feature maps. Through the strategies of disentanglement and cycle consistency, EI-Seg can accurately describe the semantic features of cerebrovasculature in the latent space using tri-planes, thereby avoiding many redundant parameters and subspaces. More importantly, the inference phase of the model is only completed under the path of tri-planar representation, guiding the nearly 2D structure to achieve 3D semantic representation, thereby saving a lot of computing power. Experimental results confirm that EI-Seg has practically no performance loss, but its cost efficiency far surpasses other competitors. Our code is available at https://github.com/USTB-MEDAI/EI-Seg.
Endpoint-only unsupervised 4D medical image interpolation synthesizes intermediate volumes from sparsely sampled sequences with only the start and end volumes available for training; however, this weakly constrained setting often yields intermediates with unstable boundaries and non-physiological motion, limiting interpretability and downstream analysis. We propose low-rank velocity fields as a structural prior, constraining motion to a structured Tucker low-rank velocity field space that decomposes motion into globally shared spatial bases and a compact sample-specific core, thereby encouraging spatially correlated, anatomy-consistent deformation while suppressing voxel-wise high-frequency artifacts. To capture global coordination and local non-rigid details, we model motion in a coarse-to-fine multi-scale scheme and compose scale-wise deformations at inference to synthesize volumes at arbitrary times. We further provide a theoretical analysis showing that, under Tucker parameterization, low-rank parameters control the smoothness energy of the velocity field, explaining why low-rank modeling promotes smoother motion. Experiments on ACDC and 4D-Lung demonstrate state-of-the-art performance, remaining competitive with methods trained with intermediate-frame supervision, and producing intermediates with improved structural coherence and more stable anatomical contours.
Conventional visual-field-loss simulations often use generic, display-fixed masks that can be bypassed by eye movements. We developed a perimetry-driven, gaze-contingent augmented reality display that converts paired monocular fields into a standardized patient-derived binocular mask. Sparse Humphrey 30-2 sensitivities were independently reconstructed using thin-plate-spline radial basis functions, aligned in a common visual-field coordinate system, integrated pointwise using the best-location rule, and converted to a display texture using a 15-dB threshold. The resulting mask was rendered around the current gaze position in a modified HoloLens 2 video-see-through configuration. Nineteen normally sighted participants completed physical object-search and navigation tasks under unmasked and simulated-loss display conditions; 21 participants with retinal disease, clinically evident field loss, and a visual field index below 30% provided a behavioral reference. Relative to the unmasked video-see-through condition, mean completion time increased by 23.29 s in object search (18.34 to 41.63 s) and by 13.84 s in navigation (36.05 to 49.89 s). The simulated display also increased stationary behavior and cumulative double-support duration in both tasks, increased head rotation during object search, and altered visual sampling. Object search showed broader changes in scanning and gait-related outcomes than navigation. The study demonstrates a traceable clinical-data-to-display pipeline and a task-sensitive framework for evaluating visual, head, and gait behavior during physical interaction.
Precise diabetic retinopathy (DR) grading is essential for developing personalized and effective treatment plans. Although deep neural networks (DNNs) have achieved promising DR grading results, constructing a precise and trustworthy DR grading model remains challenging due to limited high-quality medical image data, high computational costs, and imbalanced data distributions. To tackle these challenges, we explore the transferability of feature representations from pre-trained vision foundation models (VFMs) to fundus images through adapter learning, aiming to build an efficient imbalanced DR grading model. Unlike classical full-tuning, which fine-tunes all pre-trained parameters of VFMs, adapter learning achieves competitive performance by adding negligible finetuned parameter number. Motivated by the above analysis, we develop a Token Pyramid Pooling-Driven Style Adapter Learning (TPDSAL) to better capture task-specific feature representations from VFMs, which fully exploits pathological distribution prior of DR and the inherent fundus imaging characteristics. Besides, we propose a novel dual-view balanced loss (DVB) to improve imbalanced DR grading performance and trustworthiness, which explores the potential of training class frequencies in sample-wise predicted logit space and sample-wise loss value space simultaneously. Extensive experiments on four public fundus image datasets manifest the superiority of our TPDSAL with DVB over competitive transfer tuning and loss methods in terms of imbalanced grading performance and trustworthiness. Further analysis suggests that clinical prior knowledge utilization is beneficial for adapter learning in capturing task-specific feature representations from VFMs.