BACKGROUND:The growth of axial length (AL) can lead to high myopia and ocular deformation, especially causing microstructural changes in the fundus, which cannot be fully quantified by AL alone. We propose an optical coherence tomography (OCT)-based modified AL (Myopic Index) to represent the extent of fundus deformation caused by AL elongation and to explore its clinical significance in myopic progression prediction. METHODS:A deep learning model was trained using 27 539 cases of OCT images and referred ocular biometric data to evaluate the Myopic Index. By comparing the Myopia Index with the Measured AL, the difference of two AL indices (DAL) was calculated. We further prospectively employed 2866 cases of OCT images, which were categorised into short AL (Measured AL<22 mm), normal AL (22 mm≤Measured AL<26 mm) and long AL (≥26 mm), to evaluate the model ability of myopic progression prediction. The attention regions of images were also analysed. RESULTS:The Myopia Index was closely correlated with Measured AL (all p<0.001, R²=0.804 in all eyes). Specifically, the Myopia Index was closer to the Measured AL in eyes with long ALs, whereas in eyes with short and normal axial lengths, the Myopia Index clustered around 23-24 mm. The visualisation model demonstrated that for eyes with short and normal ALs, attention regions were primarily concentrated on the retina; conversely, for eyes with long ALs, the choroidal layer and the retinal pigment epithelium layer received more attention. Moreover, DAL was significantly correlated with AL increment (p=0.038). CONCLUSIONS:The Myopia Index reflects the real status of fundus microstructures through fundus microstructures, with a particular focus on the choroid. The Myopia Index demonstrates good predictive capabilities for high myopia progression.
BACKGROUND/AIMS:This study evaluated the performance of artificial intelligence (AI) algorithms in predicting best-corrected visual acuity (BCVA) for patients with multiple retinal diseases, using multimodal medical imaging including macular optical coherence tomography (OCT), optic disc OCT and fundus images. The goal was to enhance clinical BCVA evaluation efficiency and precision. METHODS:A retrospective study used data from 2545 patients (4028 eyes) for training, 896 (1006 eyes) for testing and 196 (200 eyes) for internal validation, with an external prospective dataset of 741 patients (1381 eyes). Single-modality analyses employed different backbone networks and feature fusion methods, while multimodal fusion combined modalities using average aggregation, concatenation/reduction and maximum feature selection. Predictive accuracy was measured by mean absolute error (MAE), root mean squared error (RMSE) and R² score. RESULTS:Macular OCT achieved better single-modality prediction than optic disc OCT, with MAE of 3.851 vs 4.977 and RMSE of 7.844 vs 10.026. Fundus images showed an MAE of 3.795 and RMSE of 7.954. Multimodal fusion significantly improved accuracy, with the best results using average aggregation, achieving an MAE of 2.865, RMSE of 6.229 and R² of 0.935. External validation yielded an MAE of 8.38 and RMSE of 10.62. CONCLUSION:Multimodal fusion provided the most accurate BCVA predictions, demonstrating AI's potential to improve clinical evaluation. However, challenges remain regarding disease diversity and applicability in resource-limited settings.
While OCT is pivotal for macular disease diagnosis, its adoption in primary care is limited by AI systems that cannot simultaneously analyze multi-sectional scans across the full spectrum of maculopathies or generate diagnostically integrated reports. Here we present iOCT, an intelligent OCT analysis system that integrates a multi-level annotation framework with data distillation to enable automated multi-sectional scan analysis and comprehensive natural language report generation. iOCT was trained on 107,790 macular OCT scans (1,296,439 images) and annotated across four levels combining case-level natural language descriptions with image/study-level diagnostic classifications. Internally, iOCT achieved BLEU-1 of 0.5995 for report generation—outperforming all baselines including R2Gen—and a mean AUC of 0.988 (95% CI: 0.985–0.991) for diagnostic classification. In prospective multicenter validation across ten hospitals (8998 cases), Integrated Reports achieved physician-level quality in 96.4% of cases (mean 2.94/3), significantly outperforming NL Reports (79.5%, 2.60/3), with the greatest gains at previously low-performing centers (1.38–2.93). iOCT matched junior ophthalmologists in speed (23.37 s vs. 24.44 s, p = 0.368) and report quality (2.91 vs. 2.88, p = 0.279), outperformed residents (p < 0.001), and approached senior specialists (19.47 s, 2.98). These findings establish iOCT as a deployable, multi-disease OCT system with performance approaching trained ophthalmologists, supporting large-scale retinal disease screening in primary care.
Early detection of colorectal cancer hinges on real-time, accurate polyp identification and resection. Yet current high-precision segmentation models rely on GPUs, making them impractical to deploy in primary hospitals, mobile endoscopy units, or capsule robots. To bridge this gap, we present the UltraSeg family, operating in an extreme-compression regime (<0.3 M parameters). UltraSeg-108K (0.108 M parameters) is optimized for single-center data, while UltraSeg-130K (0.13 M parameters) generalizes to multi-center, multi-modal images. By jointly optimizing encoder-decoder widths, incorporating constrained dilated convolutions to enlarge receptive fields, and integrating a cross-layer lightweight fusion module, the models achieve 90 FPS on a single CPU core without sacrificing accuracy. Evaluated on seven public datasets, UltraSeg retains >94
Color fundus photography (CFP) is the mainstay for large-scale retinal screening, yet its diagnostic capacity is constrained by the lack of depth-resolved structural information. Optical coherence tomography (OCT) provides cross-sectional retinal anatomy, but is less accessible in population-level screening. Here, we present EyeMVP, a cross-modal retinal foundation model that uses paired CFP–OCT pretraining to learn OCT-informed CFP representations. EyeMVP is pretrained on 674,893 strict same-eye same-day paired CFP–OCT image triples from 112,642 patients across eight hospitals in China. The model uses cross-modal masked reconstruction to enrich CFP representations with OCT-associated supervision, while requiring only CFP images at inference. To accommodate the non-aligned imaging geometry between en-face CFP and cross-sectional OCT, EyeMVP combines source-constrained cross-attention with CFP-derived structural masks. Across 16 downstream tasks, including classification, segmentation, few-shot adaptation, and cross-modal retrieval, EyeMVP outperforms representative retinal foundation models and shows consistent gains on tasks involving macular and optic nerve structure. For CFP-challenging macular diseases, EyeMVP achieves an AUROC of 0.948 for macular edema (vs. 0.852 for EyeCLIP) and 0.825 for myopic macular schisis. In an exploratory reader study, EyeMVP exceeds junior and intermediate ophthalmologist groups but does not reach senior ophthalmologist performance on macular edema, while showing numerically higher balanced accuracy than all reader groups on myopic macular schisis. These results suggest that pixel-level cross-modal reconstruction can enrich CFP representations with OCT-associated supervision, providing a practical route toward stronger CFP-based retinal analysis in screening settings.
Insulin resistance (IR) is a key precursor to diabetes and a significant risk factor for cardiovascular disease. Traditional IR assessment methods require multiple blood tests. We developed a simple AI model using only fasting blood glucose to predict IR in non-diabetic populations. Data from the NHANES (1999-2020) and CHARLS (2015) studies were used for model training and validation. Input features included age, gender, height, weight, blood pressure, waist circumference, and fasting blood glucose. The CatBoost algorithm achieved AUC values of 0.8596 (HOMA-IR) and 0.7777 (TyG index) in NHANES, with an external AUC of 0.7442 for TyG. For METS-IR prediction, the model achieved AUC values of 0.9731 (internal) and 0.9591 (external), with RMSE values of 3.2643 (internal) and 3.057 (external). SHAP analysis highlighted waist circumference as a key predictor of IR. This AI model offers a minimally invasive and effective tool for IR prediction, supporting early diabetes and cardiovascular disease prevention.
Optical coherence tomography (OCT) is an advanced retinal imaging technique that enables non-invasive cross-sectional visualization of the retina, playing a crucial role in ophthalmology for detecting various macular lesions. While deep learning has shown promise in OCT image analysis, existing studies have primarily focused on broad, image-level disease diagnosis. This study introduces the Assistive Diagnosis Framework for OCT (ADF-OCT), which utilizes a dataset of over one million macular OCT images to construct a multi-label diagnostic model for common macular lesions and a medical report generation module. Our innovative Multi-frame Medical Images Distillation method effectively translates study-level multi-label annotations into image-level annotations, thereby enhancing diagnostic performance without additional annotation information. This approach significantly improves diagnostic accuracy for multi-label classification, achieving an impressive AUROC of 0.9891 with best performance macro F1 of 0.8533 and accuracy of 0.9411. By refining the feature fusion strategy in multi-frame medical imaging, our framework substantially enhances the generation of medical reports for OCT B-scans, surpassing current solutions. This research presents an advanced development pipeline that utilizes existing clinical datasets to provide more accurate and comprehensive artificial intelligence-assisted diagnoses for macular OCT.
Neural networks have achieved remarkable success across various fields. However, the lack of interpretability limits their practical use, particularly in critical decision-making scenarios. Posthoc interpretability, which provides explanations for pretrained models, is often at risk of fidelity and robustness. This has inspired a rising interest in self-interpretable neural networks (SINNs), which inherently reveal the prediction rationale through model structures. Despite this progress, existing research remains fragmented, relying on intuitive designs tailored to specific tasks. To bridge these efforts and foster a unified framework, we first collect and review existing works on SINNs and provide a structured summary of their methodologies from five key perspectives: attribution-based, function-based, concept-based, prototype-based, and rule-based self-interpretation. We also present concrete, visualized examples of model explanations and discuss their applicability across diverse scenarios, including image, text, graph data, and deep reinforcement learning (DRL). Additionally, we summarize existing evaluation metrics for self-interpretation and identify open challenges in this field, offering insights for future research. To support ongoing developments, we present a publicly accessible resource to track advancements in this domain: https://github.com/yangji721/Awesome-Self-Interpretable-Neural-Network
Retinal artery-vein vessels are associated with systemic chronic diseases and cardiovascular diseases. Therefore, the accurate quantitative analysis of retinal artery-vein vessels is the preliminary basis of clinical diagnosis. Most of the existing artificial intelligence(AI) methods are data-driven. Although some public retinal artery-vein vessel segmentation datasets have been released, their data quality is unsatisfactory. In this paper, we establish a new fundus image dataset for AI-based artery-vein segmentation, Fundus-AVSeg. It consists of 100 high-resolution fundus images with pixel-wise manual annotation by professional ophthalmologists. We believe our Fundus-AVSeg will benefit the further development of retinal artery-vein vessel segmentation.
Retinal images play a crucial role in diagnosing various eye diseases. However, due to the shortage of ophthalmologists, many patients do not receive timely diagnosis and treatment. Though automatic eye disease diagnosis has been improved due to recent advancements in deep learning, current deep learning models often exhibit unsatisfactory performance when applied to multi-labeled ocular disease classification tasks with fundus images. This paper addresses the limitations of existing models by proposing a new retinopathy dataset, the Resk dataset, and the Multi-Resolution Attention Feature Fusion (MRAFF) architecture. The MRAFF modifies current backbones with multiple downsample steps and incorporates three types of attention mechanisms: channel attention, spatial attention, and linear attention. Extensive experiments demonstrate that MRAFF-modified models exhibit improved feature extraction and classification abilities in handling ocular disease screening. The MRAFF is expected to enhance ocular disease diagnosis and may also suit other multi-label disease screenings, thereby inspiring future computer-aided diagnostic methods.
The retinal fundus images are extensively utilized in diagnosis, and their quality may affect diagnostic results. However, due to limitations in the datasets and algorithms, current fundus image quality assessment (FIQA) methods often lack the granularity required to meet clinical demands. To address these limitations, we introduce a new benchmark FIQA dataset, Fundus Quality Score, which contains 2,246 images annotated with continuous mean opinion scores ranging from 0 to 100 and three-level quality categories. Meanwhile, we also design a novel FIQA Transformer-based Hypernetwork (FTHNet). The FTHNet can treat FIQA as a regression task to predict the continuous MOS, diverging from common classification-based approaches. Results on our dataset show that FTHNet predicts quality scores, achieving a Pearson Linear Correlation Coefficient of 0.9423 and a Spearman Rank Correlation Coefficient of 0.9488, significantly outperforming compared methods while utilizing fewer parameters and lower computational complexity. Furthermore, model deployment experiments demonstrate its potential for use in automated medical image quality control workflows. We have released the code and dataset to facilitate future research in this field.
Tabular data is one of the most common forms of data in real-life applications. In tabular data modelling, methods based on Gradient Boosting Decision Trees (GBDT) and neural networks have demonstrated unique advantages on different datasets. However, existing neural network methods still face bottlenecks in processing continuous numerical features. This study revisits the structural design of numerical embedding in tabular modelling based on neural networks to overcome these bottlenecks. We conceptually decouple the numerical embedding process into three core functional modules: numerical augmentation, normalization, and encoding. Based on this design, we have developed three novel models-TabMLPNet, TabKANet, and TabKANet-PLNL. These models apply batch normalization to numerical features and feed them into a single, shared Multi-Layer Perceptrons (MLP) or Kolmogorov-Arnold networks (KAN) encoder. We propose a Piecewise Learnable Noise Layer (PLNL), which enriches data representation by introducing piecewise noise into numerical features, further enhancing the model's generalization ability. Extensive experiments on 18 public datasets demonstrate that our models achieve significant performance improvements, outperforming other neural network models. This study underscores the significant benefits of integrating Kolmogorov-Arnold Network (KAN) with batch normalization. This combination not only enables efficient normalization of numerical features, but also allows dynamic learning of numerical distributions within each batch. As a result, it effectively mitigates performance bottlenecks that are often caused by numerical skewness. Additionally, we have confirmed that introducing augmented noise during the training process can enhance the model's feature learning capabilities. Our implementation is publicly available on GitHub and can be accessed at https://github.com/AI-thpremed/TabKANet.
The mRNA optimization is essential for mRNA vaccines, therapies, and industrial protein production. Based on current explorations, an ideal optimization approach should simultaneously (i) prevent unintended amino-acid changes, (ii) optimize multiple, biologically relevant objectives, and (iii) retain computational efficiency. However, existing methods are forced to trade off between these perspectives, forming an "impossible triangle." We present RNop, a knowledge-infused Transformer that integrates mechanism-aligned losses to address this problem. By encoding biological prior knowledge in losses, RNop makes knowledge infusion explicit and controllable across optimization focus. Trained on over 6 million sequences, in silico analyses show RNop resolves the "impossible triangle" of mRNA optimization with absolute sequence fidelity, significantly improved biological metrics, and high throughput. In in vitro validation, it can deliver up to 2.28-fold expression gain. Ablation studies reveal how each prior contributes to targeted improvements, yielding mechanism-level interpretability. RNop represents a shift in mRNA optimization methodology: by infusing explicit and interpretable knowledge, the "black-box" mRNA design can be transformed into a predictable, explainable engineering problem. RNop is designed as an extensible platform: additional biological priors can be incorporated as modular, mechanism-aligned loss functions, enabling future development and adaptation to related sequence design problems.
Cataract is one of the most common blinding eye diseases and can be treated by surgery. However, because cataract patients may also suffer from other blinding eye diseases, ophthalmologists must diagnose them before surgery. The cloudy lens of cataract patients forms a hazy degeneration in the fundus images, making it challenging to observe the patient's fundus vessels, which brings difficulties to the diagnosis process. To address this issue, this paper establishes a new cataract image restoration method named Catintell. It contains a cataract image synthesizing model, Catintell-Syn, and a restoration model, Catintell-Res. Catintell-Syn uses GAN architecture with fully unsupervised data to generate paired cataract-like images with realistic style and texture rather than the conventional Gaussian degradation algorithm. Meanwhile, Catintell-Res is an image restoration network that can improve the quality of real cataract fundus images using the knowledge learned from synthetic cataract images. Extensive experiments show that Catintell-Res outperforms other cataract image restoration methods in PSNR with 39.03 and SSIM with 0.9476. Furthermore, the universal restoration ability that Catintell-Res gained from unpaired cataract images can process cataract images from various datasets. We hope the models can help ophthalmologists identify other blinding eye diseases of cataract patients and inspire more medical image restoration methods in the future.
BACKGROUND:Insulin resistance is a key precursor to diabetes and increases the risk of cardiovascular diseases. Traditional assessment methods rely on multiple invasive tests. Developing an AI model based on minimally invasive tests, especially using only fasting blood glucose as the invasive test, can promote health monitoring in non-diabetic populations, particularly for frequent routine checks. OBJECTIVE:This study aims to develop an AI-driven model that uses only fasting blood glucose as the invasive measure to predict insulin resistance in non-diabetic populations. The goal is to facilitate health monitoring through simple, minimally invasive tests. METHODS:We selected simple and accessible input features, including age, gender, height, weight, pulse, blood pressure, waist circumference, and fasting blood glucose. Data from the National Health and Nutrition Examination Survey (NHANES, 1999-2020) were used to construct four AI-based prediction models, which were validated using data from the China Health and Retirement Longitudinal Study (CHARLS, 2015). These models were based on three commonly used insulin resistance (IR) indicators: HOMA-IR, TyG, and METS-IR. Additionally, we used SHAP values to interpret the contributions of these features to the predictions. RESULTS:The CatBoost algorithm performed excellently in classification tasks for insulin resistance. For numerical prediction of the METS-IR index, neural networks, particularly TabKANet, demonstrated superior performance in cross-dataset validation. In the NHANES test set, the AUC values for predicting insulin resistance were 0.8596 (HOMA-IR index) and 0.7777 (TyG index), with an external validation AUC of 0.7442 for the TyG index. For METS-IR prediction, our model achieved AUC values of 0.9731 (internal) and 0.9591 (external). Additionally, the AI-driven model for predicting METS-IR had RMSE values of 3.2643 (internal) and 3.057 (external). SHAP analysis identified waist circumference as a key predictor of insulin resistance, highlighting its importance in early diabetes and cardiovascular disease prediction. CONCLUSION:This study successfully developed a minimally invasive insulin resistance prediction model that relies solely on fasting blood glucose. The AI-driven models demonstrated robust performance across multiple insulin resistance assessment indicators, particularly in predicting the METS-IR index. These findings highlight the significant potential of AI in enhancing early detection and monitoring of insulin resistance in non-diabetic populations, thereby improving health monitoring strategies.
Uveal Melanoma(UM) is a highly aggressive ocular malignancy. Once metastasis occurs, the survival period is very short. An effective prognosis for UM is necessary. The BAP1 gene and Somatic Copy Number Alterations (SCNA) are two of the important indicators for clinical UM prognosis. However, current methods for detecting prognosis-related genes have limitations, including high costs and the need for advanced technical expertise, which makes them challenging to apply in routine clinical practice. Low-cost and efficient prognosis-related gene detection is worth studying. In this paper, we investigate the prognosis-related gene subtyping prediction problem based on Whole Slide Images(WSI). We propose a novel method, UMGen, for WSI-based UM prognosis-related gene subtyping prediction. Specifically, UMGen consists of a multi-region sampling module, a classifier module, and a joint decision module. The classifier module, SAGNet, applies spatial attention to extract multi-scale features. To improve accuracy and generalization, the multi-region joint decision is introduced. Comprehensive quantitative ablation experiments demonstrate that both our UMGen and SAGNet can surpass other competitors and have excellent generalization and robustness. The code and models will be released to the public for further research.
The basal diameter of uveal melanoma is critical for its prognosis and therapy, and it can indicate the metastatic risk of the tumor. Tumor segmentation is also significant for guiding clinical diagnosis, as it provides morphological and other essential information for clinicians. However, precise uveal melanoma segmentation and basal diameter prediction remain challenging for current computer-aided methods. The scarcity of large, annotated datasets and appropriate multi-task models hinders further exploration to simultaneously and accurately segment the tumors and predict their basal diameter for uveal melanoma. To address this challenge, we collected a novel dataset and proposed the Tumor Segmentation and Basal Diameter Prediction Network for Uveal Melanoma (TSBPNet-UM), which utilizes readily accessible fundus images to accomplish both tasks concurrently. Its two sub-networks, the Uveal Melanoma Segmentation and Basal Diameter Prediction networks, are designed for mutual performance enhancement. The Uveal Melanoma Segmentation network transfers spatial segmentation information to assist the Basal Diameter Prediction network. Conversely, the Basal Diameter Prediction network updates all parameters using basal diameter-derived gradient information. This model effectively overcomes the limitation that conventional models often struggle to converge on the direct prediction task of basal diameter. The superior performance and efficacy of the TSBPNetUM have been demonstrated through extensive experiments.
Text-to-SQL parsing has attracted substantial attention recently due to its potential to remove barriers for non-expert end users interacting with databases. A key challenge in Text-to-SQL parsing is developing effective encoding mechanisms to capture the complex relationships between question words, database schemas, and their associated connections within the heterogeneous graph structure. Existing approaches typically introduce some useful multi-hop structures manually and then incorporate them into graph neural networks (GNNs) by stacking multiple layers, which (1) ignore the difficult-to-identify but meaningful semantics embedded in the multi-hop reasoning path, and (2) are limited by the expressive capability of GNN to capture long-range dependencies among the heterogeneous graph. To address these shortcomings, we introduce GRL-SQL, a graph reasoning enhanced language model, which innovatively applies structure encoding to capture the dependencies between node pairs, encompassing one-hop, multi-hop and distance information, subsequently enriched through self-attention for enhanced representational power over GNNs. Furthermore, GRL-SQL incorporates an interaction module that enables joint reasoning and fusion over the question-schema representations for enhancing global context modeling. Comprehensive experiments demonstrate the effectiveness and robustness of our proposed GRL-SQL.