
We present EpitopeNet, a population-based selective learning framework inspired by B-cell dynamics, capable of classifying medical images without backpropagation.The model relies on a set of B=2133 autonomous prototypes that progressively specialize on discriminative local patches through a supervised Learning Vector Quantization (LVQ)-based learning rule.Each prototype only participates in learning if its similarity with the closest patch of an image exceeds a fixed global threshold θ, and under these conditions, it updates its internal representations or internal vector by moving closer to or further from the captured patch depending on the class label of the current image.The class of a prototype is not imposed at initialization or a priori; it is determined adaptively through the accumulation of activation counters normalized by class frequency.At inference time, the decision is made by a weighted vote of the active prototypes, where each prototype contributes proportionally to its class exclusivity.This architecture offers native interpretability that faithfully reflects the decision-making mechanism : for each analyzed image, the local regions that motivated the decision can be directly identified from the activated prototypes, without requiring post-hoc explanation techniques. Evaluated on MiniDDSM (Cancer vs. Normal classification, 3848 images), the model achieves an accuracy of 82.97 ± 1.27%, a macro F1 of 0.8209 ± 0.0115, and an AUC of 0.9026 ± 0.0056 across 5 independent runs, with balanced learning between the two classes. This work opens a perspective on non-differentiable and biologically inspired learning for medical artificial intelligence.
Background Stroke is the third most common cause of death and the fourth most common cause of disability in the world, with 11.9 million incident events each year. Although deep learning methods, particularly convolutional neural networks and transfer-learning-based models, have been increasingly applied to stroke classification, existing approaches face several challenges, including limited cross-modality comparison, class imbalance, low interpretability, and suboptimal accuracy when applied to heterogeneous imaging protocols. Moreover, no existing framework has systematically combined these techniques into modality-specific stroke classification models evaluated separately on both CT and MRI data. Objective To create and test a deep learning architecture that can be used to accurately and interpretably classify stroke from computed tomography and magnetic resonance imaging scans across imaging modalities. Methods The proposed framework comprises a five-component ensemble architecture based on transfer learning from the pre-trained convolutional networks on ImageNet, such as Xception, EfficientNetB7 and ConvNeXtSmall, ensemble voting between different architectures, explainable artificial intelligence using Grad-CAM visualization technique, systematic hyperparameter optimization and adaptive data augmentation. The framework was trained and tested on 1005 computed tomography (CT) images that were taken at Shafa Hospital and 615 magnetic resonance imaging (MRI) images from a public dataset, including three diagnostic classes: hemorrhagic stroke, ischemic stroke and normal. Results On the test set of CT, the ensemble model gave an accuracy of 95.12% with a precision of 97%, a recall of 95%, and an F1 score of 96%. The ensemble attained the highest accuracy of 94.12%, precision of 94%, recall of 95% and F1-score of 95% on MRI. CT ensemble performance was better than any of the single architectures and previous best-known methods. Clinically relevant attention patterns were confirmed with Grad-CAM visualizations, which aligned with the radiologically relevant lesion regions. Conclusion The current study is a system to systematically combine transfer learning, ensemble voting, explainable artificial intelligence, and adaptive augmentation for stroke classification. The framework is clinically interpretable and demonstrated to provide superior performance metrics, thus having the potential to be real-world deployed as a decision-support tool in emergency neuro-imaging workflows. The main contribution of this work is the systematic integration of five complementary components into a modality-specific framework, enabling accurate and clinically interpretable classification for both CT and MRI, with each modality modeled and evaluated independently.
Medical question-answering systems are commonly adapted to maximize aggregate accuracy even though the consequences of incorrect answers differ across clinical domains. MediLoRA is a physician-informed parameter-efficient fine-tuning framework in which domain criticality conditions a coupled adaptation policy comprising effective rank, target-module coverage, loss weighting, curriculum order, and tier-dependent dropout. Eight physicians rated ten specialties for safety impact, reverse-coded error reversibility, and knowledge complexity. The framework was evaluated with LLaMA-3-8B on 24,000 multiple-choice items, with a 2400-item held-out split containing 2000 MedMCQA questions and 400 clinician-validated synthetic stress-test cases. In the cross-model held-out comparison, MediLoRA achieved 84.5% exact-match accuracy (Wilson 95% CI, 83.0–85.9), compared with 77.5% for uniform LoRA, 64.2% for the frozen backbone, and 85.3% for a full fine-tuning reference. Five independent MediLoRA runs yielded 83.8% mean accuracy (SD, 1.2). For the three critical domains with available domain-level results, gains over uniform LoRA were +18.1 percentage points in pharmacology, +14.3 in emergency medicine, and +9.3 in cardiology. No predefined high-risk benchmark errors occurred for MediLoRA in the 400-case audit; the one-sided exact upper 95% bound was 0.75%. The maximal MediLoRA adapter configuration contained approximately 65.0 million trainable parameters (0.80% of the backbone), versus 6.8 million (0.08%) for uniform LoRA. Because the comparison was not parameter matched and multiple risk-conditioned components changed jointly, the results support the feasibility of the coupled policy rather than an isolated causal effect of rank allocation.
Histopathological images analysis is an essential task in diagnosing cancer, but the application of artificial intelligence (AI) in computational pathology is often constrained by the heavy computational demands of convolutional neural networks (CNNs), Vision Transformers (ViTs), and pathology foundation models due to their computational cost, memory requirements, inference time, and infrastructural needs. This review explores the role of lightweight deep learning models in histopathological image analysis in terms of their diagnostic accuracy, computational efficiency, explainability, and deployability in resource-constrained healthcare environments. Recent advancements in lightweight CNNs, compact ViTs, hybrid CNN-Transformer architectures, efficient attention mechanisms, and optimization algorithms were discussed. The paper also addresses the issue of benchmark histopathology datasets, evaluation procedures, and the necessity of reporting not only clinical performance but also computational metrics like parameters, floating-point operations per second (FLOPs), inference time, memory usage, and the platform for deploying the model. Issues such as insufficient amounts of annotated data, staining, and scanner variability, domain shift between institutions, poor external validation, inadequate explainability, and lack of benchmarking are among the key challenges mentioned. In general, the review indicates that there is an effective way towards scalable, interpretable, and clinically useful computational pathology, particularly in resource-constrained healthcare environments where accuracy must be balanced with efficiency and real-world deployability.
Esophagectomy generates thousands of pre-, intra-, and early postoperative data points per patient, yet these remain siloed, and preoperative risk scores plateau at an area under the curve of about 0.62 for anastomotic leak. Enhanced Recovery After Surgery checklists improve outcomes but are static and open-loop, unable to adapt to a patient's evolving risk. We make the case for a shift toward closed-loop monitoring with AI-supported perioperative management: a sequential, milestone-based system that recalibrates risk at defined timepoints (preoperative assessment, incision/OR start, anastomosis, postoperative day 1 endoscopy, postoperative days 2 to 5) and can indicate which additional data would most improve the next prediction, drawing on value-of-information and active-sensing principles. Preoperative, intraoperative, and early postoperative data are fused, the anastomosis serves as a summed-risk marker of technique and physiology, and a continuously updated risk class drives an escalation-capable treatment pathway. To our knowledge, after searching PubMed and Embase (esophagectomy, anastomotic leak, machine learning, intraoperative, time-series; through June 2026), no published model integrates intraoperative time-series such as hemodynamics and indocyanine-green perfusion dynamics into a leak-prediction model; no AI performs automated postoperative endoscopic assessment of the anastomosis; and multimodal fusion with prospective external validation is largely absent. Explainable models already outperform classical regression (area under the curve 0.84 versus 0.74), and intraoperative time-series integration has improved prediction in adjacent surgical fields but not in esophagectomy. Addressing automation bias, the validation deficit, interoperability, and regulation (including Regulation (EU) 2024/1689 on artificial intelligence), we argue for an explainable, externally validated tool, initially esophagectomy-specific and deployed first as read-only decision support within established frameworks (IDEAL, DECIDE-AI).
Background Coronary artery calcium (CAC) causes blooming artifacts in coronary computed tomography angiography (CCTA), leading to overestimation of luminal stenosis and reduced diagnostic reliability in heavily calcified coronary arteries. Existing convolutional neural networks have limited receptive fields, whereas transformer-based models often incur high computational complexity and inadequate local feature preservation. Objective This study proposes CoronaryFusion-Net++, a calcium-aware hybrid deep learning framework for suppressing coronary calcium blooming artifacts while preserving coronary lumen morphology and anatomical continuity. Methods CoronaryFusion-Net++ integrates a Dense Encoder Module, a Cross-Scale Fusion Module (CSFM), a Calcium-Aware Swin Transformer Bottleneck, and an Attention-Guided Decoder with Adaptive Multi-Scale Upsampling. The network was trained and evaluated on the AIMI-COCA dataset using patient-level five-fold cross-validation. Performance was assessed using PSNR, SSIM, Dice coefficient, IoU, Precision, Recall, and computational complexity. Ablation experiments quantified the contribution of each architectural component. Results CoronaryFusion-Net++ achieved a Dice coefficient of 98.83%, an IoU of 97.69%, an SSIM of 0.986, and a PSNR of 39.84 dB, outperforming representative convolutional and transformer-based baseline methods. Furthermore, the framework demonstrated superior anatomical reconstruction and blooming artifact reduction, yielding a localized Lumen Boundary Precision (LBP) of 0.21 mm and an Apparent Stenosis Error (ASE) reduction of 18.42%. Ablation studies demonstrated consistent performance improvements following the integration of dense feature extraction, cross-scale feature fusion, transformer-based contextual modelling, and attention-guided reconstruction while maintaining computational efficiency. Conclusion CoronaryFusion-Net++ provides an effective framework for calcium blooming suppression and coronary lumen preservation in CCTA images. By combining local feature representation with global contextual modelling, the proposed architecture improves image reconstruction quality while maintaining computational feasibility. These findings indicate its potential to support more reliable stenosis assessment in heavily calcified coronary arteries and provide a foundation for future clinical validation and deployment.
Early diagnosis of skin cancer is crucial for improving treatment outcomes, increasing survival rates, and enhancing overall patient prognosis. An automated method utilizing clinical markers focused on cancer lesions can significantly improve the cancer diagnosis process. This study aims to leverage significant lesion markers identified by dermatologists using a Graph Neural Network (GNN), LesionGraph-Net, to enhance performance and analyze the complex relationships among biomarkers. The region of interest (ROI) from skin cancer images is segmented using K-means clustering, and biomarkers from skin lesions are extracted using computerized features. Among them, seven key features are identified based on feature selection methods. A graph is constructed from these features using the Pearson correlation formula, with the edge weight threshold determined through experiments at 0.2. LesionGraph-Net is then employed to identify the cancer, and its performance is compared to image-based transfer learning models and feature-based approaches, including GCN, 1D CNN, and various machine learning models. Our proposed model exhibits strong performance, achieving an accuracy of 98.62% and an F1 score of 98.43%. Compared to other methods, the proposed model consistently outperforms both image and feature-based approaches, underscoring the effectiveness of graph-based analysis in diagnosing skin cancer. The study highlights the benefits of integrating automated systems that incorporate significant clinical biomarkers. The retrieved clinical indications align with computerized features, aiding early diagnosis and cancer analysis for dermatologists. The study indicates that GNN-based models outperform conventional methods in identifying cancer, highlighting their potential for clinical trials.
Accurate patient stratification and disease subtype discovery remain challenging in high-dimensional and heterogeneous clinical data. This study presents Adaptive Spectral KNN Clustering with Multi-Objective Neighbourhood Selection (ASKC-MO), a spectral-clustering configuration that performs label-free neighbourhood selection given a predefined number of clusters. ASKC-MO constructs binary k-nearest-neighbour connectivity graphs and selects the neighbourhood size using a multi-objective internal-validity score combining the Silhouette Score, Calinski–Harabasz Index, and Davies–Bouldin Index. The seven-method comparison was conducted on the reproducible WDBC and Diabetes Progression datasets; the archived RAND Health Status analysis contains results for five classical baselines and ASKC-MO but not the later self-tuning spectral baseline. No method was universally best. On WDBC, Gaussian Mixture Model achieved the highest external label agreement (mean ARI = 0.7711), while ASKC-MO was competitive at its selected neighbourhood (ARI = 0.7548); k = 5 produced a higher ARI of 0.7793 but was not selected by the internal criterion. On Diabetes Progression, ASKC-MO and the Gaussian Mixture Model attained nearly identical mean ARI (0.1287 and 0.1285) but differed sharply in stability: ASKC-MO returned the same partition on every seed, whereas the mixture model varied from 0.0060 to 0.1527 across initializations. All methods performed weakly because the discretised regression targets were poorly separated. On RAND Health Status, no method recovered clinically meaningful labels. The principal finding is methodological: internal cluster-validity indices and external label agreement can favour different neighbourhood scales. The repeated-seed analysis therefore characterises initialization stability only, whereas case-level bootstrap, subsampling, and perturbation analyses provide the more relevant evidence about sample sensitivity. Prospective clinical validation remains necessary before any decision-support use.
This study presents a lightweight two-stage hybrid CNN–Transformer architecture for EEG-based seizure detection that replaces recurrent intermediate layers used in conventional hybrid models with a direct CNN-to-Transformer pipeline, reducing model complexity to 1.2 million parameters while enhancing classification performance. The model is evaluated through component-wise ablation, attention-based interpretability analysis, Layer-wise Relevance Propagation (LRP), and noise-robustness assessment under a uniform experimental protocol on two benchmark EEG datasets, with results reported consistently across both. The framework achieves 99.12% accuracy on Bonn EEG and 98.45% on CHB-MIT, outperforming reimplemented baselines under identical preprocessing and evaluation procedures; comparison against externally reported foundation models is included for context only, given sub-1% margins that fall within approximate confidence bounds. Ablation confirms a statistically supported 1.55-percentage-point gain over CNN + LSTM. Attention visualization highlights clinically relevant central channels (Cz, C3, C4, Pz), and LRP indicates the Transformer encoder contributes 60–75% of model relevance. The model maintains 97.2% accuracy under additive Gaussian noise at 5 dB SNR, a 7.7-point improvement over CNN + LSTM; this characterizes noise robustness rather than adversarial robustness in the formal sense, as no gradient-based or bounded-norm perturbations were evaluated. Results show HyConT achieved 96.12% ± 3.82% accuracy under LOSO-CV, confirming that Transformer-based long-range temporal modeling generalizes across patients. The ∼2.3 percentage-point degradation from segment-level (98.45%) to subject-level (96.12%) is expected and consistent with cross-dataset transfer, as both measure performance under distribution shift. These analyses support the architecture's efficiency, interpretability, and noise resilience as a basis for further development toward real-time EEG monitoring, with hardware-level validation and larger heterogeneous datasets identified as necessary next steps.
Deep learning models now match the accuracy of eye doctors in diagnosing eye diseases. But these models work like black boxes. This lack of openness is a big problem for use in clinics. Clinics need clear reasons for decisions. We introduce a new model called the Temporal-Multimodal Concept Bottleneck Model, or TM-CBM. This model fits glaucoma progression prediction. We built it using data from a large group of 87,342 patients. These patients had suspected or confirmed glaucoma. The data came from the BioArc system across many clinics in the country. Our model differs from usual end-to-end systems. It first predicts middle steps that doctors understand. Examples include trends in eye pressure, the cup-to-disc ratio, and defects in the retinal nerve fiber layer. Then it gives the final diagnosis. The model uses a shared goal for training. This goal mixes data over time on eye pressure, scans from optical coherence tomography, and basic patient facts. It reaches state-of-the-art accuracy. At the same time, it lets doctors step in during use. We show how doctors can fix the model’s middle predictions. They make changes by hand. This fixes wrong final answers. In this way, the model joins the sharp eye of data methods with the skill of human doctors.
Corneal Opacity is a major cause of visual impairment and blindness, which represents a quantitative and qualitative loss of corneal transparency resulting from pathological, biological and structural alteration in the cornea. Accurate assessment of opacity in the field of image-based corneal opacity is essential for diagnosis, surgical planning, and outcome prediction driven by developments in high-resolution anterior segment imaging and computational analysis. The main objective of this review is to synthesize current advances in corneal opacity assessment using multimodal imaging and data-driven methodologies. A structured analysis of reviews is based on research published in major scientific databases, including Web of Science and Scopus, covering conventional techniques and advanced imaging modalities, such as anterior segment optical coherence tomography, Scheimpflug densitometry, ultrasound biomicroscopy, confocal microscopy, and emerging high-resolution optical methods. The literature demonstrates a clear trend towards the use of objective, automated, and depth-resolved imaging metrics derived from the most consistently validated imaging modality. Deep learning, quantitative image analysis, multimodal fusion, artificial intelligence, and explainable modeling frameworks have also been examined in recent developments. Despite significant progress in research activity observed over the past decade in overcoming challenges related to data heterogeneity, limited standardized datasets, and insufficient external validation continue to hinder clinical translation. Such development holds an informed diagnosis, provides appropriate treatment options, and develops an accurate prognosis; the accurate assessment of corneal opacity is essential. Integrating multimodal imaging with interpretable artificial intelligence holds substantial promise to support corneal opacity management.
Low birth weight (LBW) is a leading cause of neonatal mortality in Ethiopia. Most studies treat LBW as a binary outcome, overlooking the clinically meaningful severity gradient from Extremely LBW to Normal birthweight. This study predicts LBW severity across four ordinal categories using multicenter data from 2586 neonates admitted to five Ethiopian referral hospitals. Five models were compared: ordinal logistic regression (OLR), random forest, XGBoost, support vector machine (SVM), and multilayer perceptron. Feature selection used LASSO and interpretability was assessed using KernelSHAP applied directly to the best-performing SVM model. Performance was evaluated using AUC, F1-score, quadratic weighted kappa (QWK), and ordinal mean absolute error (OMAE). The SVM achieved the highest cross-validated AUC (0.821 ± 0.018) and F1-score (0.411 ± 0.010), though the AUC advantage over OLR did not reach statistical significance (DeLong bootstrap test: z=1.19, p=0.235). On the held-out test set (n=518), SVM achieved AUC =0.806, F1 =0.393, QWK =0.439, and OMAE =0.498. KernelSHAP identified gestational age as the dominant predictor (mean |ϕ|=0.139), followed by hospital site, maternal age, and respiratory rate. Hospital-site features appeared consistently in the top 10 predictors across all four severity classes, indicating meaningful institutional variation beyond patient-level clinical factors. Ordinal logistic regression corroborated these findings, confirming gestational age (AOR =1.94) and infant sex (AOR =1.51) as the strongest clinical predictors. Leave-one-hospital-out validation confirmed model generalizability (mean AUC =0.814). These findings support the use of explainable ML for neonatal risk stratification in resource-limited settings, though external validation remains needed before clinical deployment.
Background Accurate differentiation between benign and malignant ovarian cancer on ultrasound remains challenging because of overlapping sonographic characteristics, severe class imbalance, and the tendency of conventional deep learning models to produce overconfident predictions. Reliable uncertainty estimation is therefore essential for developing clinically trustworthy computer-aided diagnostic systems. Methods We developed OvCaNet, an uncertainty-calibrated deep learning framework for benign–malignant ovarian tumor classification from two-dimensional ultrasound images. OvCaNet combines an ImageNet-pretrained EfficientNet-B0 backbone with a Convolutional Block Attention Module, a taxonomy-guided auxiliary head for eight-class tumor subtype learning, a Dirichlet evidential classification head, and split conformal prediction. The framework was evaluated on the public MMOTU OTU_2D dataset containing 1469 images from 294 patients. Patient-level stratified partitioning and five-fold cross-validation were used to minimize information leakage. Performance was compared with ResNet-18, ResNet-50, DenseNet-121, MobileNetV2, and EfficientNet-B0 using discrimination, calibration, uncertainty, efficiency, and explainability measures. Results OvCaNet achieved an accuracy of 97.86%, AUC-ROC of 0.989, AUC-PR of 0.935, sensitivity of 0.947, specificity of 0.981, and F1-score of 0.939. Compared with EfficientNet-B0, the strongest baseline, OvCaNet improved AUC-PR from 0.792 to 0.935 and sensitivity from 0.804 to 0.947. It also achieved the lowest Expected Calibration Error (0.013), Brier score (0.018), and negative log-likelihood (0.071). Five-fold cross-validation produced stable performance, while ablation experiments demonstrated complementary contributions from attention, auxiliary supervision, and evidential learning. Conformal prediction achieved empirical coverage close to the predefined targets, and Grad-CAM-based analysis provided qualitative evidence of clinically relevant feature localization. Conclusion OvCaNet provides accurate, computationally efficient, and uncertainty-aware ovarian tumor classification from two-dimensional ultrasound images. The results indicate that integrating subtype-guided representation learning, evidential uncertainty estimation, and conformal prediction can improve the reliability of AI-assisted ultrasound diagnosis. Validation on larger independent and multi-center cohorts remains necessary before clinical deployment.
Liver cancer is one of the most prevalent and life-threatening diseases worldwide. Accurate tumour segmentation from computed tomography (CT) images is important for diagnosis, treatment planning, surgical navigation and monitoring the disease. The CNN- and transformer-based segmentation methods show good results, but still fail to capture higher-order contextual dependencies, structural continuity, and the correct delineation of tumours with irregular boundaries, diverse textures, and low-contrast appearances. To address these issues, this study presents a hybrid deep learning framework called H2Former-Net, which combines hypergraph learning and hierarchical transformer encoding to achieve precise liver tumour segmentation. The framework comprises a hierarchical geometric decoder for precise feature reconstruction; a hypergraph-based contextual representation module to capture intricate semantic dependencies between feature embeddings; a cross-scale spectral attention mechanism to enhance the discriminability of multi-scale features in the frequency domain; and a topology-aware refinement module to preserve anatomical continuity and boundary consistency. We extensively evaluated the proposed methodology using segmentation and classification metrics on the publicly accessible LiTS-ISBI2017 and 3DIRCADb datasets. The experimental results demonstrate that H2Former-Net outperforms existing state-of-the-art approaches. The proposed model achieved an accuracy of 99.72%, a precision of 99.48%, a sensitivity of 99.26%, a specificity of 99.81%, and an F1-score of 99.37% on the LiTS-ISBI2017 dataset. It produced the best segmentation results for VOE, RVD, ASD, and RMSE, surpassing other approaches. The Dice scores for the LiTS-ISBI2017 and 3DIRCADb datasets were 0.968 and 0.972, respectively. This study demonstrates that H2Former-Net can be used to consistently and accurately segment liver tumours. This could considerably improve clinical judgement and computer-assisted diagnosis in the treatment of liver cancer.
Background A previous single-center observational study (Struja et al., 2026) proposed a glycemic target of 160-190 mg/dL for ICU patients with sepsis, consistent with Surviving Sepsis Campaign guidelines of ≤180 mg/dL. Objective The goal of this study was to replicate the joint longitudinal-survival model analysis of Struja et al. (2026) using MIMIC-IV, and to externally validate the findings using the eICU Database, a multicenter platform spanning 335 hospitals across the United States. Methods We conducted a retrospective cohort study using two independent ICU databases. Joint longitudinal-survival models were fit to estimate associations between time-varying glucose trajectories, organ dysfunction (SOFA score, with SpO2 as proxy for the respiratory component), and three outcomes of in-hospital mortality, mild hypoglycemia (<80 mg/dL), and severe hypoglycemia (<50 mg/dL). The MIMIC-IV replication cohort comprised 8006 patients meeting Sepsis-3 criteria; the eICU validation cohort consisted of 4866 patients identified by diagnosis-based sepsis criteria. Results Across both databases, higher time-weighted average glucose was consistently associated with greater SOFA burden and worse survival. The SOFA trajectory association parameter was strongly positive in MIMIC-IV (HR 2.187, 95% credible interval 1.991-2.415) and remained significant in eICU (HR 3.019, 95% CI 2.654-3.428), with differences attributable to temporal resolution and multicenter heterogeneity. Invasive mechanical ventilation and renal replacement therapy were associated with reduced mortality in both cohorts. Predicted hazard curves showed mortality risk increasing above 120 mg/dL while hypoglycemia risk decreased monotonically. Conclusions Our findings provide the first multicenter external validation of joint longitudinal-survival models for sepsis-related glycemic control in ICUs, providing directionally consistent multicenter evidence supporting the plausibility of a 160-190 mg/dL therapeutic range that balances survival benefit with hypoglycemic risk.
Breast cancer grading based on histopathological images plays a crucial role in determining tumor progression and guiding therapeutic decisions. However, manual diagnosis remains time-consuming and prone to inter-observer variability. To address this challenge, we propose a novel Hybrid Attention-Driven Mixture of Experts (HAMoE) architecture that integrates a custom five-block convolutional neural network backbone with Squeeze-and-Excitation (SE) and Convolutional Block Attention Module (CBAM) attention mechanisms within an advanced multi-head Mixture-of-Experts framework. The proposed model employs generalized focal loss to mitigate class imbalance and emphasize hard samples, alongside cosine annealing learning rate scheduling to stabilize convergence. Furthermore, Test-Time Augmentation (TTA) and Monte-Carlo Dropout inference are utilized to improve predictive reliability and uncertainty estimation. Experiments conducted on a clinically collected breast cancer histopathological dataset demonstrate that the proposed framework achieves superior grading performance compared with conventional deep learning architectures. The model attains up to 97.83% validation accuracy and exhibits robust generalization across different tumor grades. These findings highlight the potential of the proposed HAMoE framework as an effective computer-aided diagnostic tool for automated breast cancer grading and more consistent histopathological assessment in clinical practice.
Hepatocellular carcinoma represents a major global health burden, with liver cancer ranking among the leading causes of cancer-related mortality worldwide. In Thailand, it is the most prevalent among men and fifth among women. The Barcelona Clinic Liver Cancer (BCLC) classification system aids in treatment planning but faces challenges due to retrieving and interpreting burden from the fragmented information in medical records, resulting in inconsistent BCLC staging documentation. This proof-of-concept study explored the performance of large language models (LLMs) in BCLC staging using both standard prompting, Retrieval-Augmented Generation (RAG), and traditional natural language processing (NLP) baseline approach, focusing on the benefits of open-source models for enhanced data privacy and security. The experiment utilized clinical parameters from the electronic medical records of patients with primary liver cancer treated at the National Cancer Institute in Bangkok, Thailand, during 2024. We categorized experiments into three groups based on model sizes, with the GPT-5-mini and traditional NLP serving as the benchmark comparators. Our experiments found that the GPT-oss-20b model using RAG showed the best performance, achieving 87.69% accuracy with no statistically significant difference from GPT-5-mini. Models with over 20 billion parameters achieved accuracy between 48% and 71% with standard prompting and 56% and 87% with RAG, while smaller models (7-8 billion parameters) had lower accuracy of 9-56% with standard prompting and 36-53% with RAG, most of which failed to surpass the traditional NLP baseline (62.72%). RAG implementation improved performance in the majority of models by 3–39%, although two models (Qwen 8b and Deepseek 32b) showed marginal performance decrements, suggesting that RAG benefits are model-dependent. Therefore, for institutions with sufficient computational resources, open-source LLMs with RAG show promise as a potential alternative to both traditional NLP approaches and proprietary models. However, these findings are preliminary, and prospective multicenter validation is required before clinical implementation.
The creation of precise and dependable diagnostic models is essential because breast cancer continues to rank among the world's top causes of death. With the use of attention processes and a hybrid optimization strategy, this study offers a novel deep learning-based classification framework for the identification of breast cancer. In particular, the Convolutional Block Attention Module (CBAM) is incorporated into the suggested model to enhance the network's capacity to concentrate on the most pertinent spatial and channel-wise characteristics found in images of breast tissue. Several data augmentation strategies are used to improve data generality and variety. Additionally, to guarantee model resilience and avoid overfitting, a sophisticated 10-fold cross-validation technique is applied. A Sparrow search optimization technique is used to optimize performance and hyperparameter tuning. The proposed lightweight framework is well-suited for real-time clinical decision support systems, particularly in resource-constrained healthcare environments. By enabling accurate and efficient classification of breast lesions from mammography images, it can assist radiologists in early detection, reduce diagnostic workload, and support more consistent clinical decision-making, ultimately contributing to improved patient outcomes. Achieving an accuracy of 98.5% in its base configuration and up to 99.6% with 10-fold cross-validation, experimental assessments demonstrate that the suggested attention-augmented framework performs noticeably better than traditional designs like ResNet, VGG16, and GoogLeNet.