
Diabetic Retinopathy (DR) and Hypertensive Retinopathy (HR) are major retinal diseases contributing to global visual impairment, where early screening and progression monitoring critically depend on accurate retinal Artery/Vein segmentation and diameter estimation. To address limitations in current methods regarding vessel segmentation and geometric quantification, we propose a framework integrating both segmentation and measurement. The architecture includes three novel deep learning modules: the Hierarchical Feature Extraction and Integration Module (HFEIM) for capturing multi-scale vessel structures, the Adaptive Attention Pooling Module (AAPM) for emphasizing critical vascular regions, and the Multi-Scale Attention Residual Enhancement Module (MSAREM) to enhance detection of fine vessels. Following precise arteriovenous segmentation, we apply skeletonization to extract vessel centerlines, and reconstruct orthogonal elliptical cross-sections along the vessels. The final diameter is derived using the area-equivalent principle, ensuring physiological accuracy. Our approach demonstrates excellent segmentation accuracy and robustness across various arteriovenous segmentation datasets, offering a powerful solution to current challenges in retinal vessel segmentation and retinal diseases detection, these results demonstrate the effectiveness of the proposed model.
Predicting stock price movements is a longstanding challenge in financial research due to the market’s inherent volatility and the interplay of both quantitative indicators and qualitative sentiment. Recent advances in generative artificial intelligence (AI), particularly large language models (LLMs), have opened new avenues for integrating textual information into predictive models. In this study, we propose a novel attention-based framework that leverages LLM-generated text embeddings as predictive features, combining them with numerical financial indicators to forecast stock prices. Unlike conventional cross-attention approaches, our unified self-attention strategy, which processes concatenated multi-modal features as a single sequence, achieves superior predictive accuracy by enabling richer intra-modal interactions. To further capture the temporal dependencies within financial indicators, we incorporate a long short-term memory (LSTM) module into the framework. Extensive experiments on real-world financial datasets demonstrate that our model could achieve promising performance, highlighting the effectiveness of generative AI-driven textual representations in enhancing financial forecasting. This work underscores the potential of combining structured data with LLM-derived features through attention-based architectures for more robust and interpretable stock price prediction.
Synthetic Aperture Radar (SAR) ship detection is vital for maritime traffic monitoring. However, challenges such as speckle noise, complex sea conditions, and variations in ship sizes pose significant obstacles to traditional object detection methods. This paper proposes EEC-DETR, a novel approach that leverages a lightweight backbone network, multi-scale feature fusion, and wavelet convolution to enhance model performance. The EfficientVit backbone adopts a sandwich structure to reduce parameters, while the EIFI module integrates PSConv multi-scale convolutional layers and Efficient additive attention for global modeling. Additionally, the CCAM module combines WTConv wavelet convolution with a CAFM convolutional attention fusion mechanism to strengthen target feature extraction under complex backgrounds. Experimental results on the SSDD, HRSID, and iVision-MRSSD datasets demonstrate that EEC-DETR achieves a 5.5
Parkinson’s Disease (PD) is a movement-related disease characterized by the decline of dopamine-producing neurons in substantia nigra brain regions, which causes problems with movement control. Machine learning with neuroimaging techniques, especially Magnetic Resonance Imaging, provides a non-invasive automated approach for the diagnosis of PD, enabling both classification and identification of affected brain regions. The present work proposes a Multi-Regions-of-Interest ensemble network ( EnsembleRegNet ), a decision fusion approach that aggregates the predictive powers of different brain regions. By capturing complementary information from different regions and assimilating the subtle differences across regions, the model enhances the classification of PD. Under EnsembleRegNet framework, we proposed three ensemble models, wherein two utilise clustering followed by majority voting, while one uses a neural network-based ensemble model (Neural EnsembleRegNet ). The performance is evaluated on seven large, age- and gender-matched balanced datasets, derived from multiple publicly available datasets and stratified based on parameters such as gender, disease severity, and scanner strength. All three proposed ensemble methods demonstrated better performance than decision models trained on individual brain regions across all datasets. Among ensemble models, Neural EnsembleRegNet model outperformed for all datasets except one. The highest Area Under Curve (AUC) value of 83.0 EnsembleRegNet model. Also, an AUC of 72.9% is observed for the early-stage PD dataset ( HnY = 1 ). Furthermore, the proposed EnsembleRegNet framework is designed to identify imaging biomarkers, and the significant biomarkers it uncovers are consistent with the findings reported in the existing literature.
Diabetic retinopathy (DR) is one of the serious complications of diabetes, and if left untreated, it can severely damage vision. Therefore, early detection and accurate grading are crucial for the prevention and treatment of DR. Deep learning models can assist in the diagnosis and grading of DR, thereby reducing the screening burden. However, existing methods are often limited by inaccurate localization of abnormalities, large variations in lesion sizes, the susceptibility to losing tiny lesions, and the high cost of pixel-level annotation. To address these issues, we propose a saliency-guided multi-task attention network (SMANet) that jointly optimizes saliency segmentation and DR grading tasks. Extracting multi-level features through a shared encoder and designing multi-scale spatial attention block (MSAB) to enhance multi-scale context-awareness and global dependency modeling. Designing pixel-level saliency map segmentation task to encourage model to focus on lesion regions while preserving fine-grained features and mitigating pixel-level annotation costs. Further, the ordered regularization module (ORM) is introduced in the grading task to improve the grading accuracy by exploiting the ordering of disease labels. Experiments on the DDR and APTOS2019 datasets show that SMANet achieves grading accuracies of 78.40
Scene text detection methods based on segmentation have been widely employed in the field of text detection, which refers to locate text region and mark using text box in the scene images. While, for the diverse characteristics of text in medical scenes such as complex backgrounds, text with slanted, curved, dense, multi-directional and stamp interference text, existing detection methods still struggle to achieve ideal results. In light of this, this paper presents a refined model Bidirectional-Weighted Effective Multi-Level DBNet (BEM-DBNet) based on DBNet. The architecture consists of Global and Local Attention Feature Fusion(GLA), Bi-Weighted Feature Augmentation(BFA) and Multi-level Feature Fusion(MFF). The GLA is responsible for integrating attention mechanisms at both global and local levels. It enables the model to focus on the overall structure of the text while simultaneously capturing fine-grained details. This dual-focus approach enhances the model’s ability to detect deformed text in complex backgrounds more effectively and compensates for the lack of small-scale text. BFA is a cascaded effective weighted bidirectional feature pyramid network. It enables the network to assess the importance of different input features, leading to more refined feature selection and fusion. MFF is a feature fusion structure that uses enhanced features across scales, facilitating information sharing at different spatial levels. Additionally, a differentiable binarization post-processing module is used, enabling the segmentation network to adjust the binarization threshold dynamically. Validation on multiple datasets shows the effectiveness and feasibility of our method. Codes are available at: https://github.com/wennuan/BEM-DBNet
Speech Emotion Recognition (SER) plays a pivotal role in enabling natural and responsive human-machine interactions for real-time applications. While state-of-the-art SER systems leverage large self-supervised learning models (e.g., wav2vec 2.0, HuBERT) to achieve high performance, their computational complexity and resource demands hinder deployment in latency-sensitive scenarios. To address this challenge, we propose Self-Layer-Wise-Distillation (SLWD), a novel framework for efficiently distilling knowledge from transformer-based teacher models into compact, high-speed student models. Our approach preserves the teacher’s representational power while drastically reducing computational overhead. Experiments demonstrate that the SLWD student model achieves competitive performance with the teacher—using only 36 × inference speedup, making it ideal for edge deployment. This work bridges the gap between accuracy and efficiency in SER systems, advancing their practicality for real-world applications.
Medical image segmentation is crucial for clinical decision making and treatment planning. However, it faces two challenges: First, the structures of salient objects and background details vary significantly in medical images of different modalities. Second, the misleading co-occurrence of salient and non-salient objects and the noise interference at the edges affect the segmentation accuracy of the model. To overcome these challenges, we propose RoSPER-Net, a framework designed to enhance medical image segmentation. RoSPER-Net integrates a Spatial Prompt Encoder (SPE), which generates two complementary prompts using an advanced prompt mechanism to guide the model to focus on the local-global structure of salient objects and understand the overall background information in the image, thereby improving the model’s adaptability and segmentation accuracy under different modalities and complex backgrounds. Plus, our Cross-Scale Edge Enhancement Decoder (CSED) uses noise suppression and edge enhancement mechanisms to suppress non-salient regions and highlight salient regions, thereby improving the model’s ability to detect salient objects in complex backgrounds. Comprehensive evaluations of RoSPER-Net on 5 medical image datasets verify its superior performance and versatility, demonstrating its potential in the field of medical image segmentation. Our code is available on https://github.com/Anonymous2025Paper/RoSPER-Net .
Detecting small objects in UAV aerial imagery remains a challenging task, primarily due to significant information loss during feature downsampling in existing models, insufficient multi-scale feature fusion, and poor alignment between shallow and deep features. To address these issues, we propose an improved model based on YOLOv8s, named LBTA-YOLOv8. Specifically, we introduce a novel LBTM downsampling mechanism that incorporates both appearance similarity and spatial proximity to guide the feature compression process while effectively preserving key information of small objects. Additionally, we employ TAFNet to enhance the interaction between shallow detail features and deep semantic features. Within TAFNet, we design a 3D-GLFF module based on coarse-grained block attention to achieve better alignment and fusion of multi-scale features, solving the problem of feature misalignment and inconsistency in existing fusion methods. We also propose a three-level fusion strategy to comprehensively integrate information from different scales. Extensive experiments on two UAV object detection benchmarks demonstrate that our method outperforms existing mainstream approaches and achieves state-of-the-art performance. For instance, LBTA-YOLOv8 achieves 0.434 mAP50 and 0.257 mAP50-95 on the VisDrone2021 dataset.
Machine learning methods have recently made significant breakthroughs in precision agriculture. However, deep learning models need enough training data with expert annotations to train the model efficiently. It is reported that publicly available datasets for plant diseases may suffer from data and capture bias, which may lead to incorrect predictions. To address these concerns, we present a novel deep transfer learning approach to classify tomato diseases using a custom dataset and handling capture bias. We performed ablation study to choose the candidate model for our deep learning experiments. After selecting the best model, we used a Bayesian hyperparameter optimization framework to optimize the model hyperparameters. We employed 10-fold stratified cross-validation with 80
Predicting knee joint trajectories from surface electromyography (sEMG) signals holds a significant application value in various fields such as rehabilitation engineering and prosthetics control. However, existing prediction methods often struggle to achieve satisfactory performance due to limited dataset sizes and poor cross-subject generalization capabilities. In this paper, we propose an effective framework that integrates motion decoupling with a conditional diffusion model to address these challenges. Our approach decomposes knee joint angles into shared motion patterns across subjects and individual-specific amplitude parameters, enabling dual-task collaborative modeling that considers both commonalities and individual differences. Furthermore, the conditional diffusion model is employed to generate high-quality synthetic sEMG samples, effectively expanding the available data resources. Experiments conducted on data from 11 subjects demonstrate that our approach achieves a Root Mean Square Error ( RMSE ) of 3.59 ± 0.88 ^∘ , outperforming the non-decoupled model (4.61 ± 1.58 ^∘ ), the model without diffusion (4.85 ± 1.62 ^∘ ), the Bidirectional Long Short-Term Memory (Bi-LSTM) (6.75 ± 1.33 ^∘ ) and the traditional LSTM baseline (6.88 ± 1.59 ^∘ ).
While state-of-the-art language models achieve impressive results in code synthesis tasks, they encounter fundamental challenges when processing code review dialogues—a crucial component of collaborative software development. Review feedback typically incorporates contextual assumptions, informal developer communication, and technical nuance requiring sophisticated interpretation beyond literal code understanding. Conventional evaluation strategies utilize surface-level comparison metrics and remain susceptible to training corpus overlap, limiting their diagnostic value. We introduce a systematic evaluation methodology that analyzes code review interpretation through distinct cognitive dimensions: modification intent classification, affected region identification, and solution formulation assessment. Through reformulation as structured selection problems with graduated complexity levels, our approach provides detailed capability diagnostics while minimizing memorization influences. Comprehensive evaluation of 65 contemporary models across 900 expert-validated samples from diverse programming ecosystems reveals substantial capability variations and dimension-specific limitations not captured by traditional assessments. Our findings demonstrate that even advanced models exhibit systematic weaknesses in fundamental review interpretation, particularly in code structure navigation, identifying crucial opportunities for advancing code understanding systems.
Global climate change and agricultural complexity have exacerbated crop diseases, posing significant threats to food security and agricultural productivity. Deep learning-based approaches offer promising solutions for the intelligent monitoring crop growth in real time. However, existing models often encounter excessive computational complexity, and their fine-grained features extraction accuracy is insufficient, hampering their deployment in resource-constrained agricultural environments. This study introduces LAF-YOLO, a novel lightweight network that balances detection accuracy with computational efficiency. The LAF-YOLO architecture incorporates two key innovations: a novel lightweight detection head (LGHD) that significantly reduces model parameters and computational load, and an advanced feature encoding and reconstruction module that integrates Haar Wavelet transform technology to enhance feature extraction efficiency while preserving high-resolution details. Experiments on a self-built reliable field complex background dataset encompassing nine major diseases of four primary crops demonstrate that LAF-YOLO is remarkably efficient, with only 1.97 million parameters and 4.5 GFLOPs. Achieving state-of-the-art mAP50, mAP50-95, and recall of 97.4
Building flexible and efficient robotic platforms is essential for bridging the gap between simulation and real-world reinforcement learning (RL) applications. In this work, we introduce a hybrid robotic platform that integrates a three-axis linear slide with an OpenManipulator arm. This unified system is accurately modeled in simulation and seamlessly transferred to physical hardware, enabling consistent training and deployment of reinforcement learning policies across both domains. Based on this platform, we propose a novel RL framework named Enhanced Hindsight Experience Replay (EHER) to tackle the sparse reward problem in goal-conditioned tasks. EHER extends the standard DDPG+HER baseline by incorporating a two-fold training enhancement: subtask decomposition and expert experience replay. Specifically, we leverage the structure of the robot to decompose tasks into two coordinated subtasks: (1) using the linear slide to bring the end-effector near the goal region, followed by (2) fine-tuning the arm’s motion to precisely reach the target. Successful trajectories from both subtasks are selectively reused as expert demonstrations to guide future learning. Experimental results in simulation environments demonstrate that our method significantly improves sample efficiency and accelerates policy convergence, achieving high success rate in reaching and pushing tasks. The trained policy was subsequently deployed on a real-world robotic system to validate its sim-to-real transfer performance. These emphasize the necessity of co-designing reinforcement learning algorithms in conjunction with the physical and control capabilities of robotic systems to facilitate effective real-world deployment.
Metaphors are important rhetorical devices that appear frequently in both language and visual content and are commonly used in daily life. As social media and the internet have developed, internet memes have become a significant part of cultural communication, with metaphors playing a prominent role. However, current multimodal metaphor detection faces challenges, including low-quality meme texts and insufficient interaction between modalities. To tackle these issues, this paper proposes the MetaCRN framework for multimodal metaphor detection, which is based on a large language model. The framework designs a Linguistic Insight and Knowledge Augmentation module and uses the Deepseek large language model to deeply analyze meme texts, extract the source and target domains in the texts, and generate concise explanations to help the model understand metaphors in the texts. To further optimize text-visual information fusion, this paper introduces a dynamic feature replacement fusion strategy that facilitates information exchange and weighting via a dynamic feature replacement mechanism and a spatial gated feedforward network, enhancing modality interaction. Experimental results demonstrate that MetaCRN outperforms several baseline models on the public Met-meme dataset and the multimodal sarcasm dataset, confirming its superiority in multimodal metaphor detection.
Traditional precision irrigation methods often suffer from high operational thresholds and adjustment costs, as well as limited optimality and generalizability, making actual deployment challenging. In response to these limitations, this study proposes an autonomous framework for agricultural irrigation decision-making using artificial intelligence (AI) based on a large language model (LLM). The framework introduces a crop-model-based simulation data generation approach, integrating agricultural knowledge base with historical data. It uses a tailored retrieval-augmented generation (RAG) strategy and, innovatively, an LLM performance evaluation system built upon the WOFOST simulation model. By combining the interpretability of the white-box WOFOST model with the flexibility of the black-box LLM, this framework achieves a balanced integration of usability, scalability, and precision. Experimental results demonstrate that the proposed method reduces the average total irrigation volume by 0.24 cm compared to conventional irrigation strategies, while achieving a significantly shorter average decision time of 62 s, in contrast to 986 s required by traditional methods. It is seen that augmenting data with knowledge substantially enhances model performance and mitigates hallucination issues, with only a minimal increase in token usage and decision latency.
Detecting phenotypic anomalies in microscopy images, especially those caused by genetic mutations or environmental factors, is an important task in biomedical image analysis. We propose a novel unsupervised anomaly detection method, Self-Attention Parts Guidance (SAPAG), which leverages latent diffusion models trained on normal images. SAPAG operates in two steps: (1) it segments self-attention maps into semantically distinct regions using an unsupervised segmentation model (DiffSeg), and (2) it performs region-wise reconstruction during the reverse diffusion process guided by self-attention guidance (SAG). We validated SAPAG with microscopy images of Caenorhabditis elegans embryos. Wild-type embryo images were used as normal data, and RNAi-treated embryo images, which may contain phenotypic anomalies, served as test data. SAPAG outperformed representative anomaly detection methods, including PatchCore, RD4AD, and THOR, especially in detecting anomalies in embryo shape and size. Ablation studies further confirmed that both SAG and DiffSeg contribute to detection performance. Although subtle small structural anomalies (e.g. cell nuclear-level changes) remain challenging, SAPAG demonstrates strong potential for high-throughput phenotypic screening in microscopy-based analysis.
Fetal heart rate (FHR) abnormality detection is a critical task in perinatal monitoring, providing essential support for fetal health assessment and delivery decision-making. Existing methods are mostly limited to binary or ternary classification, lacking the ability to identify more complex abnormal patterns. This paper proposes a fine-grained classification framework based on a Time Series 3D Convolutional Neural Network (TS3DCNN), capable of recognizing seven FHR patterns: acceleration, prolonged acceleration, early deceleration, late deceleration, variable deceleration, prolonged deceleration, and background. The framework employs Fast Fourier Transform (FFT) to extract periodic features and introduces a Multiscale Convolutional Unit (MCU) to enhance both local and global feature representations. The experimental results show that TS3DCNN achieves an accuracy of 65.59
The generation of radiology reports involves extracting features from medical images and converting them into textual descriptions. Transformer-based models have demonstrated outstanding performance. However, several challenges remain: the flattening operation on grid features may hinder precise lesion localization, and research on the limitations of these models in distinguishing medical terms from non-medical terms is still insufficient. To resolve these issues, this paper proposes a Spatial Information Enhancement Module (SIEM) that improves visual representations by integrating relative geometric features between grids. Additionally, we design an Adaptive Language-Vision Feature Fusion Module (ALVFF), which dynamically computes the importance of visual and language features for vocabulary prediction. We integrate these two modules into the memory-driven Transformer architecture, constructing the Adaptive Language and Vision feature fusion Network (ALV-Net) for radiology report generation. Experimental results on two radiology report datasets, IU-XRay and MIMIC-CXR, demonstrate that our method outperforms previous state-of-the-art models and achieves significant improvements in key performance metrics.
Automatic Speech Recognition (ASR) provides a powerful tool for preserving endangered languages. However, the effectiveness of ASR systems relies on the availability of high-quality speech data, which is often scarce for many languages, including Balti. This study presents the first documented effort to create a spoken Balti corpus, introducing this endangered language to the speech research community and advancing its preservation. We compile the first-ever Balti speech dataset, comprising 10,394 voice recordings of 473 isolated words (primarily nouns) spoken by 39 native speakers, paired with transcriptions. Using this corpus, we evaluate ASR performance through statistical (GMM-HMM) and neural (TDNN) approaches, along with fine-tuned pretrained Whisper models. The experiments achieve word error rates (