
Parkinson's disease (PD) is a progressive neurodegenerative disorder that leads to patients’ motor and non-motor disabilities, making early detection challenging. One emerging method for diagnosing PD involves analyzing hand-drawing patterns, such as drawing spirals, circles, or specific shapes. Recent advances in computational techniques, particularly machine learning and deep learning algorithms, have been applied to quantify and assess these drawings, allowing for the differentiation between healthy individuals and those with early-stage PD. In this paper, we propose a multi-modal deep learning model for Parkinson’s disease detection using circles, spirals, and meanders drawings from the NewHandPD database. The proposed approach involves three key stages. First, a dedicated deep learning model is assigned to the three modalities to extract features from the hand-drawn data. These extracted features are then concatenated and fed into a fully connected layer, followed by a softmax classifier for final prediction. To further enhance detection accuracy, the Binary Grey Wolf Optimizer is integrated into the training process to perform feature selection. This method evaluates different combinations of features during training and selects the best subset based on the accuracy maximization. Simulation results using the NewHandPD dataset demonstrated an improvement in Parkinson's disease detection through multimodal and multi-deep learning classification. Furthermore, the introduction of feature selection significantly improved accuracy and reduced the feature vector length.
Gait recognition has emerged as an important biometric modality due to its non-invasive nature and suitability for surveillance and security applications. However, achieving robustness under real-world variations in viewpoint, clothing, and carrying conditions remains a significant challenge. This paper introduces \textbf{AttIncGait}, a deep learning framework that integrates Inception-based multi-scale feature extraction with dual-path attention for effective gait recognition. Unlike prior methods that treat attention as an auxiliary or late-stage refinement, our approach embeds spatial and channel attention directly within Inception modules, enabling simultaneous multi-scale representation and adaptive relevance weighting. This structural integration enhances discriminative capability while preserving computational efficiency. Experiments on CASIA-B and OU-MVLP datasets demonstrate state-of-the-art performance: 97.5% accuracy on OU-MVLP and a 2.6% improvement over the best existing method under clothing variation in CASIA-B. Ablation studies further reveal that spatial and channel attention individually improve accuracy, while their joint integration yields an overall +8.5% gain on OU-MVLP. These results validate the effectiveness of attention-driven multi-scale fusion for gait recognition and highlight the potential of AttIncGait for real-world biometric identification and mobility analysis.
Visual relationship detection is crucial for semantic scene compre-hension, impacting various fields, including Human Behavior Analysis, Visual Navigation, Medicine, and Security. A key challenge in this domain is manag-ing partially visible objects and occluded object, which complicate the accurate detection of triplet relationships. This study introduces a novel approach to ad-dress these challenges by employing the ResNet152 model as the backbone for the faster R-CNN framework, paired with the Nadam optimizer – both of which have not been previously applied in this domain. Our methodology integrates region proposals and extracted features into a Graph Neural Network to con-struct scene graphs for each image. Additionally, we utilize a pre-trained Word2Vec model to encode subjects, objects, and predicate labels. To validate our approach, we conduct comparative studies evaluating the performance of ResNet152 against the widely utilized ResNet101 model and EfficientNet-B7 model, known for its efficiency in image classification, but have not yet been explored in this context. We also compare the Nadam optimizer with the com-monly adopted Adam optimizer. Our analyze focuses on the impact of these models on predicate detection and includes performance evaluations against current state-of-the-art method. Experiments on the open-access VRD dataset demonstrate that our ResNet152-Nadam combination achieves superior recall metrics, underscoring the importance of model depth and optimizer selection in enhancing predicate detection. This approach shows significant potential for advancing applications in visual relation detection, particularly in complex real-world scenario where relationships may often be unseen.
Early and accurate detection of diabetic macular edema (DME) is essential to avoid permanent loss of vision. This paper introduces Fusion-WideNet, a new hybrid classification model that combines handcrafted and deep features for the analysis of retinal OCT images. Handcrafted features—Gray Level Co-occurrence Matrix (GLCM), Histogram of Oriented Gradients (HOG), and Local Binary Pattern (LBP) are learned to extract local textural information and deep semantic features extracted from a pretrained ResNet50 convolutional neural network. These two feature sets are combined in high-dimensional space and fed into a Wide Neural Network (WideNet), which is a shallow but high-capacity network designed to handle large feature vectors. The model achieves a classification accuracy of 99.5%, outperforming traditional models and deep-only baselines. The proposed Fusion-WideNet framework not only demonstrates high diagnostic performance but also provides interpretability essential for real-world ophthalmic screening and clinical decision support.
Accurate 3-D reconstruction of garments from a single consumer-grade image remains a critical barrier to truly immersive and resource-aware virtual try-on systems. We introduce a self-supervised, multimodal pipeline that fuses visual tokens extracted by a Vision Transformer with textual garment descriptors to synthesise high-fidelity cloth geometry and texture while operating within the stringent power envelope of mobile neural-processing units (NPUs). A hybrid latent-diffusion module generates pseudo-meshes that supervise a lightweight INT8-quantised Mesh-Autoencoder, thereby eliminating the dependence on large annotated 3-D-scan corpora. To compensate for limited real data we construct SyntheCloth-300K, a dataset blending CLO-3D captures with PhysX-driven synthetic variations, and use it for joint visual–textual training. On the DeepFashion3D benchmark our method reduces Chamfer-Distance by 18% and improves SSIM by 0.03 over DressCode-NeRF, while sustaining 21 FPS at 0.32mJvertex−1 on a Snapdragon 8 Gen 3 — tripling the energy efficiency of prior art. Qualitative results reveal robust reconstruction of fine pleats and fabric drape, even under severe self-occlusion. The proposed framework thus bridges computer vision, physically based graphics, and embedded optimisation, laying the groundwork for next-generation, on-device virtual fitting applications.
Authentication plays an important role in managing security. Signature is one of the first broadly practiced method to authenticate an individual. However, existing research is solely based on the English signature detection and recognition with limited work on low-resource languages. Although being desirable for the document forensics and security purposes, it remains a challenging task to detect Urdu signatures in realistic settings due to different styles of the Urdu signatures and presence of noise, background, and other nuisance factors. Moreover, the lack of annotated datasets hindered the signature detection in natural environments. To address these challenges, this paper proposes an Urdu signature dataset consisting of more than 5000 official and real-world scanned documents of different genres, i.e., publicly available official government letters, feedback given by the public in relation to various departments, and civil court orders of the Sahiwal region. Furthermore, we proposed the YOLOv7 model for Urdu signature detection. The findings show that the proposed YOLO v7 is effective and accurate in detecting Urdu signatures even under different lighting conditions, background complexities, and signature distortions. The YOLOv7 model achieved the highest mAP@0.5:0.95 rate of 0.975.
Analyzing underwater images and improving their quality is a difficult task for researchers. Processing underwater images is extremely difficult because of the low contrast, noise, and blurriness brought on by light scattering and absorption. Since the longer wavelengths of sunshine are unable to get deep into the water due to salinity and an abundance of dissolved pollutants, underwater image becomes a crucial research topic, but the photos appear blurry, noisy, and faded. By combining bilateral filtering for decreased noise, saturation improvement for color restoration, and Laplacian sharpening for clarity, this study suggests a novel method for improving underwater image quality. The technique maintains structural integrity while improving visibility. Peak Signal-to-Noise Ratio (PSNR), Mean Squared Error (MSE), and Structural Similarity Index (SSIM) are used to assess the suggested approach, and the results show notable gains in image quality and clarity. According to the findings, this method improves underwater images, which makes them better suited for tracking the environment, marine object detection, and aquatic research.
Melanoma is a highly aggressive form of skin cancer that greatly impacts the global mortality rate related to skin cancer. Accurate identification and precise assessment of illness severity are essential for improving patient outcomes. The automatic classification of skin lesions by imaging is challenging due to the complex differences in their visual characteristics. This study employs deep learning algorithms to identify and distinguish between benign and malignant melanoma skin cancer. The malignancy is classified into seven distinct types: melanoma, melanocytic nevus, basal cell carcinoma, actinic keratosis, benign keratosis, dermatofibroma, and vascular lesion. The preliminary step of the proposed system entails a preprocessing stage where normalization and data augmentation techniques are utilized to prepare the HAM10000 dataset for the classification of benign and malignant cancer lesions. This research proposes a novel variant of the Residual Neural Network (ResNet), namely ResNet50V2.5, for enhanced picture categorization, optimizing training efficiency by circumventing unnecessary layers and improving model performance. A comparative investigation of five designs, including ResNet101, ResNet101V2, ResNet50, ResNet50V2, and ResNet50V2.5, demonstrates that the ResNet50V2.5 model attains superior accuracy, achieving a classification performance of 99.17%, thereby surpassing existing architectures in skin cancer diagnosis.
The Palm Leaf Manuscripts are a rich source of information about ancient India. It shares an enormous amount of knowledge about the past in terms of art, culture, literature and medicine. As the Manuscripts were developed organically, it is prone to getting damaged very fast. There are many mechanisms used to preserve the physical copies of the manuscripts, but because of the climatic conditions, the deterioration of the manuscripts is inevitable. This work outlines a comparative analysis of classical and deep learning-based approaches for denoising the distorted palm leaf manuscripts based on the segmentation quality of the text inscribed on the PLMs. The traditional pipeline consists of denoising, followed by binarisation and then segmentation of the entire image. We implemented this sequence using both Fast Non-Local Means and a self-trained Noise2Void (N2V) model for denoising. However, the segmented characters, particularly from the Fast NL-based approach, appeared visually distorted. In contrast, the N2V-based difference image showed better structural preservation and closer alignment with the ground truth. To tackle these limitations, we proposed a novel pipeline, which is an innovative processing pipeline that commences with denoising the Palm Leaf Manuscript images using the N2V model, proceeds with direct extraction of the text and culminates in the targeted application of binarisation exclusively on the segmented patches. This restructured approach minimises distortion, enhances text clarity, and preserves character details more effectively. Quantitative evaluation shows improved performance with lower MSE values (0.97, 1.15, 1.02), higher PSNR scores (27.17 dB, 26.61 dB, 29.09 dB) for various binarisation methods, and a structural similarity index (SSIM) of 91%, demonstrating the superiority of the proposed method over the traditional workflow.
The increase in technological advancements in unmanned ariel vehicle has lead to the challenges in the detection of drones in flight. The micro Doppler signatures obtained from radars is used to distinguish and detect different types of drones. Due to relatively similar radar spectogram image patterns or micro-Doppler signatures it is sometimes very challenging to classify different types of drones. Previously, Deep Learning methods like transfer learning and residual networks have been proposed to improve the classification accuracy. For further improving the classifying efficiency , this paper investigates the integration of channel attention mechanisms i.e. Squeeze and Excitation Net, Efficient Channel Attention and Gated Channel Transformation in the custom CNN Network (UAVDetect) with three publicly available micro Doppler spectogram UAV datasets. The paper proposes Modified SENet and Modified ECA which further improves the accuracy and better convergence.
Hyperspectral anomaly detection (HAD) poses a significant challenge as it requires modeling data with hundreds of measurements for each location in space. Many algorithms have been proposed to address problems in HAD, but most originate from one of several biases assumed of the data. This means that disparities between bias and variance can be observed among the algorithms in terms of their performance on individual datasets and more broadly across a diverse range of datasets. Ensemble learning enables the amalgamation of information across multiple biases to attenuate the trade-offs between bias and variance, improving individual dataset performance and generalizability across multiple datasets. Despite some work employing ensemble learning in HAD, amalgamating diverse HAD biases is an unexplored research direction. It is not clear whether amalgamating HAD biases improves performance, or what types and quantities of biases should be included and to what extent. To this end, this study employs 5 different ensembling methods to amalgamate 5 unique HAD biases to identify anomalies in 14 diverse datasets. The ensembling methods implemented consider equal, unequal, sparse, minimal, and mixed contributions among the biases. Results indicate that multi-biased models outperform single-biased models across all 14 datasets. In 12 of the 14 datasets peak performance was achieved by excluding or minimizing contribution to some of the biases, indicating that a mixture of sparse and minimal contributions was optimal. The results furnish empirical evidence as to the efficacy of multi-biased models to improve individual and generalized dataset performance, hence attenuating the bias-variance trade-off observed in the single-biased models. The results additionally provide direction for the most effective amalgamation strategies to construct optimal multi-biased HAD models.
Generative AI enables realistic data generation, making it essential for tasks such as text, image, and video synthesis. In robotics, predicting unknown map regions as an image generation problem is key to improving exploration and navigation. This study uniquely compares the performance of three deep generative models—Conditional Variational Autoencoder (CVAE), Vision Transformer (ViT), and Conditional Generative Adversarial Network (CGAN)—for map completion. Unlike prior studies focusing on individual models, this work provides a direct performance comparison for global map prediction. A custom dataset, derived from HouseExpo and enriched with structured time-series exploration data in Gazebo, was developed for this purpose. Experimental results demonstrate that CGAN achieves superior performance in completing unknown regions, making it more effective for robotic exploration. Additionally, we introduce a novel dataset generation methodology leveraging ROS, Gazebo, and Voronoi-based mapping. Future research will focus on scaling the dataset and integrating these models into exploration systems, further advancing generative AI applications in robotics.
This paper presents a novel approach of reconstructing topology of a deep learning model to reduce model’s trainable parameters, called Binary Feature Map-Splitting Architecture (BFMSA). The proposed approach is trained using the PlantVillage dataset for plant disease classification. A simple CNN-based BFMSA and various pre-trained models, such as InceptionV3, ResNet50, VGG19, and VGG16 models based on BFMSA, are experimented. The research has two main contributions. First, reducing the computational cost while building a CNN model from scratch based on BFMSA, where the reduction would be in the feature extraction and classification phase. Second, reducing the computational cost while building a transfer learning model, and the reduction would be in the classification phase. The study compares the proposed architecture with traditional architecture and evaluates performance using various metrics such as accuracy, loss, F1-score, precision, and recall. The findings indicate reduced overfitting and improved validation accuracy in the proposed architecture. The CNN model-based BFMSA achieved the highest accuracy of 98.31% on the validation set in comparison with traditional architecture. Whereas VGG16-based BFMSA achieved the highest accuracy among transfer learning models based BFMSA with a validation accuracy of 97.32%. Additionally, the proposed architecture decreases the trainable parameters by up to 87% compared to traditional models.
Sea turtle species identification is vital for marine biodiversity conservation, as sea turtles impact marine ecosystem balance by consuming dead seagrass and maintaining coral reefs. They help preserve the health of seagrass beds and coral reefs that benefit commercially valuable species. Therefore, to sustain sea turtle populations, detection systems that facilitate conservation efforts are essential. In developing underwater detection models, researchers must address several challenges specific to the underwater environment, including low illumination conditions, complex backgrounds, and underwater blur effects. In addition, YOLOv10-nano has emerged as the most efficient object detector in its family, though improving its performance remains a challenge. To overcome this issue, we propose an advanced deep learning approach using modified YOLOv10-nano with a new Parallel Fusion Module (PFM) integrated into the backbone alongside self-attention to enhance detection performance, named TurtleNet. The Parallel Fusion Module enhances detection performance by capturing channel-wise representational features. It emphasizes channels with relevant information through a dual-scaling process, improving feature quality. PFM is integrated into the untouched branch of the Partial Self-Attention mechanism to enrich the split half of the feature channels. Our model uses 48,302 images from Bunaken National Marine Park containing Green, Hawksbill, and Olive Ridley turtles with data augmentation applied. The method leverages YOLOv10-nano's real-time detection capabilities while the PFM optimizes feature fusion and localization accuracy. Experimental results show our model achieves an mAP50 score of 0.856 and runs at 28 FPS on CPU devices, outperforming existing approaches in precision, recall, and efficiency. This research combines computer vision with marine biology, creating an automated system that helps researchers and conservationists monitor endangered turtles.
Thermal Image Super-Resolution has become pivotal for security, autonomous driving, industrial inspection, and surveillance. This work consolidates six editions of the TISR challenge held within the Perception Beyond the Visible Spectrum workshop at CVPR from 2020 to 2025, detailing the evolution of tasks, datasets, evaluation protocols, and participation. The analysis traces a methodological shift from convolutional neural networks to transformer-based and hybrid architectures that better capture long-range dependencies, besides the emergence of cross-spectral guidance that leverages visible imagery to enhance thermal detail at large scale factors. Quantitative trends in PSNR and SSIM across editions show consistent improvement in results. Remaining challenges include robust cross-spectral alignment, computational efficiency for resource-constrained deployment, broader dataset diversity across conditions, and resilience to noise and environmental variation. The synthesis provides a unified reference for benchmarking progress and outlines actionable directions for future advances in thermal image reconstruction.
Diabetic Retinopathy caused by Diabetes Mellitus is a major vision threatening condition across the global population. Early detection and grading of diabetic retinopathy are pivotal to avoiding the associated vision impairment, and retinal image analysis serves as an effective method for the screening process. Analysing retinal images manually to detect diabetic retinopathy and grade them is a time-consuming pro cess that necessitates the involvement of experts to perform the classification. Computer-aided diagnosis of diabetic retinopathy through retinal image analysis is an effective tool to reduce the time involved in the screening process. This study proposed a convolutional neural network based Selective Feature Map Fusion architecture that leverages the fused feature maps of ResNet50 and EfficientNetV2L networks for the detection and grading of diabetic retinopathy. Feature set of ResNet50 and EfficientNetV2L captures different aspects of the underlying lesion distribution, resulting in a more comprehensive representation of the features. This approach helps the model generalize better to unseen data by incorporating diverse data aspects, thereby preventing overfitting to a specific feature set and developing a model that is more robust to noise. Selective feature fusion helps to reduce the computational overhead of the the feature fusion computation. Comparative analysis against the ResNet50 and EfficientNetV2L models individually revealed the superior performance of the fused model in both detection and grading of diabetic retinopathy. The proposed model attained an excellent level of accuracy of 95.2% in diabetic retinopathy detection with a sensitivity of 95.2% and a specificity of 95.2%, when tested on the IDRiD dataset. The model achieved a classification accuracy of 92% for diabetic retinopathy grading with a sensitivity of 88%, a specificity of 98% and an F1-score of 0.89. Moreover, the fused framework attained superior performance when compared against existing methodologies on the IDRiD dataset. The developed model exhibited robust performance also in the DeepDRiD dataset, and it can be employed successfully in the screening processes for the diagnosis and grading of diabetic retinopathy.
Chest disorders are widespread globally, encompassing conditions such as COVID-19, pneumonia, tuberculosis, and fibrosis. The diagnostic process often relies on chest X-ray (CXR) images, given the similarities in symptoms among these diseases. Manual diagnosis is a laborious and challenging endeavor due to the shared characteristics of these disorders. In contrast, leveraging deep learning technologies offers a more efficient and cost-effective approach to analyze CXR images for diagnostic purposes. This paper introduces an integrated model, utilizing both VGG16 and VGG19 architectures, coupled with Principal Component Analysis (PCA) and a feature fusion technique for the classification of multiple diseases. The model encompasses four classes: COVID-19, normal, pneumonia, and tuberculosis, making it suitable for real-time applications. The dataset employed in this study is sourced from the Kaggle repository. Our proposed model achieves an accuracy of 97.50\%, with a training time of approximately 4 seconds. Comparative analyses with other existing models are conducted to validate the effectiveness of the proposed approach.
In past few decades, handwritten verification system has got a lot of attention, but that's still a continuing process. Signature verification techniques are being used to determine whether such a signature is real or forged i.e., produced by an impostor. Many enhancements have certainly been suggested in the literature a most notable of which being use of Deep Learning algorithms to learn image attributes from signature images. As an explorative study, we trained models based on the principle of transfer learning using three state-of-the-art CNNs, namely VGG-16, VGG-19 and Alexnet as feature extractors and classifiers. The CEDAR Dataset, the ICDAR 2011 Dataset and the combination of both the CEDAR and the ICDAR 2011 Datasets are the three datasets considered for the proposed methodology. The models showed performance, with the highest precision reaching up to 100%. Among the three models, the Alexnet model exhibited the highest accuracy and lowest training cost.
The ability to detect moving objects is of great importance in a wide range of visual surveillance systems, playing a vital role in maintaining security and ensuring effective monitoring. However, the primary aim of such systems is to detect objects in motion and tackle real-world challenges effectively. Despite the existence of numerous methods, there remains room for improvement, particularly in slowly moving video sequences and unfamiliar video environments. In videos where slow-moving objects are confined to a small area, it can cause many traditional methods to fail to detect the entire object. However, an effective solution is the spatial-temporal framework. Additionally, the selection of temporal, spatial, and fusion algorithms is crucial for effectively detecting slow-moving objects. This article presents a notable effort to address the detection of slowly moving objects in challenging videos by leveraging an encoder-decoder architecture incorporating a modified VGG-16 model with a feature pooling framework. Several novel aspects characterize the proposed algorithm: it utilizes a pre-trained modified VGG-16 network as the encoder, employing transfer learning to enhance model efficacy. The encoder is designed with a reduced number of layers and incorporates skip connections to extract essential fine and coarse-scale features crucial for local change detection. The feature pooling framework (FPF) utilizes a combination of different layers including max pooling, convolutional, and numerous atrous convolutional with varying rates of sampling. This integration enables the preservation of features at different scales with various dimensions, ensuring their representa tion across a wide range of scales. The decoder network comprises stacked convolutional layers effectively mapping features to image space. The performance of the developed technique is assessed in comparison to various existing methods, including those by CMRM, Hybrid algorithm, Fast valley, EPMCB, and MODCVS, showcasing its effectiveness through both subjective and objective analyses. It demonstrates superior performance, with an average F-measure (AF) value of 98.86% and a lower average misclassification error (AMCE) value of 0.85. Furthermore, the algorithm’s effectiveness is validated on Imperceptible Video Configuration video setups, where it exhibits superior performance.
Corn is one of Indonesia's main food ingredients that contains the second largest source of carbohydrates after rice. Classification of the type and quality of corn seeds is still conducted manually by farmers. This procedure is time-consuming and can result in inaccuracies in sorting. Morphology has important characteristics to determine varieties such as size, color, area and seed shape. Some of these attributes, if measured manually, will take a long time and complexity that requires special expertise. The right way to describe these characteristics is to utilize machine learning. The machine learning used is CNN (Convolutional Neural Network). The CNN models used are ResNet101, Resnet50, VGG-19 and MobileNetV2. An analysis of the performance of the model was carried out using a confusion matrix. The results of the CNN model performance parameters for the classification of corn seed varieties with the ResNet101 model showed an accuracy of 89.8%, a precision of 86.9%, a recall of 88.3% and an F1-score of 86.4%. The ResNet50 model showed an accuracy of 86.27%, a precision of 83.2%, a recall of 84.1% and an F1-score of 83.4%. While the VGG-19 model showed an accuracy of 76.47%, a precision of 66.8%, a recall of 78.% and an F1-score of 71.1%. Meanwhile, the MobileNetV2 model showed an accuracy of 73.34%, a precision of 69%, a recall of 69.8% and an F1-score of 69.8%.