
This paper evaluates the impact of hybrid deep learning approaches on lung tumor segmentation by combining traditional image processing techniques with advanced AI-driven models. The study integrates Convolutional Neural Networks (CNNs) with preprocessing methods such as noise reduction, adaptive thresholding, and contrast enhancement to address challenges associated with complex anatomical structures and variability in medical image quality. A novel hybrid framework is proposed, leveraging traditional methods to preprocess data and improve input quality for deep learning models, ultimately enhancing segmentation accuracy and reliability.The effectiveness of the proposed approach is assessed using quantitative performance metrics, including Dice Similarity Coefficient (DSC), Hausdorff Distance, Jaccard Index, Precision, and Recall. Preliminary results indicate significant improvements in tumor boundary detection and reduced false-positive rates compared to existing methods. By streamlining segmentation workflows and enabling near-realtime applications in clinical settings, this research offers a pathway to improved diagnostic accuracy, treatment planning, and workflow efficiency.Future implications include the potential for integration into clinical imaging pipelines, fostering advancements in computer-assisted diagnosis and personalized treatment strategies. This study underscores the value of hybrid methodologies in addressing current limitations and paving the way for more precise and efficient medical image segmentation.
Cervical cancer remains a major public health concern in Vietnam, ranking second only to breast cancer among women. Early detection through screening, particularly HPV testing and Pap smears, is critical in reducing cervical cancer mortality. While deep learning has shown great promise in medical image analysis, particularly for detecting cervical cancer, challenges remain, especially when models are applied to real-world datasets with limited labeled data due to privacy concerns and expensive annotations. This research offers an effective cervical cancer diagnostic approach that combines Sliced Wasserstein Distance (SWD), Maximum Classifier Discrepancy (MCD), and the VMamba model. The VMambaDA model learns domain-invariant features and adjusts to domain shifts to address the problems of accuracy, generalization, and speed in cervical cancer screening. VMambaDA has proven to be more adept at handling the intricacies of medical images than earlier models by exhibiting better classification accuracy and sensitivity through extensive testing on both public and private datasets. This automated method could enhance the early detection of cervical cancer and intervention by pushing the limits of domain adaptation in cytopathology.
Line segment detection is a fundamental procedure in computer vision, pattern recognition, and image analysis applications. The paper proposes a novel method for wide line segment detection especially endpoints determination based on the Guided Scale Space Radon Transform and Hessian orientations. The method begins by determining the centerlines of wide lines and then exploit the image Hessian orientations around these lines to define binary region support of the line segments and then detect endpoints. The method shows to be robust against blur and noise on synthetic images where, the evaluation of the outcomes reveals the correctness of the detection by achieving low errors. In addition, results on real images are very promising.
Recent Use of Conditional Spatio-temporal Directed Graph Convolutional Networks(Cond ST-DGCN) [8] to represent human pose estimation has significantly helped in capturing varying non-local dependencies between limbs for different actions. This can be immensely helpful in Sports analytics where player pose plays key role in shot evaluation and can help in corrective action. In this article, we propose CondDGCN [8] based framework to explore use of spatial-temporal relation of batsman shot sequences (labelled and annotated 2D cricket dataset [1]) for Cricket shot action recognition by conditioning the graph network on batsman 2D poses. We achieve 97% accuracy for shot recognition and further explore visualization of conditional graph connections to establish importance of particular limbs for shots. The proposed framework uses fine-tuned 2D Pose estimator OpenPose [11] (fine-tuned for cricket dataset [1]) which in turn helps in easy adaptation of our solution to internet cricket videos for shot analytics.
Due to the increasing impact of climate change on agriculture, efficient early-stage control and yield prediction are becoming increasingly crucial. While satellite data has demonstrated significant efficiency in certain applications, deep learning solutions for plant phenotyping and UAV-based imaging have emerged as promising alternatives. Early-stage plant counting is a key strategy for yield forecasting and mitigating the risk of insufficient production. Manual crop counting, however, is time-consuming, expensive, error-prone, and labor-intensive. Automated, accurate counting offers a significant reduction in workload. This study proposes an automating plant counting method using aerial images of sorghum crops. A histogram equalization is used at the image level then a patch-wise counting is derived from dot annotations of patches using a complexity-optimized regression-based deep convolutional neural network (DCNN). The proposed approach estimates the number of plants within patches and aggregates these estimates for each image. Our findings demonstrate that this approach outperforms existing image level regression-based solutions on the same dataset, achieving a lower Mean Absolute Error (MAE) and provides similar results to more complex architectures that rely on pointwise localization while requiring fewer parameters.
The adoption of Decision Support Systems (DSS) by farmers faces significant challenges, including the need for user-friendly interfaces, comprehensive training, reliable support, and affordable pricing. Addressing these challenges is critical to advancing sustainable agriculture and meeting global food production demands. This paper introduces an advanced DSS tailored for aquaponics, leveraging cutting-edge technologies such as Artificial Intelligence (AI), the Internet of Things (IoT), and Blockchain. These technologies are integrated into the DSS to optimize resource management, ensure cybersecurity, and simplify complex processes through an intuitive interface. Real-time data is presented in an accessible format, enabling farmers of all skill levels to adopt the system with ease. The DSS features AI-driven robotic traps with 75% accuracy in real-time insect detection, autonomous robots achieving 94% precision in 3D spot spraying, and nutrient analyzers with over 92% accuracy in monitoring critical levels. A blockchain layer ensures secure data verification, traceability, and distributed AI model verification. Advanced hierarchical clustering algorithms analyze pest and nutrient dynamics, providing actionable insights for farm management. The system emphasizes interoperability and scalability, supporting seamless integration with various aquaponic setups. Field evaluations demonstrate the DSS’s capacity to reduce pesticide use by 50%, enhance crop yields, and lower sample analysis costs by 70%, highlighting its efficiency and sustainability. Data visualization latency remains below 430ms, enabling real-time responsiveness, while predictive models achieve 91% accuracy in forecasting pest population trends. These results solidify the DSS's role as a transformative tool in precision agriculture. By improving productivity, plant health, and the utilization of biopesticides and biofertilizers, this work bridges the gap between theoretical models and real-world applications. It illustrates the DSS's transformative potential for digital agriculture, offering a scalable and effective solution for sustainable farming practices globally. This study provides valuable insights into the broader applications of such systems, marking a significant advancement in agricultural innovation and technology integration.
In this study, we develop a new segmentation approach based on CycleGAN model to generate healthy lung images from pathological chest X-ray images, followed by image subtraction and binarization to produce a mask that includes pathology-affected areas. This approach enables the extraction of radiomic features from regions containing pathologies, enhancing disease classification. Our segmentation approach demonstrated effectiveness, improving AUC by 10.92% over conventional segmentation method for classifying effusion and infiltration using the XGBoost model, and outperforming previous studies. This study underscores the importance of precise pathological mask generation for accurate lung disease classification.
This work presents a novel semi-supervised dictionary learning framework that updates the dictionary by online learning and is efficient in utilizing the training data. The method employs a two-stage process to train the dictionary: initial training with limited labeled data, followed by online refinement using abundant unlabeled data. We introduce an adaptive correction weight to control the influence of new unlabeled data on the dictionary update based on its consistency with the current model estimate. This approach enables efficient use of the training data set. Moreover, results in faster dictionary convergence and improves data representation accuracy, especially in scenarios with limited training data. Experimental results demonstrate significant enhancement in the classification accuracy of the proposed method compared to the state-of-the-art semi-supervised dictionary learning methods, particularly when dealing with a limited number of training samples.
This research paper explores the development and implementation of ’NetraAI - The 3rd Eye,’ an AI-powered surveillance system aimed at enhancing public safety and security measures. The study investigates the technical architecture, real-time functionalities, and ethical considerations surrounding the deployment of NetraAI. It examines its impact on object detection, people tracking, vehicle recognition, and proactive threat detection in diverse surveillance scenarios. The findings highlight the system’s efficacy, ethical deployment, and potential contributions to the field of AI-driven surveillance.
Melanoma represents one of the most lethal forms of skin cancer, underscoring the importance of early detection for effective treatment and improved survival rates. Traditional diagnostic methods, which predominantly rely on visual inspection and biopsies, are often time-consuming and susceptible to human error. Timely diagnosis significantly enhances the likelihood of recovery and can reduce healthcare costs by minimizing the necessity for surgical, radiographic, or chemical treatments. Recently, deep learning techniques have demonstrated considerable promise in automating and improving the accuracy of medical diagnoses, including melanoma classification. In this study, we evaluate the performance of several state-of-the-art deep learning models—DenseNet, ResNet, VGG-16, VGG-19, Inception v3, and AlexNet—for multi-class melanoma cancer classification. Our objective is to identify the model that offers the best performance in terms of accuracy, sensitivity, and specificity. We conduct a comprehensive comparison using publicly available datasets, such as HAM10000 ("Human Against Machine with 10000 training images"), to ensure robust and generalizable results. This evaluation aims to advance the field of melanoma diagnosis by identifying the most effective deep learning approach, thereby facilitating early and accurate detection of this life-threatening disease.
The increasing prevalence of fatty liver disease necessitates accurate and efficient diagnostic methods. This study investigates the integration of deep learning techniques to enhance the diagnosis of fatty liver disease using ultrasound images. A dataset was utilized to train a deep learning model. The model achieved an impressive accuracy of 96% on the training dataset for distinguishing patients with fatty liver disease from healthy individuals, while maintaining a commendable accuracy of 90% on a completely new dataset. Furthermore, the model demonstrated a sensitivity of 92% in classifying different levels of liver fat in the training dataset, with an accuracy of 83% on the new dataset. These results underscore the effectiveness of combining deep learning in medical imaging, providing a robust framework for the early detection and classification of fatty liver disease. The findings suggest that this hybrid approach can significantly improve diagnostic accuracy, ultimately contributing to better patient management and outcomes in clinical practice. Further research is warranted to validate these results across diverse populations and enhance model generalizability.
Noise as an unwanted interference can significantly degrade speech signals, especially those recorded by many microphones. This interference is modeled as additive noise that originates from a range of sources including White Gaussian Noise (WGN), babble, crowd, large city, and traffic noises. These disturbances can alter the characteristics of speech signals reducing both their quality and intelligibility. This paper introduces a novel approach designed to reduce noise and enhance the quality and intelligibility of speech signals. The proposed method combines Wavelet Transform with Adaptive Filters, specifically the Wiener filter and RLS filter. The evaluation process involves testing noisy speech signals under realistic conditions with different signal-to-noise ratios (SNRs) and different types of additive noise. The objective measure is used for evaluation, including the perceptual evaluation of speech quality (PESQ). Results show that combining Wiener or RLS filtering with Wavelet Transform significantly improves noise reduction, outperforming the use of Wavelet Transform alone.
object detection based on event vision has been a dynamically growing field in computer vision for the last 16 years. In this work, we create multiple channels from a single event camera and propose an event fusion method (EFM) to enhance object detection in event-based vision systems. Each channel uses a different accumulation buffer to collect events from the event camera. We implement YOLOv7 for object detection, followed by a fusion algorithm. Our multichannel approach outperforms single-channel-based object detection by 0.7% in mean Average Precision (mAP) for detection overlapping ground truth with IOU = 0.5.
The demand for adaptive learning in education is an important application for Computer Vision (CV)-based detection models which extract learners’ faces to classify engagement. However, a loss in visual data due to the partial occlusion of learners’ faces imposes challenge for practical use-cases. In this paper, we propose an occlusion-aware framework to improve the robustness of non-occlusion-aware models in engagement detection. Firstly, our framework consists of an occlusion-aware data augmentation pipeline that aims to simulate the types of partially occluded faces in-the-wild. Secondly, we investigate the application of masked auto-encoders (MAE) for occlusion recovery. Extensive experiments have been performed to measure the effectiveness of the proposed occlusion-aware framework. Specifically, on the challenging FER-2013 dataset our occlusion-aware data augmentation achieves a 5.24% improvement in accuracy on baseline models and a 1.54% improvement against models with augmentation methods used in the state-of-the-art. Alongside the application of MAE, our proposed framework improves baseline classification accuracy by 13.3% and 4.59% on the FER-2013, and DAiSEE dataset for engagement detection. Our investigation into occlusion-aware engagement detection using the DAiSEE dataset provides important insight into the limitations of occlusion-awareness in engagement detection.
Recent advancements in deep neural networks have shown remarkable improvements in image quality during the demosaicking process, surpassing conventional algorithms. However, these deep neural network techniques are often characterized by heavy computational requirements, rendering them unsuitable for deployment on resource-constrained platforms. This presents a critical challenge in the field of image demosaicking: while deep learning approaches excel in enhancing image quality, their computational intensity poses a significant hindrance to their wider adoption. Consequently, there is a pressing need for methodologies that can strike an optimal balance between achieving superior image quality and maintaining computational efficiency. In this work, we propose a new deep framework, hyper-prior dependent demosaic neural network, HPDNet that utilizes the significant concepts of the conventional algorithm and the characteristic of image sensor data. We designed the network that exploits three concepts, those are pixel gradient prior attention, phase separation, and multi-level sparse and dense feature extraction. We designed the deep neural network that extracts the optimal gradient prior and the multi-level extracted features are fused and attended by gradient prior. It can fully utilize spatially variant information. Also to make the network deployed in the mobile platform, we devised self-pruned image convolution that adopts image filter characteristic and reduce computations. Experiments show that proposed network outperforms SOTA demosaic networks both in terms of image quality and computation.
Activity recognition is one of the major tasks in computer science and engineering that aims to recognize and understand the actions of one or more agents from a series of observations. It assists different sectors in ensuring safety in human-machine interaction. This paper investigates and tests an encoder classifier based on an LSTM model with various attention mechanism structures over a synthetic time-series motion dataset generated with the NVIDIA Isaac simulator. The simulator simulates the motion of an excavator as an example of a complex kinematic system in a construction environment, gathering data from position sensors connected to the excavator’s main joints. The results show that fewer sensors can be used for certain types of motion classification with a large accuracy range between 66.7% and 83.3%. The encoder-LSTM model with scaled dot-product attention gives the most accurate results compared to other attention mechanism types, with around 33.4% when using data from only two sensors and around 22.2% for the four main sensors. These results are in comparison to models that used other attention mechanisms.
We introduce a novel method to address intra-class imbalance in 3D point cloud segmentation of wheat, focusing on distinguishing between ear and non-ear parts. Variability in plant structure, influenced by factors such as curvature and shape, often leads to data imbalance which complicates segmentation tasks. Our approach utilizes Monte Carlo Dropout to identify and prioritize uncertain samples at the end of each training epoch, employing uncertainty-driven sampling to select samples with the lowest confidence. These samples undergo augmentation through scaling and leaf crossover techniques, enhancing their representation in the training set. Our comparative evaluations demonstrate that this strategy significantly improves the mean Intersection over Union (mIoU) and segmentation accuracy, thereby increasing model robustness for complex 3D plant structures.
This paper presents an advanced encryption algorithm specifically designed to enhance the security of volumetric medical image data, crucial for the Internet of Medical Things (IoMT). The algorithm encrypts a stack of 256 images, each with dimensions of 256×256 pixels, through a meticulous multi-stage process. It begins by segmenting an image cube into Red, Green, and Blue channels, which are then encrypted through three phases: XOR operations with keys from an 8D hyperchaotic system, substitution with S-boxes derived from Gold sequences, and a final transformation using Fibonacci Q-matrices. This approach significantly improves security by improving entropy, decreasing cross-correlation, and strengthening resistance to statistical attacks, with various cryptographic seeds used at each stage to enhance the robustness of the encryption. Specific performance evaluation metrics include a pixel cross-correlation of approximately 0, an NPCR of 99.61%, a UACI of 34.45%, an entropy of 7.998, and a key space that exceeds 22870. Extensive cryptographic assessments confirm the effectiveness of this algorithm, making it a vital tool for securing medical image data within IoMT, ensuring safe transmission and storage in healthcare systems.
Automated white blood cell segmentation is crucial for diagnosing and monitoring severe conditions like leukemia and lymphoma. However, this task is hindered by the scarcity and quality of available public datasets. To make better use of these limited datasets, a semi-supervised learning framework is required. In our study, we introduce the C2GMatch framework, which leverages weak-to-strong dual-view guidance to enhance segmentation performance from partially labeled data. Our approach surpasses two out of three leading semi-supervised segmentation methods on three public datasets, including BCCD, LISC, and LiveCells, achieving IOU scores of 54.12% and 68.13% and Dice scores of 69.97% and 82.01% on LISC and LiveCells, respectively. Our framework also shows better segmentation masks compared to previous works such as Pseudo-Seg or Unimatch in qualitative results.
Given a template image T, the task in template matching is to search for T in a larger image region S. While simple in concept, much effort in recent years has been channeled on improving its robustness to complex transformation, with attention being largely on the similarity measures. Much lesser attention has been on the efficiency of the search computation itself. This paper attempts to fill the gap by proposing Pseudo-Normalizable Fourier Transform (PNFT), a novel complex transform that enables efficient transform-invariant, vectorized template matching. The concept of normalization is overloaded to denote the transformation of a function over an arbitrarily located patch to that about the origin, enabling the efficient use of image integral to perform template matching. The paper discusses on the ’normalizability’ – whether the function can be normalized – and further introduces pseudo-normalization as an approximation. The approach proposed requires no learning and runs fast on lower-end CPU-only platform. Further, it is conceptually simple and can be implemented in approximately 200 lines of Python/Numpy code. The efficacy of the approach is empirically analyzed using a complex example image.