Heart-sound recordings are highly susceptible to environmental and physiological noise, which complicates clinical interpretation and reduces the reliability of automated diagnostic systems. Effective denoising is therefore essential for preserving waveform morphology and enabling accurate feature extraction. This study proposes a heart-sound-specific wavelet approach and evaluates its denoising performance in comparison with conventional wavelets and support vector machine (SVM)-based methods. The method was assessed using publicly available datasets, including PASCAL and PhysioNet, which provide diverse normal and pathological phonocardiogram (PCG) recordings. Uniform and Gaussian white noise were added at varying signal-to-noise ratios (SNRs) to simulate realistic acquisition environments. Denoising performance was quantified using cross-correlation coefficients, SNR improvement, root-mean-squared error (RMSE), and mean absolute error (MAE). Results demonstrate that the proposed heart-sound wavelet achieved superior noise-suppression capability and a 7% performance gain over commonly used Db and Bior wavelets, while maintaining waveform integrity. Subsequent classification experiments showed that denoising quality directly influenced diagnostic performance: the model achieved 0.87 accuracy, 0.81 precision, and 0.83 sensitivity on the PASCAL dataset, and 0.997 accuracy, 0.946 sensitivity, and 0.944 precision on PhysioNet. These findings highlight the potential of tailored wavelet-based denoising to enhance automated heart-sound analysis and support more robust clinical and embedded diagnostic applications.
The incorporation of artificial intelligence (AI) in health care has created new opportunities for enhancing diagnostic precision, especially in the identification of respiratory diseases. This work examines the computational efficiency and generalizability of AI-driven lung sound analysis models, focusing on their performance across varied patient populations and datasets. The scalability and resilience of these models are rigorously assessed in practical clinical environments, addressing challenges such as demographic variability, data integrity, and environmental interference. Emphasis is placed on maintaining an essential balance between achieving high-diagnostic accuracy and maximizing computational efficiency, illustrating how effective models can enhance clinical processes and facilitate swift, precise decision-making across various healthcare settings. The findings underscore the potential of AI-driven lung sound analysis to revolutionize respiratory care, offering insights that can improve model adaptation and ensure equitable healthcare delivery among diverse patient populations.
Accurate and timely diagnosis of respiratory diseases remains a critical challenge in clinical practice, particularly in resource-limited and remote healthcare settings. This study proposes a web-based automated respiratory disease classification system leveraging a hybrid Convolutional Neural Network–Long Short-Term Memory with Time-Distributed (CNN-LSTM-TD) architecture for lung sound analysis. The proposed model integrates three complementary time-frequency representations—Mel-Frequency Cepstral Coefficients (MFCCs), Mel-spectrograms, and Chroma Short-Time Fourier Transform (Chroma-STFT)—to comprehensively capture both local spectral characteristics and long-range temporal dependencies inherent in respiratory cycles. Specifically, the TimeDistributed CNN block extracts localised acoustic features from sequential frames, while the LSTM layer models their temporal evolution, enabling robust identification of pathological acoustic signatures such as wheezes and crackles. The model was rigorously evaluated on the benchmark ICBHI 2017 dataset across six diagnostic categories: healthy, asthma, chronic obstructive pulmonary disease (COPD), pneumonia, upper respiratory tract infection (URTI), and bronchiectasis. The CNN-LSTM-TD model achieved an F1-score of 0.94, recall of 0.91, precision of 0.97, overall accuracy of 96.40%, and an AUC-ROC of 0.96, significantly outperforming standalone CNN, LSTM, and CNN-LSTM baseline models. The accompanying web interface supports audio file upload, real-time visualisation of waveforms and spectrograms, and confidence score reporting, collectively facilitating clinical decision support and telemedicine integration. These results demonstrate that the synergy of temporally aware deep feature extraction and accessible web deployment positions the proposed system as a clinically viable, scalable tool for automated respiratory disease diagnosis and remote patient monitoring.
For several centuries, research has been carried out to address respiratory ailments, which are among the most detrimental to human health. The advent of the stethoscope in the 19th century has facilitated the identification of respiratory sounds. This innovation represents a significant advancement in the identification and diagnosis of numerous respiratory ailments. In Malaysia, public hospitals have traditionally employed stethoscopes in their emergency departments. However, the precision of readings obtained through this method is susceptible to interference from ambient noise, uneven terrain, and suboptimal acoustic performance, particularly during medical transportation. Consequently, this can result in erroneous diagnoses and inappropriate treatment. Potential remedies for addressing the challenges associated with assessing respiratory sounds during medical transportation include advancements in stethoscope technology, novel auditory techniques, and reduced levels of background noise within the transportation environment. The present investigation concerns the effects of developing a new machine learning (ML) algorithm for the assessment of lung sound in conditions of high ambient noise. The objective is to devise a ML algorithm that can categorize acute respiratory illnesses based on their level of urgency in the presence of ambient noise.
This paper presents two novel datasets, Hajj-Crowd-2021, focused on crowd density and crowd anomalies during the Hajj pilgrimage. The datasets are derived from videos and images captured during the 2015-2019 Hajj and Umrah seasons, specifically in the Tawaf area surrounding the Kaaba. The crowd density dataset contains 30,000 images classified into five categories: very low, low, medium, high, and very high. The crowd anomaly dataset consists of 200 videos (100 normal and 100 anomalous) with a total of 60,000 frames. The datasets aim to address the lack of comprehensive and specific data for analyzing crowd behaviors during Hajj and Umrah. The paper describes the data collection and annotation process, as well as a comparison with existing state-of-the-art datasets. These datasets have the potential to support the development of crowd monitoring and management systems for large-scale religious gatherings.
Detecting small objects in aerial images poses several challenges, including issues with resolution limitations, scale variability, background clutter, and object occlusion. Annotated datasets for small objects in aerial images are often scarce, complicating the training and validation of detection models. This article introduces a new dataset specifically designed for small object detection in low-altitude aerial images. It addresses the challenges posed by shadows, including their impact on object visibility, by including images that capture small objects obscured with shadows. The dataset also features ground-truth shadow maps to support research in shadow detection. This dataset offers potential for future research and serves as a resource for transfer learning.
Lung diseases, such as asthma, pneumonia, and chronic obstructive pulmonary disease (COPD), pose considerable global health issues, making early diagnosis essential for effective treatment. Traditional diagnostic techniques for these ailments, especially auscultation using a stethoscope, are subjective and susceptible to inaccuracies. This study introduces CALMNet, a deep learning model that uses Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) networks, and a TimeDistributed layer to accurately analyze lung sounds by considering both their spatial and time-related features. The model is trained and assessed utilizing the ICBHI dataset, which encompasses a varied collection of lung sounds from patients with various respiratory ailments. Preprocessing techniques, including double denoising (FFT and high-pass filtering), were utilized to enhance the audio data quality, while feature extraction methods such as Mel Spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), and Chroma Short-Time Fourier Transform (ChromaSTFT) were employed to capture the fundamental characteristics of lung sounds. The results show that CALMNet achieves an impressive accuracy of 97.65%, with an F1-score of 0.909, precision of 0.911, and recall of 0.90, outperforming other models like CNNs and LSTMs. CALMNet exhibits enhanced efficacy in differentiating various lung illnesses, as evidenced by the confusion matrix and ROC curve analysis. The model’s exceptional accuracy and strong performance indicate its promise as an effective instrument for the automated classification of lung illnesses, providing a dependable, objective, and scalable option for clinical use. Future endeavors will concentrate on augmenting the dataset, enhancing model interpretability, and executing real-time diagnostic applications to further elevate its applicability in healthcare environments.
Agriculture is one of the main economic pursuits of the 64 Bangladeshi districts.: in Bangladesh, about seventy percent of the workforce rely on agriculture for their living. Bangladesh's gross national revenue is much enhanced by the country's rice farming but attacks of insect pests have a great impact on rice harvests. Different insect pests require different management measures, so the accurate identification of paddy field insects is a crucial task that allows, e.g., the application of the appropriate poison for specific insect pests. This will also prevent the wasteful use of ineffective insecticides. The main challenge addressed in this work is to detect and instantly segment small harmful insects in paddy fields. To address this problem, we have used deep convolutional neural network (DCNN) learning, based on Mask-RCNN, therefore enabling a technique for visual localisation and classification of agricultural pest insects. We have also developed our own dataset of harmful insect annotated images. In the proposed Mask-RCNN model we used a ResNet101 backbone, which can detect and segment at the same time. The proposed model achieves an AP@0.5 of 85.7%, a mAP of 63.8%, and an AR@10 of 68.5, therefore generating an anticipated accuracy of 75%. ResNet101 performs better on all measures. The suggested approach should be able to identify and classify the small harmful insects, with suitable accuracy, in real-world deployments.
This work contributes to the recent advancements in crowd management research, specifically concerning high-density settings, in the context of large gatherings. Video analysis and visual monitoring have become essential for enhancing the safety and security of pilgrimages worldwide. Multi-person posture estimation is crucial for several computer vision applications and has significantly progressed in recent years. Nonetheless, only few methods have tackled the challenge of pose estimation in congested settings, which remains difficult and inescapable in many scenarios. Moreover, current approaches do not provide adequate evaluation criteria for such situations. This study introduces a novel and effective method for tackling the challenge of posture estimation in extensive crowds, accompanied with a new dataset for enhanced algorithm assessment. Our methodology combines several computer vision methods supported by a Mask R-CNN model to precisely separate and evaluate multi-person postures, facilitating the automatic detection of behavioural patterns in large crowds. Our proposed method, with a ResNet101 backbone, on our HAJJ-Crowd videos dataset achieved 70.0 mAP. Our new HAJJ-Crowd video dataset can be used for assessment and testing purposes as it includes instance segmentation and prediction outcomes for several common methodologies.
The accurate and early detection of respiratory diseases is vital for effective diagnosis and treatment. This study presents a new approach for classifying lung sounds using a double denoising method combined with a 1D Convolutional Neural Network (CNN). The preprocessing uses Fast Fourier Transform to clean up sounds and High-Pass Filtering to improve the quality of breathing sounds by eliminating noise and low-frequency interruptions. The Short-Time Fourier Transform (STFT) extracts features that capture localised frequency variations, crucial for distinguishing normal and abnormal respiratory sounds. These features are input into the 1D CNN, which classifies diseases such as bronchiectasis, pneumonia, asthma, COPD, healthy, and URTI. The dual denoising method enhances signal clarity and classification performance. The model achieved 96% validation accuracy, highlighting its reliability in detecting respiratory conditions. The results emphasise the effectiveness of combining signal augmentation with deep learning for automated respiratory sound analysis, with future research focusing on dataset expansion and model refinement for clinical use.
With applications in animal monitoring, veterinary diagnostics, behavioral analysis, and robotics, real-time estimation of animal posture is an area of increasing interest in computer vision. In this work, we propose a method based on deep learning approaches to assess animal positions in real time. Selected for their applicability in both home and agricultural settings, the study centres on four animal classes: chicken, dog, horse, and cow. An important output of this work is a bespoke dataset created especially for posture estimation activities, including annotated videos. Every video in the dataset records continuous movement under different lighting and environmental contexts and runs for fifteen seconds. Keypoints marking important body joints for all types of animals were added to the extracted frames. Using the YOLOv8n posture architecture - which provides a balanced trade-off between speed and accuracy - we performed posture estimation. Although YOLO models are usually optimized for object recognition, we fine-tuned YOLOv8n-Pose to predict both bounding boxes and body keypoints, therefore enabling the real-time identification of intricate postural information. Trained on an annotated dataset using supervised learning, the model was tested on another test set from the same distribution. The proposed model achieves a PosePR mAP at the IoU threshold 0.5 of 99.5% in all classes, according to the experimental data. The dog class showed lower precision and F1 scores; the dog and horse classes showed decreased recall. The model maintains strong performance in real-time even if interclass posture variability and occlusion in video frames present natural difficulties. The system handles video input at an average frame rate enough for monitoring systems to be live. This study emphasizes the need for custom data sets tailored to real-world activities and the viability of employing YOLOv8n-Pose for the estimation of animal posture based on key points. Future directions include growing the data set, increasing keypoint accuracy, and including temporal consistency across frames. The data set is available at this link https://drive.google.com/drive/folders/1xci52bt9IxcYQrq36r2fQBaLvx3cSGHH#
In computer vision, identifying anomalies in crowded Hajj situations is challenging due to severe inter-object occlusions, varied crowd concentrations, and complex dynamics of the human mass. Two primary goals are addressed by this FCNN-based architecture: feature representation and wrong movement outlier identification. We proposed a crowd anomaly hajj monitor (CAHM) method using video sequences of crowded scenes to identify and localize abnormal behavior. The main contribution of our approach is the combination of anomaly detection-based optical flow features and classification based on spatial-temporal features using fully convolutional neural networks (FCNN), which has never been achieved in the past based on our extensive research in this domain. Using FCNN and spatial-temporal data, a pre-trained supervised FCNN detects (global) anomalies in Hajj crowded scenes. Additionally, this research intends to develop a new dataset based on the Hajj pilgrimage scenario in order to solve the issues. This architecture enables us, through spatial-temporal convolutions, to capture features in both spatial and time dimensions and to extract knowledge about the presence and motion of features encrypted in continuous frames. Two primary goals are addressed by this FCNN-based architecture: feature representation and cascade outlier identification. The suggested technique outperforms current methods in terms of detection and localization accuracy, as shown in the experimental results on the benchmark datasets. We employ the UCSD and Subway dataset in our study and present the Hajj-Crowd-2021 as a new video dataset. Extensive testing on the proposed Hajj-Crowd-2021 anomaly dataset demonstrates that it provides cutting-edge recognition performance and outstanding durability in crowd analysis. To verify our work, we used the available datasets to compare the proposed model to current models. In all of these situations, our model outperforms. Crowd analysis, for instance, improves classification accuracy by achieving 96
Global food security is seriously threatened by wheat leaf disease, which makes effective and precise disease detection and classification techniques necessary. For efficient disease control and the best possible crop health, timely identification and precise classification are essential. However, the limited availability of datasets for wheat leaf diseases hinders the development of effective and robust classification models. This research emphasizes the importance of precise wheat leaf disease diagnosis for global food security. The existing methods face challenges with limited data and computational demands. The research explores the potential of deep learning for automated disease detection, considering these challenges. CycleGAN proved to be the most effective among various augmentation techniques, enhancing the performance of classifiers DenseNet121, ResNet50V2, DenseNet169, Xception, ResNet152V2, and MobileNetV2. ADASYN also significantly improved classification accuracy, with MobileNetV2 consistently outperforming across different augmentation methods. This technique excels in overcoming challenges posed by limited datasets and class imbalances. Using CycleGAN for data augmentation notably enhanced classifier performance, addressing the scarcity of real-world samples. Evaluation through confusion matrix analysis revealed a minimal number of misclassified images—possibly as low as 0 to 3 images over the test dataset. The exceptional 100% accuracy achieved by the MobileNetV2 model on both CycleGAN and ADASYN augmented datasets highlights the potential of these techniques to unlock new levels of accuracy in wheat disease classification. This augmentation technique fine-tuned the classifier, reducing errors and highlighting the crucial role of CycleGAN in enhancing the accuracy and precision of wheat disease classification models. The proposed method establishes CycleGAN’s effectiveness in augmenting wheat leaf disease classification and recognizes ADASYN’s potential. The developed technique shows promise for automated disease detection in agriculture, enhancing global food security. Future research may optimize computational efficiency and explore integrating emerging technologies such as edge computing.
As healthcare systems transition into an era dominated by quantum technologies, the need to fortify cybersecurity measures to protect sensitive medical data becomes increasingly imperative. This paper navigates the intricate landscape of post-quantum cryptographic approaches and emerging threats specific to the healthcare sector. Delving into encryption protocols such as lattice-based, code-based, hash-based, and multivariate polynomial cryptography, the paper addresses challenges in adoption and compatibility within healthcare systems. The exploration of potential threats posed by quantum attacks and vulnerabilities in existing encryption standards underscores the urgency of a change in basic assumptions in healthcare data security. The paper provides a detailed roadmap for implementing post-quantum cybersecurity solutions, considering the unique challenges faced by healthcare organizations, including integration issues, budget constraints, and the need for specialized training. Finally, the abstract concludes with an emphasis on the importance of timely adoption of post-quantum strategies to ensure the resilience of healthcare data in the face of evolving threats. This roadmap not only offers practical insights into securing medical data but also serves as a guide for future directions in the dynamic landscape of post-quantum healthcare cybersecurity.
Low-altitude aerial images contain multiple shadowed regions, and a shadowed region frequently covers several non-homogenous regions of objects and surfaces. Using the information of neighboring shadow-free regions, a shadowed region can be efficiently compensated with reliable output. In this paper, using the transfer of statistical color properties, a new region-based shadow compensation approach is proposed. The proposed method demonstrates a significant improvement in the overall color mean of shadowed regions, increasing from 76.49 to 141.67, nearing the value of 172.51 observed in shadow-free regions. Additionally, the overall standard deviation of the shadowed region also improves from 31.01 to 43.29, approximately 93.68% matching the 46.21 standard deviation observed in the shadow-free region. The visual results are comparable to those achieved by state-of-the-art deep-learning methodologies and are reasonably acceptable for challenging aerial urban environments.
The global spread of SARS-CoV-2 has prompted a crucial need for accurate medical diagnosis, particularly in the respiratory system. Current diagnostic methods heavily rely on imaging techniques like CT scans and X-rays, but identifying SARS-CoV-2 in these images proves to be challenging and time-consuming. In this context, artificial intelligence (AI) models, specifically deep learning (DL) networks, emerge as a promising solution in medical image analysis. This article provides a meticulous and comprehensive review of imaging-based SARS-CoV-2 diagnosis using deep learning techniques up to May 2024. This article starts with an overview of imaging-based SARS-CoV-2 diagnosis, covering the basic steps of deep learning-based SARS-CoV-2 diagnosis, SARS-CoV-2 data sources, data pre-processing methods, the taxonomy of deep learning techniques, findings, research gaps and performance evaluation. We also focus on addressing current privacy issues, limitations, and challenges in the realm of SARS-CoV-2 diagnosis. According to the taxonomy, each deep learning model is discussed, encompassing its core functionality and a critical assessment of its suitability for imaging-based SARS-CoV-2 detection. A comparative analysis is included by summarizing all relevant studies to provide an overall visualization. Considering the challenges of identifying the best deep-learning model for imaging-based SARS-CoV-2 detection, the article conducts an experiment with twelve contemporary deep-learning techniques. The experimental result shows that the MobileNetV3 model outperforms other deep learning models with an accuracy of 98.11%. Finally, the article elaborates on the current challenges in deep learning-based SARS-CoV-2 diagnosis and explores potential future directions and methodological recommendations for research and advancement.
With an increasing number of people on the planet today, innovative human-computer interaction technologies and approaches may be employed to assist individuals in leading more fulfilling lives. Gesture-based technology has the potential to improve the safety and well-being of impaired people, as well as the general population. Recognizing gestures from video streams is a difficult problem because of the large degree of variation in the characteristics of each motion across individuals. In this article, we propose applying deep learning methods to recognize automated hand gestures using RGB and depth data. To train neural networks to detect hand gestures, any of these forms of data may be utilized. Gesture-based interfaces are more natural, intuitive, and straightforward. Earlier study attempted to characterize hand motions in a number of contexts. Our technique is evaluated using a vision-based gesture recognition system. In our suggested technique, image collection starts with RGB video and depth information captured with the Kinect sensor and is followed by tracking the hand using a single shot detector Convolutional Neural Network (SSD-CNN). When the kernel is applied, it creates an output value at each of the m $\times $ n locations. Using a collection of convolutional filters, each new feature layer generates a defined set of gesture detection predictions. After that, we perform deep dilation to make the gesture in the image masks more visible. Finally, hand gestures have been detected using the well-known classification technique SVM. Using deep learning we recognize hand gestures with higher accuracy of 93.68% in RGB passage, 83.45% in the depth passage, and 90.61% in RGB-D conjunction on the SKIG dataset compared to the state-of-the-art. In the context of our own created Different Camera Orientation Gesture (DCOG) dataset we got higher accuracy of 92.78% in RGB passage, 79.55% in the depth passage, and 88.56% in RGB-D conjunction for the gestures collected in 0-degree angle. Moreover, the framework intends to use unique methodologies to construct a superior vision-based hand gesture recognition system.