Mask-wearing detection plays a crucial role in controlling and preventing the spread of infectious diseases. However, this task presents several challenges due to the occlusion of facial features, the diversity of mask designs, and environmental factors such as lighting and camera angles. In this study, we propose an intelligent automated system capable of classifying three distinct mask-wearing states: correctly worn, not worn, and improperly worn. A significant contribution of this work is the enhanced identification of improper mask usage, which is often a major challenge for manual monitoring and existing automated systems. Additionally, the system integrates a real-time distance measurement algorithm to identify and alert when social distancing violations occur. The model was trained and evaluated on a dataset of 1,655 manually annotated images collected under diverse real-world conditions. Experimental results demonstrate that the proposed system achieves a classification accuracy of 97.22
The classification of tomato leaf diseases is critical for advancing smart agriculture; however, deep learning model performance is often limited by low-data regimes. This study introduces and evaluates a Combined Data Augmentation (CDA) strategy that integrates traditional augmentation techniques (TDA) with data generated via a Stable Diffusion model (GDA) to improve classification performance under severely limited data conditions, specifically with only 20 training images per class. Experiments were conducted on subsets of the PlantVillage ( D^PV ) and PDR2018 ( D^PDR ) datasets, utilizing EfficientNet-B0 as the backbone across three training configurations: transfer learning with partial freezing, training from scratch, and fine-tuning all layers. The statistical significance of the results was assessed using the Wilcoxon signed-rank test with Holm correction and Cohen’s d_z effect sizes across 15 cross-validation folds. Findings indicate that Combined Data Augmentation (CDA) provides its most decisive benefit under the from-scratch configuration, rescuing the model from training collapse on both datasets ( D^PV : +21.95 percentage points, p_holm = 0.009; D^PDR : +7.91 percentage points). Under pretrained configurations, the improvement is modest on D^PV and absent or negative on D^PDR , where augmentation strategies including GDA were found to significantly reduce accuracy. Further analysis demonstrates that the effectiveness of generative augmentation is highly dependent on the Strength parameter; Strength = 0.35 is the only regime preserving performance parity with the Baseline, while higher values introduce semantic drift and statistically significant performance degradation. A learning rate sensitivity analysis at η _0 = 10^-3 confirmed that these conclusions are stable across hyperparameter choices. This study provides empirical and statistical evidence characterizing the conditions under which combined data enhancement strategies are beneficial, neutral, or detrimental in limited data contexts, and informs the development of robust plant disease diagnostic systems.
Fire and smoke detection in image data is a crucial application of computer vision; however, practical deployment in real-world environments remains challenging. While previous research has made significant progress; however, detection systems often struggle to maintain stability when faced with changing lighting conditions, partial occlusions, and the complex visual characteristics of indoor spaces. This study proposes an automated fire and smoke detection system based on the Real- Time Detection Transformer (RT-DETR) architecture. Unlike traditional models that focus solely on accuracy, this system is engineered to address the practical need for early warning, achieving a high recall of 91.6 % to minimise missed fire events. The system is designed for versatile integration into existing CCTV surveillance, Smart Building ecosystems, and IoT/Edge. Evaluated on the Home-Fire dataset using a five-fold cross-validation strategy, the model achieves an mAP@0.5 of 94.5 %. These results demonstrate that the proposed system offers a robust, scalable, and reliable solution for real-world fire safety monitoring, providing a robust foundation for autonomous fire safety monitoring and rapid emergency response.
Chest X-ray lesion detection remains challenging due to severe class imbalance, subtle lesion appearance, and the risk of over-optimistic evaluation caused by improper data splitting. In this study, we propose a sensitivity-oriented detection framework based on YOLOv11 for robust chest X-ray screening under clinically realistic conditions. The proposed approach integrates patient-wise data partitioning, enhanced data augmentation, and prediction fusion to improve generalization while mitigating data leakage. Experiments are conducted on the VinDr-CXR dataset using a strict patient-level split to ensure full separation between training and validation sets. A series of internal fine-tuning scenarios is designed to analyse the trade-offs among precision, recall, and localization accuracy. Based on internal validation, the medium-scale YOLOv11-m configuration (denoted as M3) is selected as the reference model, as it provides the most stable balance between sensitivity and localization performance. Under rigorous evaluation, M3 achieves a precision of 0.431, a recall of 0.416, an mAP@0.5 of 0.387, and an mAP@0.5:0.95 of 0.193. Compared with representative baselines, M3 demonstrates improved robustness under patient-wise evaluation, outperforming transformer-based DETR by a large margin (mAP@0.5: 0.387 vs. 0.232) and achieving performance comparable to YOLOv7 while exhibiting substantially higher sensitivity to small and diffuse lesions. Further comparison with recent studies shows that the proposed method achieves higher overall mAP@0.5 (0.387 vs. 0.362-0.378) while improving detection performance on clinically challenging abnormality classes. These results indicate that the proposed YOLOv11-based framework provides a reliable and clinically meaningful baseline for chest X-ray lesion screening and future methodological advancements.
Amid a rapidly developing era, people can inevitably have problems with stress, depression, pressure, or difficulty sleeping due to frequent overthinking. To overcome the above problems, yoga will be an excellent solution to help adjust thoughts and harmonize body and soul, helping us relax, relax the mind, and retain positive thoughts. Negative and evil auras will be pushed away, and the worldview will improve. Yoga practice has incorrectly caused many unwanted injuries for practitioners. Therefore, we present an approach grounded in skeleton-based feature extraction and neural networks to find a solution to the recognition of yoga postures, creating a premise for researching a smart virtual trainer that supports home workouts for users from input image data converted into skeleton data through MoveNet. The classification models were used to train recognition and classification of yoga poses. The models were trained and evaluated on a dataset of 3939 images of 10 yoga poses. Experimental results show that the proposed algorithms are entirely suitable for the classification task when achieving good results on different metrics such as Precision, Recall, F1-score, and Accuracy.
In recent years, the gold market has witnessed one of the most volatile gold price eras ever. For decades, developing an accurate gold price prediction model has been a complex problem for researchers. Within the scope of this research, we apply several machine learning methods to the problem of forecasting time series data using a combination of three models, including Long short-term memory (LSTM) - Convolutional Neural Network (CNN) - Random Forest Regression (RF) to forecast future gold prices. Two performance metrics, including Mean Absolute Error (MAE) and Mean Squared Error (MSE), are used to evaluate the performance of the different models developed. Experimental results show that the proposed Random Forest model has outstanding performance in predicting and changing trends of gold prices in both the short and long term, with the lowest MAE and MSE of all three models. In general, RF has improved prediction accuracy, which is also a potential bright spot for building an increasingly optimal prediction model to forecast gold prices and contribute to forecasts in other fields.
Educational activities have also strongly promoted the use of artificial intelligence to improve efficiency. Online exams from Internet Service-Based Courses have become popular and may become a trend to prevent disease transmission and facilitate international students. However, from there, the question of how to detect cheating through online exams also arises. In this study, we collected videos from some Vietnamese colleges to deploy deep learning techniques for cheating detection in e-exams. We take advantage of the Mediapipe to extract skeletons from videos, and then, such skeletons are fetched into the model training phases. Through that, we also propose a Long Short-Term Memory Network (LSTM) for learning the extracted skeletons. In this way, the model contributed an astonishing accuracy of over 85
During information technology development in most fields, music also develops rapidly and has many diverse genres. With the increasing number of songs, finding a favorite song becomes more and more complicated when we need help to remember the name or genre of that song clearly. Our study aims to seek songs based on analyzing the characteristics of an audio clip with lyrics or analyzing data about words-key (lyrics) provided by the user. More specifically, this study has attempted three approaches. The first method uses Google Speech to Text Application Programming Interface (API) with the speech recognition library to output text from the user’s audio inputs or to directly enter text to search. Then, we applied the Inverted Index structure to process and store the original lyrics text. The second method is to extract audio features, Mel Frequency Cepstral Coefficients (MFCC), and then leverage the audio-Approximate Nearest Neighbors (ANN) algorithm to support neighborhood search. The third approach is Audio Fingerprint, used to identify and classify audio segments by converting the audio signal into a unique data string, also known as a hash function. The experiments are evaluated on Vietnamese Song. The proposed approach is expected to provide a potential method for Vietnamese music search engines.
Metagenomic data has recently become crucial for precision or personalized medicine. However, these data are often complex, challenging to observe and require sophisticated visualization approaches such as clustering algorithms. Additionally, leveraging the robustness of a simple deep learning architecture, such as a shallow convolutional neural network, has attracted many scientists. Therefore, this study utilized well-known clustering algorithms such as density-based spatial clustering of applications with noise (DBSCAN), balanced iterative reducing and clustering using hierarchies (BIRCH), and ordering points to identify the clustering structure (OPTICS) to identify patterns in complex data and generate visualizations from species abundance composition of various diseases. The study then integrated a shallow convolutional neural network to perform disease prediction tasks on clustering-based visualizations. Experimental results showed that BIRCH outperformed some studies in diagnosing Type 2 diabetes, while DBSCAN performed well in diagnosing Colorectal cancer and Inflammatory Bowel Disease.
In the fight against COVID-19, accurate and timely patient diagnosis is crucial to control the disease and prevent its spread effectively. A recent study examined transfer learning from various architectures, such as Densenet, Gernet, and SeNet, and employed decoder architectures like UNet++, Deeplabv3, and Deeplabv3+ to reproduce pulmonary and COVID-19 infection regions from the features and achieve the optimal results. Remarkable results were obtained using two public datasets that included both positive and negative slices. Specifically, Densenet161 integrated with UNet++ achieved the highest scores in specificity, sensitivity, Dice coefficient, and Intersection over Union (IoU), with values of 87.6
Lungs are highly susceptible to attacks from various agents around us, and we often suffer from diseases that can be life-threatening. This study presents a diagnosis approach based on a combination of the well-known convolutional neural network architecture, VGG-16, and model interpretation techniques such as Grad-CAM and LIME. This approach helps visualize the lung areas infected with COVID-19 and other considered anomalies such as Pleural thickening and Pulmonary fibrosis. Also, it utilizes model-explanation techniques to visualize lung lesion areas. Also, we have attempted to provide explanations of prediction via all layers of VGG-16 by Grad-CAM and investigated the number of superpixels with LIME. The experimental results evaluated on Computed Tomography (CT) images collected from COVID-19 patients and healthy lungs reveal the promising combination between the image classification of VGG-16 and interpretation methods of Grad-CAM and LIME in lung disease diagnosis.
The advancement of information technology has significantly improved environmental monitoring by enabling the use of automated devices that measure environmental impact indicators and transmit data through various communication protocols. The water environment monitoring platform is specifically designed to support efficient device management and long-term data storage. The platform targets two types of users: system administrators and end users. End users, consisting of device owners and users with shared permissions to view devices' data, can make identity requests to store data, manage created devices, monitor and track data, and share permission to view any sensors' data on the device with other users. Meanwhile, system administrators verify device creation requests and renew a device's access token. Our platform leverages several technologies, including Blockchain Hyperledger Fabric to store device data, Firebase to store device information and manage users, Nodejs to create server-side APIs that communicate with the Blockchain network, and ReactJS and Reactnative to develop client-side interfaces. The effectiveness of the platform has been evaluated through a series of experiments, and the results show reasonable throughput. These technologies have enabled effective and efficient environmental monitoring and data management in modern agriculture and aquaculture.
Overcrowding in hospitals in Vietnam has caused many disadvantages in receiving and treating patients. Especially at the stage of receiving and diagnosing procedures taking patients to the treatment departments in the hospital takes up much time. This study proposes a text-based disease diagnosis using text processing techniques (such as Bag of Words, Term Frequency- Inverse Document Frequency, and Tokenizer) combined with classifiers (such as Random Forests (RF), Multi-Layer Perceptron (MLP), Embeddings and Bidirectional Long Short-term memory (LSTM)) on symptoms. As observed from the results, deep Bidirectional LSTM can reach 0.982 in AUC in the classification of 10 diseases on 230,457 samples of pre-diagnosis collected from Vietnam hospitals used in the training and testing phases. The proposed approach is expected to provide a way to automate patient flow in hospitals to improve healthcare in the future.
Cancer incidence is usually relatively low, skin diseases are not paid enough attention, and most patients, when admitted to the hospital, are in a late state which is already much damage and making it difficult to treat. Some diseases have so many similarities that it is difficult to distinguish between diseases when viewed with the naked eye. Currently, the trends of applying artificial intelligence techniques to support medical imaging diagnosis are vigorously applied and achieved many achievements with deep learning in image recognition. Deep architectures are complex and heavy, while shallow architectures also bring good performance with some appropriate configurations. This study investigates configurations of shallow convolutional neural networks for binary classification tasks to support skin disease diagnosis. Our work focuses on studying and evaluating the effectiveness of simple architectures with high accuracy on the problem of skin disease identification through images. The experiments on eight considered skin diseases have revealed that shallow architectures can perform better on small image sizes (32 × 32) rather than larger ones (128 × 128) with more than 0.75 in accuracy on all considered diseases.
Diabetic retinopathy is a highly prevalent disease with a global increase in its occurrence. It is characterized by progressive damage to the retina, the light-sensitive lining at the back of the eye. If left untreated, it can ultimately result in permanent blindness. However, accurately determining the stage of diabetic retinopathy is a complex task that necessitates the expertise of experienced medical professionals. In this study, renowned contemporary architectures such as DenseNet121 and InceptionV3 were adapted and modified to predict diabetic retinopathy stages on the dataset obtained from the Kaggle competition - APTOS 2019 Blindness Detection. An explanation technique was employed to localize regions of distinct lesions to facilitate predictions for ophthalmologists. The findings of this study demonstrate that DenseNet121 outperforms other models, achieving a validation classification accuracy of 83.2%.
Advancements in machine learning in general and in deep learning in particular have achieved great success in numerous fields. For personalized medicine approaches, frameworks derived from learning algorithms play an important role in supporting scientists to investigate and explore novel data sources such as metagenomic data to develop and examine methodologies to improve human healthcare. Some challenges when processing this data type include its very high dimensionality and the complexity of diseases. Metagenomic data that include gene families often have millions of features. This leads to a further increase of complexity in processing and requires a huge amount of time for computation. In this study, we propose a method combining feature selection using perceptron weight-based filters and synthetic image generation to leverage deep-learning advancements in order to predict various diseases based on gene family abundance data. An experiment was conducted using gene family datasets of five diseases, i.e. liver cirrhosis, obesity, inflammatory bowel diseases, type 2 diabetes, and colorectal cancer. The proposed method provides not only visualization for gene family abundance data but also achieved a promising performance level.
Metagenomic is now a novel source for supporting diagnosis and prognosis human diseases. Numerous studies have pointed to crucial roles of metagenomics in personalized medicine approaches. Recent years, machine learning has been widely deploying in a vast amount of metagenomic research. Usually, gene family data are characterized by very high dimension which can be up to millions of features. However, the number of obtained samples is rather small compared to the number of attributes. Therefore, the results in validation sets often exhibit poor performance while we can get high accuracy during training phrases. Moreover, a very large number of features on each gene family dataset consumes a considerable time in processing and learning. In this study, we propose feature selection methods using Ridge Regression on datasets including gene families, then the new obtained set of features is binned by an equal width binning approach and fetched into either a Linear Regression and a One-Dimensional Convolutional Neural Network (CNN1D) to do prediction tasks. The experiments are examined on more than 1000 samples of gene family abundance datasets related to Liver Cirrhosis, Colorectal Cancer, Inflammatory Bowel Disease, Obesity and Type 2 Diabetes. The results from the proposed method combining between feature selection algorithms and binning show significant improvements in both prediction performance and execution time compared to the state-of-the-art methods.
In recent years, Metagenomic data, or “multi-genome” data, has been increasingly used for research in “personalized medicine” approaches with the purpose of improving and enhancing effectiveness in human health care. Many studies have experimentally analyzed this data and proposed many methods to improve the accuracy of the analysis. Applying and integrating information technology to process and analyze Metagenomic data for personalized medicine approaches are necessary because of the enormous complexity of Metagenomic data. The potential advantages of Metagenomic data have been proven through many studies. Within the scope of this research, we introduce and evaluate useful tools for studying Metagenomic data in supporting the diagnosis of human disease and health conditions. From these studies, we may develop extensive and in-depth studies from previous studies to explore the important effect of the microbial ecosystem that is a rich set of microbial features for prediction and biomarker discovery in the human body. Moreover, there are trends diagnosis, appropriate treatments to improve and enhance human health.
Trong những năm gần đây, dữ liệu Metagenomic hay còn gọi là dữ liệu “hệ đa gen” được sử dụng ngày càng nhiều cho các nghiên cứu trong các tiếp cận “Y học cá thể hóa” với mục tiêu cải thiện và nâng cao tính hiệu quả trong việc chăm sóc bảo vệ sức khỏe con người. Nhiều nghiên cứu đã thực nghiệm phân tích trên bộ dữ liệu này và đề xuất nhiều phương pháp để cải thiện độ chính xác trong phân tích. Việc ứng dụng công nghệ thông tin để xử lý và hỗ trợ phân tích dữ liệu này phục vụ cho Y học cá thể là không thể thiếu bởi khối lượng công việc xử lý và độ phức tạp là rất lớn. Với những lợi ích đầy tiềm năng của dữ liệu Metagenomic đã được chứng minh qua nhiều nghiên cứu. Trong phạm vi bài báo này, nhóm nghiên cứu giới thiệu và đánh giá những công cụ rất hữu ích phục vụ cho việc nghiên cứu dữ liệu Metagenomic trong hỗ trợ chẩn đoán bệnh cho con người. Từ các nghiên cứu này, chúng ta có thể phát triển những nghiên cứu mở rộng và sâu hơn để khám phá những ảnh hưởng quan trọng của hệ sinh thái vi sinh vật trong cơ thể con người ảnh hưởng đến sức khỏe và từ đó đề xuất những xu hướng chẩn đoán và điều trị phù hợp để nâng cao và cải thiện sức khỏe con người.
Advancements in machine learning have been applying deeply and widely in numerous fields. Especially, computer vision tasks for object detection in recent years have achieved great performances which are even better than human recognition ability. This work leverages machine learning methods and information systems to present a framework for student attendance combining machine learning-based face recognition algorithm and relational databases to store, recognize, and record student attendance. The proposed method is tested on various scenarios and is expected to apply in practical cases. We investigate the Histogram of Oriented Gradients (HOG) for face detection and to use cosine distance to recognize faces. The purpose of this study is face recognition in real-time i.e. using a webcam, camera of the mobile device, and from a photograph or from a set of faces tracked in a video. We measured the distance between the landmarks and compared the test image with different known encoded image landmarks in the recognition stage. Face Recognition includes extracting features and then recognizing it, in any case, such as brightness, transformations as translation, rotation, and scale image. We recognized that using the HOG algorithm to detect faces improves more and more efficient model and avoids time-consuming.