
AI quality inspection is a key enabler of intelligent manufacturing, generating computation-intensive and latency-sensitive tasks that may overload resource-constrained inspection devices and affect production efficiency. Although Multi-Access Edge Computing (MEC) can reduce task execution delay by offloading workloads to nearby servers, dynamic wireless conditions, stochastic task arrivals, and human-in-the-loop intervention requirements make real-time scheduling highly challenging. To address this issue, this paper investigates a human-factor-state-aware MEC-enabled AI quality inspection system, where task offloading and computing resource allocation are jointly optimized with consideration of both system states and operator-related collaboration states. Specifically, the task scheduling problem is formulated as a Markov Decision Process (MDP), in which queue states, channel conditions, computing resource states, and human-factor indicators such as manual re-inspection pressure and collaboration urgency are jointly incorporated into the state space. To solve the resulting mixed discrete-continuous optimization problem, a tailored Deep Deterministic Policy Gradient (DDPG) algorithm is developed to generate real-time scheduling decisions for lightweight AI inspection tasks. In addition to minimizing long-term task latency and system energy consumption, the proposed method further suppresses collaboration delay and high-risk task backlog under dynamic industrial environments. Simulation results show that the proposed strategy converges rapidly and consistently outperforms benchmark schemes in latency and energy efficiency, while providing stronger adaptability to human-machine collaborative quality inspection scenarios.
This study addressed the limited language expression abilities of rural left-behind preschool children and the limitations of conventional anxiety assessment methods, which primarily rely on parent reports and expert observation. A One-Dimensional Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) classification model integrating Electrocardiogram (ECG) and Electrodermal Activity (EDA) signals was developed. A total of 160 rural left-behind preschool children aged 3-6 years were enrolled. At the participant level, 112 children were assigned to the training set and 48 to the test set, ensuring that all signal segments obtained from the same child were included exclusively in a single dataset subset. Raw ECG and EDA signals were acquired at a sampling rate of 1,000 Hz. Following filtering, normalization, and resampling, the processed data were converted into input tensors with dimensions of 300 × 2. Anxiety levels were determined jointly using the Preschool Anxiety Scale (PAS), the Child Behavior Checklist for Ages 1.5-5 (CBCL/1.5-5), and blinded behavioral assessments conducted by three child psychologists. Physiological signals were not involved in the labeling process. In the participant-independent test, the CNN-LSTM model achieved an accuracy of 87.3%, outperforming the standalone CNN model, which achieved an accuracy of 83.8%. Furthermore, five-fold participant-level cross-validation yielded a mean accuracy of 87.4% ± 0.9%, with a 95% confidence interval ranging from 86.34% to 88.46%. These findings indicate that jointly modeling local waveform features and temporal dependencies facilitates the discrimination of anxiety levels within the current sample. The proposed model is intended as an auxiliary screening tool, and its generalizability across different geographical regions, acquisition devices, and real-world settings requires further external validation.
The increasing accessibility of advanced image editing and generative models has strengthened the need for reliable passive image forgery detection and localization. This paper reviews representative methods for identifying manipulated images and regions without access to the original image or embedded authentication information. We introduce an evidence- and reasoning-oriented taxonomy that organizes existing approaches into five categories: acquisition consistency and noise, compression and spectral consistency, spatial correspondence and structural reasoning, multi-evidence and contextual reasoning, and generative-editing and foundation-model reasoning. The review further examines commonly used datasets, evaluation practices, and robustness settings, showing how annotation quality, manipulation realism, post-processing, and generator diversity influence reported performance. Drawing on the limitations identified across the literature, we outline the Multi-Evidence Reliability-Aware Forgery Analysis framework as a conceptual reference framework for condition-aware evidence extraction, reliability-aware fusion, uncertainty estimation, and evidence-grounded reporting. The analysis indicates that future progress depends not only on stronger architectures, but also on transferable forensic representations, realistic evaluation, calibrated uncertainty, and reliable operation under unseen manipulations and distribution conditions.
For the specific task of analyzing tourist satisfaction in film tourism scenarios, this study proposes a deep learning (DL) model that integrates multiple technologies. The model takes contextual word vectors as input, which are generated by a pre-trained language model. It combines Bidirectional Long Short-Term Memory (Bi-LSTM) to capture contextual semantic dependencies in text, while an attention mechanism is employed to dynamically focus on key emotional expressions. In addition, structured domain features, including film and television work tags and location entities, are incorporated to enhance the model’s ability to understand film and television contexts. Experiments are conducted on a film tourism review dataset. The results show that the proposed method performs excellently in the three-classification task of positive, neutral, and negative emotions, achieving an accuracy of 93.4% and an average Macro-F1 of 91.2%. This method significantly outperforms baseline models such as Support Vector Machine (SVM), Text Convolutional Neural Network (TextCNN), and Long Short-Term Memory (LSTM), and demonstrates particularly notable advantages in identifying negative reviews with a recall of 89.4%. This study thus verifies the effectiveness of multi-module collaborative modeling in sentiment analysis for vertical domains and offers a practical technical path for public opinion monitoring and service optimization in smart cultural tourism.
Accurate and energy-efficient crowd pattern recognition is important for edge-intelligent monitoring in large public spaces. Existing vision-based methods often require continuous image acquisition and high-bandwidth transmission. Device-based methods may also involve identity-related signals. These limitations make them less suitable for privacy-sensitive and resource-constrained edge scenarios. To address these problems, this paper proposes TTGSG, a target-time guided spatiotemporal graph learning framework for lightweight event-driven crowd pattern recognition. The framework uses a dual-beam infrared sensing scheme. It converts pedestrian movements into compact event-level observations, including crowd count, movement direction, traversal speed, inflow, and outflow. Environmental attributes are also introduced as auxiliary inputs. To model irregular event streams, TTGSG designs a target-time guided learning mechanism. To capture dynamic spatial interactions, it develops a C-diffusion-based bidirectional graph propagation module and a global attention-based topology updating mechanism. In addition, a future-oriented stochastic latent representation is constructed for uncertainty-aware high-density early warning. Experiments on four real-world public-space datasets show that TTGSG achieves better crowd-state recognition and early-warning performance than representative baselines. Edge deployment results further demonstrate that the proposed framework can run efficiently on resource-constrained hardware. These results indicate that TTGSG achieves a practical balance among recognition accuracy, warning timeliness, model compactness, and energy efficiency. It provides a deployable edge AI solution for privacy-preserving event-driven crowd analytics.
To address the limitations of virtual tourism, including homogeneous cultural narrative expression, fragmented audio–visual content, and insufficient interactive experience, this study proposed an artificial intelligence-driven audio–visual animation generation framework based on diffusion models and incorporated a cultural semantic control mechanism to improve generation controllability and semantic consistency. Through multimodal semantic alignment, the framework enabled cross-modal collaborative generation of text, images, and audio. Under the guidance of a cultural knowledge graph, a three-level cultural semantic representation comprising cultural symbols, narrative structure, and emotional imagery was constructed, thereby facilitating the generation of virtual tourism content with coherent narrative logic and rich cultural connotations. Experimental results demonstrated that the proposed framework consistently outperformed the comparison models across multiple evaluation dimensions. In terms of generation quality, the model achieved Fréchet Inception Distance (FID) values of 18.67–19.31 across different data subsets, substantially lower than those of the baseline methods. The Inception Score (IS) increased to 8.13–8.42, indicating superior image clarity and content diversity, while the Contrastive Language–Image Pretraining Score (CLIPScore) reached 0.40–0.44, reflecting stronger semantic alignment between textual and visual modalities. Regarding cultural narrative consistency, marked improvements were obtained in Cultural Semantic Consistency (CSC) (0.79–0.85), Narrative Coherence Score (NCS) (0.77–0.83), Emotional Alignment Index (EAI) (0.79–0.86), and Cultural Authenticity Score (CAS) (87.68–90.33). This indicated that the cultural semantic control mechanism effectively enhanced the accuracy of cultural symbol representation, narrative coherence, and emotional consistency. Subjective evaluation further supported the effectiveness of the proposed approach. Higher Mean Opinion Score (MOS) ratings (4.23–4.41), immersion scores (8.48–8.73), and user engagement values (0.86–0.89) were achieved compared with the competing methods. In addition, the audio–visual synchronization rate exceeded 92%, demonstrating superior audio–visual coordination and interactive experience. Ablation experiments further revealed that removing the cultural semantic control module resulted in substantial reductions in CSC and CAS, whereas excluding the audio–visual synchronization mechanism noticeably weakened immersion and temporal consistency, thereby confirming the effectiveness and necessity of both components. These findings indicated that the proposed framework improved the quality of virtual tourism content generation and substantially strengthened cultural expression, narrative coherence, and user experience. The integration of cultural semantic modeling with diffusion-based generative models was therefore demonstrated to be an effective solution, providing a practical technical paradigm for multimodal content generation in cultural communication.
Classification methods based on deep neural networks (DNNs) lack interpretability, making them difficult to gain full trust in critical fields such as finance, healthcare, and law, which greatly limits their applications. Most existing studies focus on the interpretability of unimodal data, while challenges remain in the interpretability of multimodal data, especially when multimodal pattern recognition models are deployed on resource-constrained edge devices. To address this problem, this paper proposes a lightweight edge-oriented multimodal explainable image classification method based on visual attributes and decision-tree reasoning. The method integrates attributes extracted from different visual modalities, such as visible-light images and depth maps, into model training, and explains the decision process through visual attributes, hierarchical decision trees, and uncertainty-aware multimodal fusion. The attribute representation is compressed by global pooling, and the decision-tree reasoning module performs category-level inference with low computational overhead, making the framework compatible with lightweight backbones and energy-efficient edge inference. Although introducing interpretability usually leads to a decrease in model accuracy, the proposed method maintains good interpretability while achieving high classification accuracy. On three datasets, NYUDv2, SUN RGB-D, and RGB-NIR, the model achieves significantly improved accuracy compared with unimodal explainable methods and performance comparable to multimodal non-explainable models. Additional edge-oriented analysis further shows that replacing the backbone with lightweight CNNs can reduce theoretical parameter and FLOP costs, indicating the potential of the proposed framework for lightweight and energy-efficient edge pattern recognition.
Retinal vessel segmentation is a fundamental component in computer-aided diagnosis of ophthalmic diseases, and its accuracy directly affects the automatic screening and quantitative analysis of various conditions such as diabetic retinopathy. Due to the slender tubular structure of blood vessels, segmentation requires two key priors: directional long-range dependency modeling and multi-scale contextual fusion. However, U-Net is limited by the restricted receptive field of 3 & times;3 convolutions and insufficient feature fusion in skip connections, making it difficult to adapt to the anisotropic morphological characteristics of vessels. Based on the encoder-decoder symmetric architecture and skip connection mechanism of U-Net, and leveraging its strengths in spatial detail recovery and multi-scale feature fusion, this study further introduces targeted improvements for directional modeling and contextual fusion of fine vessels. A lightweight and efficient improved network, termed SLKF-UNet, is proposed. In the encoder, a strip convolution module (StripBlock) is introduced to establish anisotropic long-range dependencies along horizontal and vertical directions, thereby effectively enhancing the continuity modeling of slender vessels. In the decoder, a large-kernel fusion module for skip connections (LKFBlock) is designed, which incorporates multi-scale large-kernel convolutions at skip connections to achieve efficient fusion of global and local context, thereby compensating for boundary and detail loss during upsampling. To ensure scientific rigor, the standard segmentation protocol of the DRIVE dataset is adopted, and a three-level validation framework consisting of comparative experiments, single-module ablation studies, and positional ablation studies is constructed to evaluate module effectiveness, combined gains, and optimal architecture. Extensive experiments on the DRIVE dataset demonstrate that both modules yield consistent performance improvements, and their combination outperforms the baseline U-Net and other representative improved methods across multiple metrics, including Dice coefficient, mIoU, IoU_fg, IoU_bg, and overall accuracy, thereby validating the effectiveness of the proposed model. The connect-first, then-refine design paradigm proposed in this study provides theoretical support and a novel technical pathway for the segmentation of slender tubular structures in medical images, and lays a foundation for the clinical application of retinal vessel segmentation.
The cultivation of Dendrobium officinale from Huoshan involves complex interactions among multiple environmental parameters, posing significant challenges to conventional single-parameter control strategies. To address this issue, this paper proposes a rule-based dynamic control algorithm that integrates multi-sensor data fusion with a hierarchical expert-rule mechanism, enabling the identification and prioritized handling of coupled environmental anomalies and solving coordinated multi-parameter regulation. The proposed algorithm effectively prevents conflicts between control actions and enhances system robustness under complex environmental scenarios. Experimental results demonstrate that the proposed approach outperforms static threshold-based methods and conventional PID control in terms of cooperative control effectiveness, adaptability to complex conditions, and resource efficiency, while avoiding the high computational cost and low interpretability associated with complex AI-based models.
Brain tumor segmentation (BraTS) aims to accurately identify and delineate tumor regions from brain imaging modalities, such as magnetic resonance imaging (MRI). This review focuses on two central topics in the field: (1) the technical challenges and solutions associated with missing multimodal imaging data, and (2) the development and application of 2D and 3D U-Net architectures along with their variants for BraTS. By systematically summarizing key methodological advances, this paper provides a comprehensive reference for both research and clinical practice in brain tumor analysis.
Breast calcification and abnormal tissue formation were identified as major indicators of breast cancer, where early and accurate screening played a crucial role in reducing mortality rates. In this manuscript, Enhanced Self-attention Based Hierarchical Dilated Convolutional Neural Network for Advanced Detection and Precise Diagnosis of Breast Tumors (SBHDCNN-BTDD) is proposed to improve diagnostic accuracy. Initially, input images are gathered from the BUS_UC - Breast Ultrasound Dataset and the input images undergo preprocessing using the Distributed Minimum Error Entropy Kalman Filter (DMEEKF), which effectively removes noise and enhances image quality. Tumor regions were then segmented using Accuracy-Enhanced U-Net (Acc-UNet), producing binary masks for precise region-of-interest extraction. The segmented images were subsequently classified using a Self-Attention-Based Hierarchical Dilated Convolutional Neural Network (SBHDCNN) into benign and malignant categories. To address the limitations of fixed parameter selection in conventional deep learning models, the Secretary Bird Optimization Algorithm (SBOA) was employed to optimally tune the SBHDCNN parameters. The segmentation model achieved an F1-score of 99.30%, while the final classification network attained an accuracy of 99.61% and an F1-score of 99.41%. Comparative experimental results demonstrated that the proposed framework outperformed existing approaches in terms of accuracy, robustness, and overall diagnostic performance.
Conditions affecting the retina, such as Diabetic Retinopathy (DR), Age-Related Macular Degeneration (ARMD), and glaucoma, are significant causes of irreversible vision impairment globally. The identification of these conditions quickly and precisely is invaluable. Unfortunately, noise, poor contrast, and complex anatomical features are challenges in retinal image analysis, particularly when detecting small diseases and small vessel margins. This work presents a new deep learning framework that combines a Residual Convolutional Neural Network (Residual CNN) and Long Short-Term Memory networks (DiaCNN-LSTM) to improve disease detection accuracy. During preprocessing, a Smart Adaptive Retinal Pre-processing (SARP) method is applied that integrates Contrast Limited Adaptive Histogram Equalization (CLAHE) and bilateral filtering, followed by Canny edge detection to highlight structural details while sacrificing noise and artifacts. For feature extraction, multiple descriptors of Statistical Intensity Histogram (SIH), Weber Local Descriptor (WLD) enhanced with Discrete Cosine Transform (DCT), Local Optimal Oriented Pattern (LOOP) and Histogram of Oriented Gradients (HOG) are used. Data augmentation methods, such as rotation, flipping and cropping, improve the variability of the training samples. Clustering is performed using k Means through Empirical Bregman Divergence (EBD), and classification is done by a customized DiaCNN-LSTM architecture for diabetic retinal disease detection. This architecture allows for accurate classification in difficult cases, particularly early or ambiguous cases. On the RFMiD and APTOS2019 datasets, experimental results show that the proposed model outperforms state-of-the-art methods with an accuracy of 98.86% and Cohen’s Kappa value of 98.24 (%). The AI model that is being presented is scalable and interpretable, making it suitable for clinical use, especially for real-time monitoring by teleophthalmology and low-resource healthcare facilities.
The human eye visually perceives surrounding objects, and the retina plays a crucial role in capturing light and converting it into electrical signals for further processing by the brain. The eye consists of multiple layers and is divided into anterior and posterior segments. Numerous automated and manual procedures have been developed to identify retinal disorders, but these methods are often uncomfortable for patients and time-consuming. To address these drawbacks, this paper proposes a new technique named Enhanced Central Serous Retinopathy Classification Using an Optimized Digital Twin Enabled Domain Adversarial Graph Network for Advanced Analysis of Retinal Images (CSRC-ODTAGN-RI). Initially, input data is collected from Optical Coherence Tomography (OCT) image and Fundus image dataset. Then the collected data is given to the Adaptive Square Root Cubature Kalman Filter (ASRCKF) to reduce the noise and enhance the quality of the image. Then, the preprocessed images are given to the Digital Twin Enabled Domain Adversarial Graph Network (DTAGN) for Central Serous Retinopathy (CSR) classification. The Sea Horse Optimization (SHO) algorithm, inspired by seahorse behavior, is an evolutionary optimization technique used to fine-tune parameters within the DTAGN, enhancing CSR classification. The proposed method attains 26.64%, 15.27%, and 30.45% higher accuracy and 26.72%, 10.08%, and 30.64% higher F1-score compared to the existing models: Detect the CSR utilizing DL by Retinal Imageries (DCSR-OCT-CNN), Intra and Inter Expert Validation of an Automatic Segmentation Technique for Fluid Regions Associated with Central Serous Chorioretinopathy in OCT Imageries (IIEV-FRCSC-OCTI), and MacularNet: Towards Fully Automated Attention-Dependent Deep CNN for Macular Disease Categorization (MNFA-DCNN-MDC), respectively.
Hidden emotions are often expressed through subtle facial movements known as micro-expressions (ME). However, accurate Micro-Expression Recognition (MER) remains challenging because of the low intensity, short duration, and limited availability of ME datasets. To address these challenges, this manuscript proposes an advanced MER framework called MER-HEFM-GGNN for accurate hidden emotion classification. Initially, input video samples are collected from the CASME II dataset and pre-processed utilizing Cauchy Robust Correction-Sage Husa Extended Kalman Filtering (CRCSHEKF) to enhance data quality through cropping, alignment, motion amplification, and normalization. Subsequently, Integrated Fast Hough Transform (IFHT) and Image Difference Sequence Features (IDSF) are employed to extract discriminative motion features. These features are then classified using a Gegenbauer Graph Neural Network (GGNN). The GGNN model further is optimized using Harbor Seal Whiskers Optimization (HSWO), which enables adaptive parameter tuning and enhances the classification accuracy of subtle micro-expression patterns. The proposed framework achieves superior performance with an accuracy of 94.62%, Unweighted Average Recall (UAR) of 92.83%, and Unweighted F1-score (UF1) of 91.97%, while also reducing computational time compared to existing methods. This framework can be effectively applied in areas like lie detection, clinical diagnosis, and security systems, contributing to advancements in emotion recognition technologies.
Photonic radar systems improve environment sensing and multi-target detections in self-driven vehicles but are hampered in their adoption due to high system cost, complexity of integration, and environment dependencies. To overcome these limitations, an Advanced Hybrid Mode-Division Multiplexing-Polarization-Division Multiplexing Photonic Radar is proposed in this research for target detection in 5G-enabled self-driven vehicles. The primary objective is to develop a hybrid multiplexing photonic radar system that integrates mode-division multiplexing (MDM) and polarization-division multiplexing (PDM) to enable precise multi-target detection under adverse operating conditions, thereby improving the reliability, safety, and sustainability of autonomous vehicle navigation systems. The MDM-PDM configuration supports four-channel transmissions in a compact form factor using X- and Y-polarized streams with phase-shifted donut modes. Linear Frequency Modulation (LFM) chirp signals improve Doppler tolerance, while beat signal extraction and signal-to-noise ratio (SNR) estimation allow precise range and velocity measurements. Radar data are processed using a Logarithmic Differential Convolutional Neural Network (LDCNN) implemented in MATLAB, enhancing target discrimination and detection efficiency. The proposed system demonstrates achieves an accuracy of 96.4% and probability of detection (Pd) of 0.984, while maintaining a low false alarm rate (Pfa) of 0.016. Additionally, a low SNR error of 0.015% is observed, serving as an auxiliary indicator of signal fidelity under simulated conditions. Simulation results demonstrate superior multi-target localization, resilience to noise, and reliable operation, supporting safe and intelligent autonomous navigation. This work introduces a novel hybrid MDM-PDM photonic radar architecture integrated with deep learning-based signal processing, providing a compact, high-precision solution for real-time target detection in complex driving environments.
A deep learning model of segmentation and analysis of facial components combining efficient preprocessing, hybrid segmentation, multi-feature extraction, and cooperative feature selection. The preprocessing step uses Contrast-Limited Adaptive Histogram Equalization (CLAHE), Tensorial Adaptive Anisotropic Diffusion (TAAD) and Active Appearance Modeling (AAM) to improve the image clarity and structural integrity. The segmentation is done with a SqueezeNetAttention U-Net hybrid architecture that combines the performance lightweight of SqueezeNet with space attention of Attention U-Net to ensure that the component is properly delineated. Multi-source feature extraction combines Local Phase Congruency (LPC), Histogram of Oriented Gradients (HOG), Color Histograms, Gradient Orientation Tensor (GOT) and ResNet embeddings to extract a variety of facial signals. A hybrid feature-selection method, termed Hybrid BFO-SBO Feature Selection (HBFS), is proposed, which consecutively employs Bacterial Foraging Optimization (BFO) for global exploration and Secretary-Bird Optimization (SBO) for local exploitation. Finally, facial component detection is performed using the proposed DEEPFACS framework, in which a ResAttNet-LSTM network jointly integrates a ResNet backbone, self-attention blocks, and LSTM units to effectively learn spatial relationships and achieve accurate classification. The experimental assessments on CelebAMask-HQ and Helen datasets show that the segmentation accuracy (98.7%) and robustness are better than the baseline feature-selection methods. The framework has a low inference latency and used in real-time facial component analysis and human computer interaction.
Objectives: This study sought to determine the radiological factors associated with early neurologic deterioration (END) and long-term functional outcomes (LFO) in acute ischemic stroke (AIS) patients following thrombolysis treatment. Methods: We retrospectively included patients with symptomatic anterior circulation large vessel stenosis or occlusion stroke. The National Institutes of Health Stroke Scale (NIHSS) was used to measure the severity of the symptoms both at admission and 72 hours later. We used a binary logistic regression model to identify the functional independence image factors. Results: There was a substantial correlation between the occurrence of END and unfavorable scores on the mRS at 90 days. The significant CT parameter predictors for END and an unfavorable 90-day mRS score were lower regional leptomeningeal collateral (rLMC) score, larger volume of perfusion lesion, infarct core, and ischemic penumbra. As shown by area under the receiver operator characteristic curves (AUCs) for predicting the long-term functional outcomes, the optimal thresholds of the rLMC score, perfusion lesion, infarct core, and ischemic penumbra were 17.5, 90 mL, 14.5 mL, and 57.0 mL, respectively. Conclusions: In patients with symptomatic AIS treated with thrombolysis, CT imaging parameters, including rLMC, perfusion lesions, infarct core, and ischemic penumbra have independent predictive significance for END and LFO. Advances in Knowledge: Our findings may help to improve the understanding of the prognostic value of these factors, identify potential therapeutic targets, and guide clinical decision-making for AIS patients undergoing thrombolysis.
Proper hand hygiene is critical for infection prevention, yet real-world compliance and quality assessment remain challenging in busy clinical environments. This paper presents a lightweight end-to-end hand hygiene monitoring system that integrates egocentric hand pose estimation, temporal step recognition, and soap-foam verification for real-time embedded deployment. Hand keypoints are obtained using either MediaPipe or our custom GhostHandEgoNet, which leverages Ghost convolutions with attention to improve robustness under occlusion while maintaining a compact footprint; it achieves an end-point error below 14 pixels with only 2.56 M parameters and improves keypoint accuracy by over 90% compared with MediaPipe in occluded scenes. A three-layer Long Short-Term Memory (LSTM) model continuously skeletonizes sequences to classify the standardized seven-step handwashing protocol, achieving accuracies of 93.5%, 97.1%, and 94.3% on the Kaggle, METC, and in-house hospital datasets, respectively. To assess washing quality beyond motion analysis, a modified Fast-SCNN module performs foam segmentation and sufficiency classification in a single forward pass, reaching a foam mIoU above 85%. Compared with the Transformer and ST-GCN-based backbones, our method achieves the lowest average latency of 87.66 ms/frame and the highest throughput of 11.41 FPS on the NVIDIA Jetson AGX Xavier platform. These results demonstrate that our approach not only delivers competitive classification accuracy but also satisfies the stringent real-time and memory constraints required for practical embedded deployment, making it particularly suitable for on-device hand hygiene monitoring in clinical and public health settings.
Feature selection is a fundamental process in machine learning and data analysis that significantly affects the performance of machine learning models, especially in high-dimensional datasets that often contain irrelevant or redundant features. These additional features can lead to reduced model accuracy and long computational time. An innovative and effective approach for feature selection in high-dimensional data is used in graph-based methods. This study comprehensively surveys different graph-based feature selection techniques for high-dimensional data to reduce complexity and improve accuracy. We have explored a wide range of graph-based feature selection methods and organized the existing works into traditional and modern methods. Also, we have compared existing methods in terms of prediction accuracy, computational complexity, scalability, flexibility, and speed. We also examined graph-based feature selection method applications and discussed the challenges and limitations associated with high-dimensional datasets. Finally, we specify unresolved challenges and potential future research paths to facilitate progress in this fast-changing domain.
Detecting lung cancer at an early stage can reduce mortality, which is one of the most dangerous diseases. However, manual interpretation by health professionals can produce variable results. In this paper, an Advanced Federated Learning Technique for Multi-Institutional Lung Disease Detection using Chest Radiographs and CT Scans (FL-MILDD-CR-CTS) is proposed. The input images are collected from chest X-ray and CT scan datasets and preprocessed using an Adaptive Tracking Dual Nested Kalman Filter (ATDNKF) for noise removal and enhancement. Dimensionality reduction is performed using Principal Component Analysis (PCA) to maintain essential information while reducing complexity. Feature extraction and disease classification are conducted using SVM, CNN, and Random Forest, optimized for identifying complex lung patterns. The framework aggregates locally trained models using Federated Learning (FedAvg and FedProx) to preserve data privacy. The experimental results show that FL-MILDD-CR-CTS achieves 26.68% better accuracy and 30.94% better precision when compared with the existing models.