
The power system is a core infrastructure for the national economy and energy security. As key assets in the grid, transmission-line components must be accurately inspected from UAV images despite small targets, large scale variation, complex backgrounds, and partial image degradation. This paper proposes KED-GNN, a graph neural network enhanced by space-to-depth convolution, Kolmogorov–Arnold Network-based nonlinear transformation, and efficient multi-scale attention, for UAV-based transmission-line component and fault detection. Graph-based image representation methods, such as Vision GNN, provide a way to represent an image as a graph by treating image patches as nodes and their relationships as edges, which has potential for UAV transmission-line inspection. However, direct use of this representation still faces small-target information loss, multi-scale feature variation, and complex background interference. Therefore, KED-GNN introduces an SPD-Conv-based detail-preserving sampling algorithm, a KAN-based nonlinear feature transformation mechanism, and an EMA-based efficient multi-scale attention mechanism. Under the same Faster R-CNN detection framework, KED-GNN achieves a final mAP @ 0.5 of 0.86 and precision of 0.85, improving mAP @ 0.5 by 3.61
Channel estimation errors significantly degrade signal detection performance in massive multiple-input multiple-output (massive MIMO) systems, limiting the effectiveness of conventional linear, near-optimal, and learning-based detectors under practical channel uncertainty. To address this challenge, this paper proposes a channel state information (CSI)-aware deep neural network-assisted detector (DNN-ML) that explicitly incorporates channel estimation error variance into the detection process. The proposed framework employs CSI feature extraction, dimensionality reduction, error-adaptive feature fusion, and reliability-aware decision refinement to jointly exploit the received signal, estimated CSI, and CSI uncertainty. The detector is evaluated for a 256×256 massive MIMO system employing 256-QAM modulation over Rayleigh fading channels with 10
Remote healthcare monitoring systems face significant challenges in managing the substantial data volumes generated by continuous physiological signal acquisition on battery-powered wearable Internet of Things (IoT) devices. This paper introduces a novel Tchebichef–Compressed Sensing–Primal–Dual (TCP-PD) framework for electrooculogram (EOG) signal compression in IoT-based healthcare applications. Unlike conventional compressed sensing approaches that rely on Fourier- or wavelet-based sparsifying transforms, the proposed methodology employs discrete Tchebichef polynomials, exploiting their inherently discrete formulation and superior energy compaction properties on EOG signals. The TCP-PD framework integrates Gaussian random measurement matrices that satisfy the Restricted Isometry Property with an optimized primal–dual reconstruction algorithm enhanced by a backtracking line search. A comprehensive evaluation on the Sleep-EDF and DROZY databases shows that the proposed method achieves a 70
Plastic pollution has become a persistent problem in natural ecosystems. Driven by insufficient waste management, limited recycling efficiency, and high costs for recycling, plastic waste in the environment poses significant ecological and health risks. Recycling requires accurate identification of polymer types, but existing optical sorting systems often struggle to distinguish common household plastics. In this work, we present a novel classification approach based on a multispectral imaging system consisting of nine cameras equipped with near-infrared bandpass filters. The system is designed to discriminate the seven most common household plastics. From the resulting multispectral images, we extract the spectral fingerprints and derive features such as intensity differences between specific wavelength pairs and their slopes, as well as false-color image representations. A dedicated preprocessing pipeline aligns and normalizes the data before classification. We recorded a multispectral household plastic database ( https://github.com/FAU-LMS/MHPM ) and trained four different classifiers Gradient Boosting, Extreme Gradient Boosting, Light Gradient Boosting Machine, and CatBoost. The best-performing model achieves a classification accuracy of 86.7 s per pixel, enabling efficient processing of high-resolution images. The entire setup is built from off-the-shelf hardware components, which makes replication straightforward and allows direct integration into industrial sorting pipelines.
To address the limitations of insufficient user segmentation accuracy and strategy rigidity in traditional electricity marketing, this study proposes an AI-driven framework integrating deep clustering and adaptive decision-making. A hybrid model combining deep embedded clustering and reinforcement learning (RL) has been developed to automatically extract high-dimensional behavioral patterns and generate dynamic marketing strategies. The framework introduces a feature attention mechanism for adaptive feature selection and employs a multi-agent RL system to optimize personalized strategies based on real-time user feedback. Experimental results on simulated datasets demonstrate a 23.5 % improvement in clustering purity and a 31.8 % increase in strategy response rates compared to conventional methods, validating the effectiveness of AI-enhanced components.
Sleep Apnea Syndrome (SAS) is a common disorder characterized by repeated cessation of airflow during sleep. Accurate classification of its main subtypes—Obstructive Sleep Apnea (OSA), Central Sleep Apnea (CSA), and Normal Breathing (NB)—is essential for effective clinical management. This study proposes a classification framework based on electroencephalography (EEG) signals and ensemble learning techniques. EEG data from C3-A2 and C4-A1 channels in 25 subjects were extracted and then segmented into 10-s windows. From each segment, four features such as Sample Entropy, Higuchi Fractal Dimension, Variance, and Standard Deviation were extracted across five frequency bands. Principal Component Analysis (PCA) was applied for dimensionality reduction prior to classification. Classification was performed using ensemble models with decision trees as base learners. Using subject-wise cross-validation, Boosting showed the best performance among the methods tested, with classification accuracies of 92.33
Vibration signals in rotating machinery are frequently contaminated by noise, which compromises data quality and leads to the loss of valuable signal information. Consequently, effective noise reduction is essential for accurate signal analysis. This study presents a comparative analysis of various noise reduction methods applied to vibration signals in rotating machinery. The methods evaluated include compressed sensing with dictionary bases and multiple reconstruction algorithms, discrete and wavelet packet denoising using different mother wavelets and thresholding techniques, filtering methods, hybrid approaches combining wavelet denoising with compressed sensing, and mode decomposition based methods including empirical mode decomposition (EMD), ensemble empirical mode decomposition (EEMD), and complete ensemble empirical mode decomposition (CEEMD) integrated with wavelet denoising. Detrended fluctuation analysis (DFA) and filtering techniques are also investigated. A key challenge in compressed sensing is the uncertainty associated with signal sparsity estimation, particularly in noisy or near sparse signals, which can affect reconstruction accuracy. Performance metrics including mean squared error (MSE), mean absolute error (MAE), signal to noise ratio (SNR), peak signal to noise ratio (PSNR), cross correlation, and computational speed are employed to evaluate denoising performance and optimal sparsity levels. The results indicate that compressed sensing based approaches enable effective noise reduction while significantly reducing data size, preserving signal reconstruction quality under high compression conditions. The results further demonstrate that hybrid wavelet denoising and compressed sensing techniques provide superior denoising performance compared to individual methods, and that the choice of an appropriate dictionary matrix significantly enhances both sparsification and reconstruction accuracy.
As one of the most valuable crops in the world, tomatoes are essential to the economies of many countries. However, these harvests are still vulnerable to multiple diseases that can diminish and even eradicate the production of healthy crops; thus, it is imperative to accurately and promptly identify these diseases. Consequently, a rapid and precise approach for the classification of plant diseases is presented by the deep learning (DL) method. Therefore, in this study, we have presented an automatic tomato leaf disease classification using convolutional neural network with squeeze-and-excitation blocks (CN2-SE) with the improved deep residual shrinkage network method. The suggested method consists of preprocessing, segmentation, feature extraction, and classification stages. Data are first preprocessed using a median filter; followed by preprocessed images are segmented using the enhanced fuzzy C-means clustering (EFCM) method. Next, the proposed convolutional neural network with squeeze-and-excitation blocks (CN2-SE) is used to extract significant features. The extracted features are inputted into a proposed IDRSN classifier to classify a tomato leaf as diseased or non-diseased. To enhance the classifier effectiveness, the proposed adaptive golden eagle optimization (AGEO) algorithm is used to optimize the network’s initial weights and biases. According to the results, the suggested method outperforms existing methods in terms of classification accuracy.
Speaking one’s language is the most common and efficient way to convey ideas and information. People can express themselves verbally by putting their ideas, emotions, and experiences into words. However, with the rapid advancements in technology, criminals often use voice changers as anti-forensic tools to conceal their voices while committing illegal activities. In this proposed work, normal voice (NV) is disguised using two methods: objects in the mouth (OM) and whispering (WP), which are considered for experimental purposes. Features are extracted using Mel Frequency Cepstral Coefficients (MFCC) and their derivatives, as well as Tonal Frequency Cepstral Coefficients (TFCC) and their derivatives, for both normal and disguised voices. After feature extraction, several machine learning classifiers are employed to differentiate between disguised and normal voices. Based on above abstract write a suitable paper title. The classifiers used include Support Vector Machine (SVM), Decision Tree (DT), Linear Discriminant Analysis (LDA), Naïve Bayes (NB), and Logistic Regression (LR). For speaker identification from WP-disguised voices, the proposed method achieved accuracies of 82.50
In recent years, the application of Unmanned Aerial Vehicle (UAV) in production activities involving outdoor operations has gradually become widespread and routine. However, challenges such as the wide size variation of small objects, frequent occlusions, and complex backgrounds have emerged as key difficulties in UAV aerial image object detection. This paper proposes PASR-YOLO, an enhanced YOLO11 algorithm optimized for small objects. To address the issue of insufficient representation of small objects in the detection layer, the High Resolution Detail Injection Unit (HDI) is introduced. Building upon the addition of a high resolution detection branch, it innovatively extracts small object features from low level layers and injects them into higher layers. Furthermore, incorporating the Cross Stage Partial network (CSP) based on three different size branch representations, it aggregates multi-scale and fine-grained information, achieving a lightweight and high precision small object feature pyramid. To further reduce the number of parameters, a low-redundancy mapping named n-Scale Recursive Split-Enhance Fusion Block (NSSE) is proposed that accommodates both large and small models, balancing compactness with equivalent feature extraction and gradient flow capabilities. On the VisDrone dataset, PASR-YOLO achieves 6.6
Over the past few years there has been growing interest in properly measuring audio loudness to provide consistent sound to a user no matter what platform they are accessing through including music streaming, broadcast, and hearing devices. Existing traditional approaches based on signal-based processing, although powerful in a controlled set-up, lack the capacity to generalize to other audio genres and to real-world practice. The study suggests a data-based framework of audio loudness estimation based on a combination of classical machine learning models and advanced deep learning networks, such as TD-transformer models (transformer-based feature-token regression model) acoustic relationships and interactions between the acoustic descriptors based on multi-head self-attention, facilitating learning of higher-order, nonlinear interactions between the features and achieving robust integrated loudness prediction. Employing the GTZAN music genre dataset containing tabular features (MFCCs, RMS energy and spectral centroid) extracted in tabular format we have trained and tested a few regression models. The obtained results demonstrate that the proposed TD-Transformer model outperformed by recording the lowest value of the MAE 0.012, RMSE 0.017, and the highest value of the R^2 score 0.920, proving to be a high-capacity model in accurately predicting audio loudness and also in comparison to the best baseline model MLP which recorded an MAE of 0.0150, RMSE of 0.019, and R^2 score of 0.89. To implement the model interpretability and facilitate the transparency of the decision-making, we applied an explainable AI approach of local interpretable model-agnostic explanations and Shapley additive explanations to visualize and determine how the single features in the model input affected the model prediction. The proposed framework shows best performance in loudness estimation on the GTZAN dataset and suggests the potential of transformer-based feature interaction modeling in structured audio descriptors.
As a specialized case of compressed sensing, 1-bit compressed sensing retains only the sign information of linear measurements, thereby achieving extremely low-sampling cost and high sampling speed. It also exhibits stronger robustness against nonlinear distortions and additive noise. The Binary Iterative Hard Thresholding (BIHT) algorithm has been proven efficient for solving 1-bit compressed sensing problems. However, the discontinuity of its hard-thresholding function makes its performance highly sensitive to inaccuracies in the estimated signal sparsity level. An inaccurate sparsity estimate leads to a fixed threshold that can cause significant reconstruction bias. To address this issue, this paper proposes a novel Arctangent-Regularized Binary Iterative Thresholding (ARBIT) algorithm. By introducing an arctangent penalty term, the algorithm constructs a continuously adjustable thresholding function that preserves sparsity while effectively mitigating thresholding-induced bias. The algorithm further incorporates an adaptive parameter updating mechanism and a Nesterov momentum strategy to enhance convergence speed and stability. Experimental results demonstrate that the ARBIT algorithm significantly outperforms traditional iterative thresholding methods on both synthetic and real seismic datasets in terms of support set recovery rate, noise robustness, and reconstruction efficiency. This work provides a viable new solution for efficient compression and high-quality reconstruction of seismic data.
Although movable antenna (MA) systems offer significant performance gains via spatial reconfiguration, they suffer from severe degradation under practical imperfect channel state information (CSI). To address this, this paper proposes a robust joint design of antenna position optimization and precoding. We model stochastic errors in angular parameters and path gains and develop a Robust Monte Carlo Orthogonal Matching Pursuit (RMC-OMP) algorithm. The algorithm effectively suppresses pseudo-correlation peak interference caused by CSI errors through statistically averaging matching scores across multiple channel error realizations, while strictly ensuring the half-wavelength spacing constraint. Simulation results demonstrate that the RMC-OMP retains 85.6% of the sum rate achievable with perfect CSI. This validates the proposed algorithm’s strong robustness for 6G MA systems.
To address the issues of uncontrollable jamming regions and susceptibility to partial cancellation in existing suppression jamming methods against single-point-source synthetic aperture radar ground moving target indication (SAR-GMTI), this paper proposes a two-dimensional controllable suppression jamming generation method using multi-point-source-cooperative composite-modulation for SAR-GMTI. This method employs motion modulation and noise-product modulation to control the azimuth-direction jamming position and range, and frequency-shift modulation and noise-convolution modulation to control the range-direction jamming position and range. Additionally, distributed multi-jammer cooperation effectively mitigates the problem of partial cancellation. Theoretically, the paper derives an analytical expression for the SAR-GMTI jamming imaging results and conducts a detailed analysis of jamming performance. Both theoretical analysis and experimental results demonstrate that the proposed jamming method can generate controllable suppression jamming patches while preventing cancellation by SAR-GMTI processing, thereby achieving moving target protection. Moreover, due to the precise controllability of the generated jamming regions, the method achieves higher jamming efficiency and exhibits good engineering practicality.
Quality assurance in industrial manufacturing relies heavily on PCB defect detection, yet it faces challenges such as complex backgrounds, minute target sizes, and significant scale variations. To address these challenges, we propose an enhanced adaptive multi-scale object detection model built upon RT-DETR, termed EAM-DETR. Firstly, an improved Faster-ELA Block is introduced, which reduces the model’s computational overhead while simultaneously enhancing detection efficiency. Secondly, we introduce the AIFI-ASSA module, which integrates adaptive sparse self-attention (ASSA) to prioritize salient features within clutter-filled backgrounds. Finally, we develop a multi-scale collaborative feature pyramid (MCFP) that preserves fine-grained features and substantially improves detection of small targets across various scales. On the PKU-Market-PCB dataset, EAM-DETR reaches 97.1
Turbo product codes (TPCs) are a class of forward error correction codes that achieve high coding gain at high code rates through iterative decoding. To meet the flexible and scalable requirements of new communication systems, in this paper, a high-speed software TPC decoder on graphics processing unit (GPU) is proposed. The presented decoder exploits both inter-frame and intra-frame parallelism. Moreover, optimized data structure and on-chip memory allocations are conducted to improve memory access bandwidth. In addition, the asynchronous data transfer approach is employed to hide data transfer latency. The proposed (64,57)^2 TPC decoder on RTX3070 achieves 231.7 Mbps throughput with 5 iterations. Experimental results show that the proposed decoder obtains 3.14× throughput speedups compared with the existing method, at the cost of degraded latency performance.
Multi-person pose estimation in crowded scenes remains difficult because heavy overlap and occlusion frequently break independently predicted keypoint heatmaps. Contextual Instance Decoupling (CID) improves crowded-scene estimation by generating instance-aware feature maps, yet its final joint heatmaps are still produced channel by channel without explicit structural coupling. This paper presents SimpleCID, a lightweight refinement built on top of CID. After the Global Feature Decoupling stage, we model the keypoints of each person as nodes in a human body graph and propagate responses through a fixed normalized adjacency matrix. The refined response is fused with the original heatmap by a residual connection with a small coefficient, allowing adjacent joints to provide structural support while preserving the baseline prediction. The module introduces no additional trainable parameters, keeps the original training pipeline unchanged, and adds only lightweight matrix multiplication along the joint dimension. On crowded-scene benchmarks, SimpleCID consistently improves the baseline: it raises AP by 1.2 points on CrowdPose and improves OCHuman AP from 41.4 to 43.3. Qualitative comparisons further show more complete limb recovery and fewer anatomically inconsistent predictions under severe occlusion. These results demonstrate that explicit yet simple skeleton reasoning is an effective complement to contextual instance decoupling.
Synchrosqueezed time–frequency analysis is widely used to improve energy concentration and instantaneous frequency estimation in nonstationary signals. Despite its effectiveness, standard synchrosqueezing relies on non-causal kernels and reassignment rules that may introduce time–frequency energy prior to the physical onset of a signal, potentially compromising temporal interpretability in applications where causality is essential. In this work, I propose a causality-preserving formulation of synchrosqueezed time–frequency analysis in which temporal admissibility is enforced directly within the reassignment and thresholding framework. Causality is treated as a structural property of the time–frequency representation rather than as a post-processing constraint. To quantify the effects of causality enforcement, I introduce diagnostic metrics that explicitly measure causality violations, ridge stability, and energy localization. Using analytically defined synthetic signals with known ground truth, I perform a systematic comparison between standard and causality-preserving synchrosqueezed representations. The results demonstrate that enforcing causality eliminates non-physical pre-onset energy by construction and yields improved ridge stability and reduced off-ridge energy leakage, while maintaining comparable instantaneous frequency accuracy across a broad range of noise levels. A sensitivity analysis further shows that the resulting trade-offs can be controlled through thresholding and smoothing parameters without requiring fine tuning. The proposed framework provides a transparent and physically interpretable extension of synchrosqueezed time–frequency analysis, particularly suited for applications where temporal causality and onset consistency are critical.
Simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) and nonorthogonal multiple access (NOMA) are two highly promising candidate technologies for sixth-generation (6G) wireless communication systems and have garnered considerable attention from academia and industry over the past several years. In this work, a novel double-STAR-RIS-assisted uplink NOMA system is put forward, with the aim of minimizing transmit power under the constraints of phase shift matrices and minimum quality of service (QoS). To settle this optimization problem, an alternating optimization (AO) algorithm is put forward. First, the transmit power equation is obtained by deriving the formula pertaining to the QoS. Next, the transmit power minimization problem is split into several subproblems, and the optimal power allocation and passive beamforming are obtained via iteration. Last but not least, the simulation results indicate that the proposed algorithm is feasible, effective, and converges quickly, revealing that compared to conventional orthogonal multiple access (OMA) methods, the proposed scheme can decrease the total transmit power.
Alzheimer’s disease (AD) is a neurodegenerative disorder that damages brain cells if not recognized and treated early. AD is a chronic disease, but proper treatment in the early stage can help to manage AD. Deep learning models can examine magnetic resonance imaging (MRI) scans to predict significant modifications in the brain associated with AD. Thus, this work develops an automated model for AD detection using deep learning. Initially, MRI images are collected from standard datasets, such as the OASIS-based Kaggle repository and the augmented Alzheimer MRI dataset. These images are then given to the Adaptive Squeeze-and-Excitation Attention Module-based Conv-Transformer Network (ASEA-ConvTNet) for the AD detection process. The developed ASEA-ConvTNet incorporated the convolutional neural network (CNN) element with a transformer network to achieve global contextual data. Moreover, the Squeeze-and-Excitation (SE) block is used to develop a feature response in a channel manner to learn the interdependency between feature details. In this, the ‘Adaptive’ specifies the utilization of the renovated constant parameter of flower fertilization optimization (RCP-FFO) to optimize the model parameters during training. The RCP-FFO-based optimization process is influenced by the network to get the fine-tuned parameter value that helps to reduce the parameter usage and improve the performance. Thus, from the developed model, the AD-detected outcome is achieved. Then, performance and comparative analyses are performed using various measures and different algorithms. The developed system achieves accuracy rates of 94.9