
Infrared (IR) object detection poses unique challenges due to low texture, weak edges, sensor noise, and a strong dependence on global context. Modern convolutional object detectors, including YOLOv8, are primarily optimized for RGB imagery and often underperform when directly applied to large-scale infrared datasets. In this paper, we present a simple yet effective architectural enhancement to YOLOv8 by (1) replacing the standard Spatial Pyramid Pooling Fast (SPPF) module with a learnable multi-scale variant (SPPFPlus), and (2) integrating the Convolutional Block Attention Module (CBAM) into both the backbone and neck. The proposed approach introduces explicit parallel multi-scale context modeling and adaptive channel–spatial attention, addressing key representational limitations of infrared imagery. Extensive experiments on a large infrared dataset demonstrate an improvement of approximately 15% in detection performance over the YOLOv8 baseline, with notable gains in recall and small-object detection. The modifications are lightweight, modular, and can be seamlessly integrated into existing YOLOv8
Industrial Internet of Things (IIoT) networks supporting mission-critical control and automation require ultra-reliable, low-latency communications (URLLC). Non-stationary interference, including narrowband tones, impulsive bursts, and chirp-like signals, significantly degrades Quality of Service (QoS) by increasing Block Error Rate (BLER), Hybrid Automatic Repeat Request (HARQ) retransmissions, and worst-case latency. Conventional mitigation techniques based on Fast Fourier Transform (FFT) and Short-Time Fourier Transform (STFT) exhibit limited adaptability and insufficient time-frequency resolution for transient interference. This paper presents a system-level integration and cross-layer evaluation of wavelet-based interference mitigation, integrating the Continuous Wavelet Transform (CWT) for multiscale detection, the Stationary Wavelet Transform (SWT) for shift-invariant suppression, and the Median Absolute Deviation (MAD) for adaptive thresholding. Evaluation is conducted using real 5G New Radio (NR) traces from a Firecell testbed, augmented with heterogeneous synthetic interference. Experimental results demonstrate a 3 dB gain in Effective Signal-to-Noise Ratio (ESNR), a 45% reduction in 99th-percentile latency, a 62% reduction in BLER, and a 39% increase in throughput relative to FFTnotch and STFT-based baselines. These results confirm the practical value of wavelet-domain cross-layer interference mitigation for 6G-oriented IIoT deployments.
Biomedical image segmentation plays a pivotal role in diagnostic radiology and computational pathology, enabling precise delineation of anatomical and pathological structures. However, despite advancements in deep learning-based segmentation, challenges persist in interpretability, computational tractability, and scalability. This paper proposes an advanced computational framework that integrates Large Language Models (LLMs) with segmentation architectures, quantum databases for accelerated query performance, and optimized image compression techniques. The proposed system leverages mathematical principles of variational optimization, tensor decomposition, and quantum search complexity to enhance segmentation efficiency, reduce latency, and improve decision support. A rigorous comparative analysis is performed using benchmark datasets, demonstrating superior segmentation accuracy, reduced query response time, and improved data storage efficiency. The integration of LLMs provides an interpretable interface for clinicians and radiologists, enhancing the usability of automated segmentation in real-world medical workflows.
Today, drone-based attacks represent serious threats to the security and safety of public infrastructures. For successfully detecting a malicious drone in a given zone, there are three phases: signal collection (sensing), features extraction and classifications. Signal collection can be performed using available sensing technologies such as radar, acoustics sensors and electro-optic technologies, among others. The classification phase is often achieved using general-purpose algorithms such as Naive Bayes and support vector machine (SVM). On the other hand, the features extraction phase is very problem-specific, and its performance depends on several factors such as the used sensory technology, environment, and the drone characteristics. Features engineering is a designing stage that aims at identifying the most distinctive information carriers which capture the drone's discriminative characteristics. In this paper, we present effective drones' features extraction techniques for the most popular sensory technologies available which are radar, RF analyzers, acoustic sensors, and electro-optic sensors. We focus on identifying the most distinctive features of drones and show how to extract them out of the collected signals.
The rapid expansion of 5G use cases and dense IoT deployments has intensified pressure on scarce spectrum, making reliable wideband sensing indispensable for dynamic spectrum access (DSA). This work targets energy detector-based cooperative wideband spectrum sensing (CWSS) over multiple sub-bands in complex AWGN channels, where noise uncertainty and very low SNR commonly affect performance. We propose a CWSS architecture that applies maximal-ratio combining (MRC) at the pre-detection stage and employs a k-out-of-N fusion rule at the decision center. MRC boosts the effective per-sub-band SNR before thresholding, while the fusion mechanism aggregates local binary decisions to improve global reliability with modest reporting overhead. Comparative evaluations indicate that the MRC-aided ED with k-out-of-N consistently achieves higher detection probability for a given false-alarm rate across challenging low-SNR conditions, outperforming non-cooperative sensing and conventional CSS baseline. The results demonstrate that combining MRC with k-out-of-N fusion mitigates noise-uncertainty effects, strengthens wideband hole detection, and provides a practical sensing frontend for policy-compliant DSA in 5G environments
Existing monitoring systems in developing countries are deficient in detecting diseases and other factors that adversely affect agricultural productivity. This paper presents a low-cost method of disseminating information to farmers by transmitting digital data over an FM channel. The low-cost method is employed in a poultry pen where temperature variations are broadcasted using the radio-text field of the FM-RDS protocol. The system uses a FLIR Lepton 3.5 sensor embedded in a tCam-Mini Rev4 thermal camera to capture temperature values. The FM-RDS transmitter program reads the temperature file every 60 seconds and displays the temperature in the radio-text field of the RDS protocol on an FM station. Differential BPSK is applied to the output of the subband encoder to place the text file on the 57 kHz subcarrier on the FM spectrum. The SDR Touch app decodes the FM-RDS signal using a combination of hardware and software processing. Experimental results show that using the RDS protocol, temperature values could be reliably transmitted with 100% signal quality in a noise-free environment as shown on the advanced RDS features of the SDR Touch app. A signal quality of 95% or more is deemed adequate.
Multi-focal image fusion occupies a place in image processing research. It allows, from several images of the same scene with different blurred regions, to give a fused image without blur. This allows fusing photos taken by drones at different heights by zooming in each image a different object. Several methods are developed in the literature but which are made independently of the nature of the images. The aim of our work is to propose a method adapted essentially to images of significant fluctuations (of very large variance) considered as an alpha stable signal. For these images, we propose a method consisting of combining the Laplacian pyramid and Dumpster-Shafer theory using the alpha stable distance as a selection rule. Indeed, we decompose the multifocal images into several pyramidal levels, and apply the Dumpster Shafer method with the alpha stable distance at each level of the pyramid. The motivation of this work is to exploit the power of the dumpster Shafer fusion method and that of the Laplacian pyramidal decomposition and the fineness of the alpha stable distance. This kind of image-specific method gives better fusion because it uses a metric more suited to the nature of the data. This work was applied to some experimental images and it provides a comparison, using statistical tests, between our method and other known methods in the field of image fusion. We deduce that this method gives good fusions and that it is significantly better.
Image recognition, which comes under Artificial Intelligence (AI) is a critical aspect of computer vision, enabling computers or other computing devices to identify and categorize objects within images. Among numerous fields of life, food processing is an important area, in which image processing plays a vital role, both for producers and consumers. This study focuses on the binary classification of strawberries, where images are sorted into one of two categories. We Utilized a dataset of strawberry images for this study; we aim to determine the effectiveness of different models in identifying whether an image contains strawberries. This research has practical applications in fields such as agriculture and quality control. We compared various popular deep learning models, including MobileNetV2, Convolutional Neural Networks (CNN), and DenseNet121, for binary classification of strawberry images. The accuracy achieved by MobileNetV2 is 96.7%, CNN is 99.8%, and DenseNet121 is 93.6%. Through rigorous testing and analysis, our results demonstrate that CNN outperforms the other models in this task. In the future, the deep learning models can be evaluated on a richer and larger number of images (datasets) for better/improved results.
The Prony method for approximating signals comprising sinusoidal/exponential components is known through the pioneering work of Prony in his seminal dissertation in the year 1795. However, the Prony method saw the light of real world application only upon the advent of the computational era, which made feasible the extensive numerical intricacies and labor which the method demands inherently. The Adaptive LMS Filter which has been the most pervasive method for signal filtration and approximation since its inception in 1965 does not provide a consistently assured level of highly precise results as the extended experiment in this work proves. As a remedy this study improvises upon the Prony method by observing that a better (more precise) computational approximation can be obtained under the premise that adjustment can be made for computational error , in the autoregressive model setup in the initial step of the Prony computation itself. This adjustment is in proportion to the deviation of the coefficients in the same autoregressive model. The results obtained by this improvisation live up to the expectations of obtaining consistency and higher value in the precision of the output (recovered signal) approximations as shown in this current work and as compared with the results obtained using the Adaptive LMS Filter.
The structure of retinal blood vessels is crucial for the early detection of diabetic retinopathy, a leading cause of blindness worldwide. Yet, accurately segmenting retinal vessels poses significant challenges due to the low contrast and noise present in capillaries.The automated segmentation of retinal blood vessels significantly enhances Computer-Aided Diagnosis for diverse ophthalmic and cardiovascular conditions. It is imperative to develop a method capable of segmenting both thin and thick retinal vessels to facilitate medical analysis and disease diagnosis effectively. This article introduces a novel methodology for robust vessel segmentation, addressing prevalent challenges identified in existing literature. The methodology PSO-HRVSO comprises three key stages: pre-processing, main processing, and postprocessing. In the initial stage, filters are employed for image smoothing and enhancement, leveraging PSO optimization. The main processing phase is bifurcated into two configurations. Initially, thick vessels are segmented utilizing an optimized top-hat approach, homo-morphic filtering, and median filter. Subsequently, the second configuration targets thin vessel segmentation, employing the optimized top-hat method, homomorphic filtering, and matched filter. Lastly, morphological image operations are conducted during the post-processing stage. The PSO-HRVSO method underwent evaluation using two publicly accessible databases (DRIVE and STARE), measuring performance across three key metrics: specificity, sensitivity, and accuracy. Analysis of the outcomes revealed averages of 0.9891, 0.8577, and 0.0.9852 for the DRIVE dataset, and 0.9868, 0.8576, and 0.9831 for the STARE dataset, respectively. The PSO-HRVSO technique yields numerical results that demonstrate competitive average values when compared to current methods. Moreover, it sur-passes all leading unsupervised methods in terms of specificity and accuracy. Additionally, it outperforms the majority of state-of-the-art supervised methods without incurring the computational costs associated with such algorithms. Detailed visual analysis reveals that the PSO-HRVSO approach enables a more precise segmentation of thin vessels compared to alternative procedures.
This paper presents a new method on the use of the gammachirp auditory filter based on a continuous wavelet analysis. The gammachirp auditory filter is designed to provide a spectrum reflecting the spectral properties of the cochlea, which is responsible for frequency analysis in the human auditory system. The impulse response of the theoretical gammachirp auditory filter that has been developed by Irino and Patterson can be used as the kernel for wavelet transform which approximates the frequency response of the cochlea. This study implements the gammachirp auditory filter described by Irino as an analytical wavelet and examines its application to a different speech signals. The obtained results will be compared with those obtained by two other predefined wavelet families that are Morlet and Mexican Hat. The results show that the gammachirp wavelet family gives results that are comparable to ones obtained by Morlet and Mexican Hat wavelet family.
The spectral density of stable signals with p-adic times is already estimated under various conditions. The estimate is made by constructing a periodogram that is subsequently smoothed by a spectral window. It is clear that the convergence rate of this estimator depends on the bandwidth of the spectral window (called the smoothing parameter). This work gives a method to select the smoothing parameter in an optimal way, i.e. the estimator converges to the spectral density with the bestrate. The method is inspired by the cross-validation method, which consists in minimizing the estimate of the integrated square error.
Deep neural network (DNN) image classification has grown rapidly as a general pattern detection tool for an extremely diverse set of applications; yet dataset accessibility remains a major limiting factor for many applications. This paper presents a novel dynamic learning approach to leverage pretrained knowledge to novel image spaces in the effort to extend the algorithm knowledge domain and reduce dataset collection requirements. The proposed Omni-Modeler generates a dynamic knowledge set by reshaping known concepts to create dynamic representation models of unknown concepts. The Omni-Modeler embeds images with a pretrained DNN and formulates compressed language encoder. The language encoded feature space is then used to rapidly generate a dynamic dictionary of concept appearance models. The results of this study demonstrate the Omni-Modeler capability to rapidly adapt across a range of image types enabling the usage of dynamically learning image classification with limited data availability.
End-to-end learned image and video codecs, based on auto-encoder architecture, adapt naturally to image resolution, thanks to their convolutional aspect. However, while coding high resolution images, these codecs face hardware problems such as memory saturation. This paper proposes a patch-based image coding solution based on an end-to-end learned model, which aims to remedy to the hardware limitation while maintaining the same quality as full resolution image coding. Our method consists in coding overlapping patches of the image and reconstructing them into a decoded image using a weighting function. This approach manages to be on par with the performance of full resolution image coding using an endto-end learned model, and even slightly outperforms it, while being adaptable to different memory sizes. Moreover, this work undertakes a full study on the effect of the patch size on this solution’s performance, and consequently determines the best patch resolution in terms of coding time and coding efficiency. Finally, the method introduced in this work is also compatible with any learned codec based on a conv/deconvolutional autoencoder architecture without having to retrain the model.
Investigation of the dissolution of tablets is an important area of pharmaceutical research. Such research aims to predict the dissolution process as accurately as possible without destroying the tablets. Several methods have been published that can estimate dissolution with approximate accuracy, but they are primarily complex learning algorithms that are time-consuming and require many samples to train. This article seeks to answer whether these complex models are necessary or whether a similar result can be achieved with the help of more straightforward methods. Therefore, during this work, a simpler linear regression model was created and analysed its effectiveness in estimating the dissolution curves. The investigation concluded that the results are not as accurate as in the case of more complex methods, but the model is more robust and can be used in the case of fewer samples. Thus, by further developing these and combining the methods, we can achieve better results in the future.
The transcription accuracy of automatic speech recognition (ASR) system may suffer when recognizing accented speech. The resulting bias in ASR system towards a specific accent due to under representation of that accent in the training dataset. Accent recognition of existing speech samples can help with the preparation of the training datasets, which is an important step toward closing the accent gap and eliminating biases in ASR system. For that we built a system to recognize accent from spoken speech data. In this study, we have explored some prosodic and vocal speech features as well as speaker embeddings for accent recognition on our custom English speech data that covers speakers from around the world with varying accents. We demonstrate that our selected speech features are more effective in recognizing nonnative accents. Additionally, we experimented with a hierarchical classification model for multi-level accent classification. To establish an accent hierarchy, we employed a bottom-up approach, combining regional accents and categorizing them as either native or non-native at the top level. Furthermore, we conducted a comparative study between flat classification and hierarchical classification using the accent hierarchy structure.
When two people are on the phone, although they cannot observe the other person's facial expression and physiological state, it is possible to estimate the speaker's emotional state by voice roughly. In medical care, if the emotional state of a patient, especially a patient with an expression disorder, can be known, different care measures can be made according to the patient's mood to increase the amount of care. The system that capable for recognize the emotional states of human being from his speech is known as Speech emotion recognition system (SER). Deep learning is one of most technique that has been widely used in emotion recognition studies, in this paper we implement CNN model for Arabic speech emotion recognition. We propose ASERS-CNN model for Arabic Speech Emotion Recognition based on CNN model. We evaluated our model using Arabic speech dataset named Basic Arabic Expressive Speech corpus (BAES-DB). In addition of that we compare the accuracy between our previous ASERS-LSTM and new ASERS-CNN model proposed in this paper and we comes out that our new proposed mode is outperformed ASERS-LSTM model where it get 98.18% accuracy.
Medical Field, Robotic vision, Pattern recognition, Hurdle detection, and smart city are examples of areas that require image processing to achieve automation. Detecting an edge is an important stage in any computer vision application. The performance of the edge detecting algorithm is largely affected by the noise present in an image. An Image with a low signal-to-noise ratio (SNR), imposes a challenge to locate its edges. To improve the observable image boundaries, an adaptive filtering technique is proposed in this article. The proposed algorithm uses convolution of Gabor filter with Gaussian (GoG) operator to clean the noise before non-Maxima suppression. Furthermore, using variable hysteresis thresholding can further improve edge locating. The implementation of the algorithm was done by Python and Matlab. The obtained results were compared to a number of reviewed algorithms such as the Canny method, Laplacian of Gaussian, The Marr-Hildreth method, Sobel operator, and the Haar wavelet-based method. Three performance factors were used; PNSR, MSE, and processing time. The simulation result shows that the proposed method has higher PNSR, lower MSE, and shorter processing time when compared to the Canny detector, the Marr-Hildreth, Haar wavelet-based, Laplacian of Gaussian, and the Sobel operator methods. The higher PNSR, lower MSE, and shorter processing time mean improved edge details of the processed image.
Lung nodules are tiny lumps of tissue that are common in the lungs. The nodule may be benign or malignant; malignant nodules are cancerous and can grow rapidly. For a long time, X-ray images of the chest have been utilized to diagnose lung cancer. We developed in this paper a computer aid diagnosis system (CAD) to atomically classify a set of lung x-ray images into normal and abnormal (with nodule and no-nodule) cases. We used 180 images in this work, the images are in full size no filtering or segmenting process were applied, 75 of them are for normal cases and the other 105 are for abnormal cases, at the same time 120 of the images have been used to train the classifier and 60 for testing. Our classifiers were fed with a variety of features, including LBP (local binary pattern) and statistical features. And a classifier was able to identify cases with nodule from cases without nodule with an accuracy (ACC) of 86.7%.
The recognition of handwritten digits has aroused the interest of the scientific community and is the subject of a large number of research works thanks to its various applications. The objective of this paper is to develop a system capable of recognizing handwritten digits using a convolutional neural network (CNN) combined with machine learning approaches to ensure diversity in automatic classification tools. In this work, we propose a classification method based on deep learning, in particular the convolutional neural network for feature extraction, it is a powerful tool that has had great success in image classification, followed by the support vector machine (SVM) for higher performance. We used the dataset (MNIST), and the results obtained showed that the combination of CNN with SVM improves the performance of the model as well as the classification accuracy with a rate of 99.12%.