IntroductionThis study presents a Field-Programmable Gate Array (FPGA)-based convolutional neural network (CNN) accelerator for preliminary Parkinson's disease (PD) handwriting classification using hand-drawn circle images, with emphasis on arithmetic-level optimization through efficient multiplier architectures. Although optimized multipliers have been extensively studied for machine learning acceleration, their application-specific effects on inference consistency and hardware efficiency in healthcare-oriented FPGA implementations remain underexplored.MethodsTo address this issue, a lightweight binary CNN classifier, trained on Google Colab, is deployed on FPGA hardware and evaluated with three multiplier architectures: standard multipliers, approximate logarithmic multipliers, and Karatsuba multipliers. The desktop CNN model was evaluated using both non-augmentation validation and standard augmentation strategies. The primary evaluation methodology used a non-augmentation validation approach, in which augmentation was applied exclusively to the training set, resulting in a software validation accuracy of 92.86%. The standard augmentation strategy achieved a validation accuracy of 97.83% and was used to compare the effects of augmentation before splitting. The trained model was quantized to Q4.12 fixed-point precision and implemented on FPGA hardware, where dense-layer computations were performed using different multiplier architectures. Hardware inference was validated on the NewHandPD hand-drawn circle dataset, and FPGA outputs were compared with software inference results via graphical analysis and Mean Absolute Deviation (MAD) to assess numerical consistency.ResultsExperimental results indicate that the FPGA-based implementation achieved classification behavior closely aligned with software inference while improving hardware efficiency. Under the non-augmentation validation approach, the FPGA implementation achieved 89.73% accuracy compared to 92.86% in software, whereas the standard augmentation strategy achieved 95.40% accuracy compared to 97.83% in software using the Approximate Logarithmic Multiplier.DiscussionThe results indicate the potential feasibility of lightweight CNN deployment with optimized multipliers for resource-efficient edge healthcare applications. However, due to the limited dataset size, the presented findings should be interpreted as a preliminary proof-of-concept study rather than definitive clinical validation.
In a noisy industry environment, to predict machine faults using vibration signals, a specially designed Deep Convolution Neural Network (DCNN) with an additional noisy layer has been recently demonstrated. On the contrary, this paper presents a noise-susceptible fault classification using frequency spectrums in standard DCNN. The study involves two types of spectrum representation (a) Short-time Fourier Transform (STFT)of raw original signal and (b) Hilbert Huang Transform (HHT) of Empirical mode decomposition-intrinsic mode function (EMD-IMFs) and two different datasets (i) CWRU Bearing dataset and (ii) Nasa Milling Dataset which is non-stationary. Three binary DCNN classification problems are performed. For bearing data both representations maintained 100% classification accuracies for noise range up to 10 dB. HHT-EMD-IMF performs better for the non-stationary milling dataset. HHT-EMD-IMF is extended to bearing data set multiclass fault classification problem. The performance is compared with state-of-the-art work. The study compares all IMFs of clean and noisy signals to quantify the impact of noise on EMD for 8 different specific faults of the CWRU bearing dataset. The analysis of average normalized noise shows that EMD-IMF 0 has minimum deviation due to noise for all noise ranges. The DCNN model maintains 100% classification accuracy at high noise levels up to 10 dB and is better than the state-of-the-art noisy CNN approach. An ablation study shows that the proposed method is highly susceptible to impulse noise as well. It is also shown that the proposed method does not need additional computation time for training as in noisy layer CNN.
Research on Compressive Sensing (CS) framework for ECG signals are of 3 categories. Category 1) involves CS for a single frame of ECG signal and performance analysis of reconstructed frame. Category 2) involves CS for every frame of the complete ECG signals and arrhythmia classification using deep learning approach on the reconstructed ECG signals. Category 3) is same as category 1) but the classification involves direct compressed measurements without reconstruction. The key merit of category 2) is that the signal can also be visually inspected by doctors. This paper falls in category 2) and proposes a hybrid CS algorithm for ECG signal reconstruction using discrete wavelet transform (DWT) and discrete cosine transform (DCT) both as sensing and sparsifying operator and the classification of reconstructed signal is achieved using DCNN for two different image representations (a) CWT and (b) STFT. Using MIT-BIT ECG dataset, for the proposed hybrid CS implementation the reconstruction is achieved with PRD=2.591, SNR=36.530dB and SSIM=0.745587 and the classification accuracy achieved for binary class problem with CR=0.5 is 93% for CWT image representation and 98.4% for STFT image representation. The proposed hybrid CS method performs better than the conventional non hybrid method for all the performance metrics and the classification results are better than other existing state of art methods for category 2) & 3). The same experiment is extended to a 3-class DCNN classification problem and the accuracy is 93.33% for the STFT image.
Parkinson's Disease (PD) is a progressive neurodegenerative condition that effects motor and non- motor functions, such as speech and movements. Early diagnosis of PD continues to be a clinical challenge because of its subtle early symptoms. In this paper, we present a machine learning-based approach to early diagnosis of Parkinson's Disease from voice recordings, improved with nature-inspired feature selection methods. Audio features from the MDVR-KCL dataset are preprocessed and evaluated using several classifiers, such as Support Vector Machine (SVM), K-Nearest Neighbors (KNN), XGBoost, and Naïve Bayes. For enhanced classification accuracy and model simplicity, we use Firefly Algorithm (FFA) and Particle Swarm Optimization (PSO) for feature selection. From experimental findings, we have the best accuracy of 86.67% with the combination of XGBoost and PSO, proving the efficiency of bio-inspired optimization in improving PD detection. The suggested method offers a affordable and scalable solution for early Parkinson's Disease diagnosis, which can also be used in real-world telehealth and mobile-based screening systems.
The Hilbert spectrum images of intrinsic mode functions (IMF) of empirical mode decomposition (EMD) analysis and variational mode decomposition (VMD) analysis of faulty machine vibration signals are used in deep convolutional neural network (DCNN) for machine fault classification in which the DCNN automatically learns the features from spectral images using convolution layer. Though both EMD and VMD analysis suit well for non-stationary signal analysis, VMD has the merit of aliasing free IMFs. In this paper, the performance improvement of DCNN classification for a non-stationary vibration signal dataset using VMD is brought out. The numerical experiment uses the Hilbert spectrum images of 4 EMD-IMFs and 4 VMD-IMFs in DCNN to classify 10 different faults of the Case Western Reserve University (CWRU) bearing dataset. The confusion matrices are obtained and the plot of model accuracies in terms of epochs for the DCNN is analysed. It is shown that the spectrum images of one of the four EMD-IMFs, IMF0, give a validation accuracy of 100% and in the case of VMD the spectrum images of two of the four VMD-IMFs, IMF0, and IMF1 give a validation accuracy of 100%. This reveals that non-aliasing IMFs of VMD are better at classifying bearing faults. Further to bring out the merits of VMD analysis for non-stationary signals the numerical experiment is conducted using VMD analysis for binary fault classification of the milling dataset which is more non-stationary than the bearing dataset which is proved by plotting the statistical parameters of both datasets against time. It is found that the DCNN classification is 100% accurate for IMF3 of VMD analysis which is much better than the 81% accuracy provided by EMD analysis as per existing literature. The performance comparison highlights the merits of VMD analysis over EMD analysis and other state-of-the-art methods and ensemble learning methods.
A complex motor speech disorder, dysarthria makes diagnosis and its severity classification extremely challenging, thereby affecting suitable therapy and intervention strategies. This paper presents a deep learning-based method based on TORGO dataset to overcome these challenges. Moreover, the problem statement focuses on the difficulty of exactly spotting dysarthria and assessing its degree of severity using traditional methods, which usually lack precision and efficiency. This work presents a new method combining advanced acoustic feature extraction techniques, such Mel-frequency cepstral coefficients (MFCC) and spectrogram analysis, with state-of- the-art neural network and its hybrid architectures such convolutional neural networks (CNNs), long- and short-term memory (LSTM) with CNN, and gated recurrent unit (GRU) combined with CNN. It offers an extensive framework for assessing the degree of dysarthria and also uses short-time Fourier transform (STFT) images obtained from a dataset for severity classification. The proposed CNN model obtained an accuracy of 98.2% using Mel-spectrogram for detecting the dysarthria and the hybrid CNN-GRU model reached an accuracy of 97% using the STFT images for classifying dysarthria based on its severity. Moreover, this work highlights the ability of proposed deep learning models to offer tailored therapy approaches depending on degree of severity and automates dysarthria diagnosis process.
The automated generation of a NLP of an image has been in the spotlight because it is important in real-world applications and because it involves two of the most critical subfields of artificial intelligence, computer vision and natural language processing. The image captioning process involves generating descriptive explanations of the content and activities contained within an image. Image Captioning is an approach that employs descriptive language to make explanations of images. Image captioning is an extremely useful tool in many applications, including the analysis of large sets of unlabeled photographs and discovering hidden patterns, which are of great importance in machine learning applications. Besides, image captioning is also crucial in the creation of software used to navigate autonomous cars and also for assisting visually impaired people. The Image Captioning task can be achieved through the use of Deep Learning Models. The achievements in deep learning and Natural Language Processing have made it possible to create captions for images. This research utilizes Neural Networks as an approach to image captioning. The Convolutional Neural Network (ResNet) is used as an encoder, which removes the features of the pictures. The Recurrent Neural Network (Long Short Term Memory) acts as a decoder, creating captions for the images based on the removed image features and a built vocabulary.
Compressive Sensing (CS) is a form of lossy compression that enables signal reconstruction using a reduced number of samples. This technique for ECG signal reconstruction offers the advantage of improved reconstruction quality while minimizing noise in the recovered signal. This study presents a comprehensive study of CS for ECG signal reconstruction using two well known fast transform, Discrete Wavelet Transform (DWT) and Discrete Cosine Transform (DCT). Two different cases are examined: case (1) DCT as the sensing operator and DWT as the sparsifying operator and case (2) DWT as the sensing operator and DCT as the sparsifying operator, for both the cases 5 different various wavelet kernel such as Haar, Daubechies 3 (db3), Symlet 6 (Sym6), Coiflet 5 (coif5) and Biorthogonal (bior) wavelets are utilized. A detailed study was carried out to determine the most preferable wavelet kernel using 5 different performance measures, Root Mean Square Error (RMSE), Percent Root Mean Square Difference (PRD), Signal-to-Noise Ratio (SNR), Structural Similarity Index Measure (SSIM) and the correlation coefficient (r). Based on the findings, the bior wavelet demonstrated superior performance compared to all other wavelets, achieving a lower PRD value of 0.087373 and a higher SNR value of 55.873135 dB.
In speech analysis-based Parkinson’s disease (PD) classification, unlike binary classification very small number of research had been focused on multiclass classification of severity grading of PD. In this paper, a novel scheme is proposed in which convolution neural network (CNN) learning on line spectral frequency (LSF) spectrum is proposed for PD detection and CNN learning on short time Fourier transform (STFT) spectrum is proposed for PD severity grading. To justify the proposed scheme, the Kings college London (KCL) running speech PD dataset that consists of speech recordings of PD with severity level labelled using Unified Parkinson’s Disease Rating Scale (UPDRS)-III-18 standard and Healthy control in the form of read text and spontaneous dialogues is used. The dataset is arranged in 10 different combinations that includes 4 binary classification detection problem and 6 multiclass classification severity grading problem and CNN learning experiments are conducted using LSF, STFT and Mel frequency cepstral coefficients (MFCC) spectrums. The numerical implementation reveals that CNN learning on LSF speech spectrum outperforms the CNN learning that is done on both STFT and MFCC spectrums for all 4 binary class combinations with the best validation accuracy of 87.50 %. The other important inference made is that CNN learning of STFT spectrums performs the best for severity grading in all 6 multiclass combinations and the best validation accuracy is 93.75 % for a three class PD severity grading problem which is a very high validation accuracy achieved and as per comparison it is inferred that the method outperforms the state-of-the-art methods.
Recent advancements in computer techniques have shown promise in aiding in the early diagnosis of neurodevelopment conditions in growing children. Our research focuses on an innovative approach to detect Autism Spectrum Disorder (ASD). By integrating MambaVision as a robust feature extraction tool with machine learning algorithms such as Random Forests and XGBoost, we demonstrate enhanced accuracy on ASD classification. Our approach leverages the powerful feature extraction capabilities of the MambaVision framework that encompasses patterns and structural facial cues using state space models for long range visual dependencies associated with ASD, which then serve as inputs to a machine learning algorithm best suited for handling high dimensional data. Finally, we assess the computational efficiency of our proposed method and demonstrate that it is suitable for real time applications such as early-stage ASD screening tools in educational and medical institutions.
Dysarthria, a speech disorder stemming from neurological conditions, affects communication and life quality. Precise classification and severity assessment are pivotal for therapy but are often subjective in traditional speech-language pathologist evaluations. Machine learning models offer objective assessment potential, enhancing diagnostic precision. This systematic review aims to comprehensively analyze current methodologies for classifying dysarthria based on severity levels, highlighting effective features for automatic classification and optimal AI techniques. We systematically reviewed the literature on the automatic classification of dysarthria severity levels. Sources of information will include electronic databases and grey literature. Selection criteria will be established based on relevance to the research questions. The findings of this systematic review will contribute to the current understanding of dysarthria classification, inform future research, and support the development of improved diagnostic tools. The implications of these findings could be significant in advancing patient care and improving therapeutic outcomes for individuals affected by dysarthria.
It is proposed to use Deep Convolution Neural Network (DCNN) which is a good classifier of natural images to learn speech spectrum images of sustained phonation to detect Parkinson's Disease (PD) as an alternative to the existing feature-based machine learning method. It is shown that the proposed method yields very high accuracy without the need for separate feature computation stage. The speech spectrum representations proposed are Short Time Fourier Transform (STFT) spectrum of size Nx256 and Line Spectral Frequency (LSF) spectrum of size Nx16. LSF reflects the speech production mechanism and it is a novel idea to use LSF spectrum in DCNN to detect PD speech. The spectrum images look like random patterns and the performance is improved when using an additional deeper hidden layer of tampering pattern in the last stage of a fully connected layer. Using a standard PD-sustained phonation dataset the training accuracies achieved are 98.50% and 92.50% for STFT and LSF method, respectively. The validation accuracies achieved are 84.38% for STFT and 100% for LSF. The STFT method results in a sensitivity of 97.05%, a specificity of 88.63%, a precision of 86.84%, an F1-score of 91.66, a false positive rate (FPR) of 11.36%, and a false alarm rate of 12.82%. The LSF method results in a sensitivity 97.05%, a specificity of 95.45%, a precision of 94.28%, an F1-score of 95.65, an FPR of 4.50%, and a false alarm rate of 5.71%. The LSF based method performs better and the performance comparison with the state-of-the-art methods brings out the merits of the LSF spectrum image-based DCNN learning in PD detection using sustained phonation.
In response to the growing need for enhanced energy management in smart grids in sustainable smart cities, this study addresses the critical need for grid stability and efficient integration of renewable energy sources, utilizing advanced technologies like 6G IoT, AI, and blockchain. By deploying a suite of machine learning models like decision trees, XGBoost, support vector machines, and optimally tuned artificial neural networks, grid load fluctuations are predicted, especially during peak demand periods, to prevent overloads and ensure consistent power delivery. Additionally, long short-term memory recurrent neural networks analyze weather data to forecast solar energy production accurately, enabling better energy consumption planning. For microgrid management within individual buildings or clusters, deep Q reinforcement learning dynamically manages and optimizes photovoltaic energy usage, enhancing overall efficiency. The integration of a sophisticated visualization dashboard provides real-time updates and facilitates strategic planning by making complex data accessible. Lastly, the use of blockchain technology in verifying energy consumption readings and transactions promotes transparency and trust, which is crucial for the broader adoption of renewable resources. The combined approach not only stabilizes grid operations but also fosters the reliability and sustainability of energy systems, supporting a more robust adoption of renewable energies.
The basic functioning of heart can be read through Electrocardiogram (ECG) Signal, this signal gives an idea whether the functioning of heart is normal or abnormal and type abnormality can also be identified, which helps to diagnose the patients in time. This work investigates a deep-learning model using 2DCNN to classify various category of ECG signal. This proposed CNN model is trained and tested to classify three different classes of heart arrhythmia such as cardiac arrhythmia (ARR), congestive heart failure (CHF) and normal sinus rhythms (NSR). The time domain ECG signal is preprocessed and further it is transformed in to time-frequency scalogram by utilizing continuous wavelet transform (CWT), these scalogram is remodeled and saved as RGB images with necessary dimensions. Later these converted RGB images are fed to the input of various 2DCNN models such as alexnet, vgg16, squeezenet and googlenet to classify arrhythmia type. ECG Recordings from MIT BIH database were chosen and used for training and testing dataset. The performance of proposed scheme is evaluated on various CNN networks, a reasonable classification accuracy of 99.33 % was acheived by alex net.
Speech deterioration is a well-established indicator for early detection of Parkinson disease (PD). The remote monitoring of PD symptoms is practicable by analyzing the person's speech. This study demonstrates that deep learning of speech spectrum images of Running speech from Parkinson's disease (PD) patients and healthy controls (HC) results in a convolution neural network model with high accuracy for classifying PD. Earlier work for PD detection was totally based on the computation of speech features. Complicated speech feature computations take time and require a large amount of memory access. Using a spectrum image representation of speech in a deep convolution neural network gives very high accuracy without the need for feature computation, and selection in this work. This is possible because the convolution layer learns the features from the spectrum automatically and in this work, it is proposed to work with PD speech database for Running speech. In this case, the short time Fourier transform (STFT) is a good choice for representing speech as images. The training accuracy achieved in the experiment is 99.48% and validation accuracy are 79.10%. This method performs best and is a well-founded method for classifying PD from running speech signals.
Maintaining the water level in a dam is really important. The water level in the dam affects the amount of potential energy present to generate electricity, higher water level gives high potential energy. Also, the water level should not rise above the crest of the dam as it increases the risk of bursting the dam and causing major floods in nearby villages and cities. To overcome this problem, we have created a smart dam system using RS-485 serial communication protocol between a single-chip embedded controller ARM cortex STM32F103C8T6 and ESP8266. The water level will be measured using a capacitive sensor connected with STM32. STM32F acts as the master and serially communicates with ESP8266. Servo motor and Infrared sensors will be connected with ESP8266 to open the floodgate and detect the obstacle near the floodgate. The floodgate will be opened automatically based on the water level. The water level and status of the floodgate will be updated lively in a website that we created and a notification will be sent to the locals in case of any emergency due to the water level using IFTTT.
Early diagnosis means an individual gets an indication about the disease on his or her own at very early stage of the disease. Today, many people across the globe are suffering from Parkinson's disease (PD). Early detection of Parkinson's disease can be a better choice to treat the disease much early. Vocal cord disorder, speech impairments/speech disorders are the early indicators of PD. The initial stage of PD affects the human speech production mechanism. The speech impairments are not apparent to common listeners. We should monitor carefully for the initial stage of PD by using proper expert systems. In this review, we mainly focused on speech signal analysis for the identification of PD with the help of machine learning techniques. The voice sample of affected people from PD can be used in an early detection algorithm using various classification models with different accuracy, sensitivity, specificity, etc. In our review, we found that mainly two types of techniques have been used in this problem (a) conventional feature-based techniques and (b) machine learning-based techniques. The detailed review using these types of algorithms is presented in this paper. In feature-based applications, mel frequency cepstral coefficient (MFCC) and linear predictive coding (LPC) are the mostly used features. Machine learning-based algorithms used intelligent architecture like artificial neural network (ANN), convolution neural network (CNN), hidden Markov model (HMM), XG boost, support vector machine (SVM), etc. It is found that machine learning-based algorithms are doing better in terms of highest accuracy but with some limitations.
A spectrum-image based representation of machine vibration signals with deep convolution neural network is proposed for machine fault classification in which the convolution layer is used for automatic feature extraction as an alternate to the conventional feature-based methods. Two different forms of spectrum representations are proposed, one based on the short time Fourier transform of the original signals and the other based on the short time Fourier transform of the intrinsic mode functions acquired by empirical mode decomposition. Empirical mode decomposition has its own merits in discriminating non stationary signals and the novelty of the work is to use the short time Fourier transform of intrinsic mode functions with deep convolution neural network model. The classification and validation accuracy of the model are investigated with respect to epochs. It is demonstrated that both spectrum-based techniques perform good with 100% model accuracies in a numerical experiment of binary classification on a bearing dataset that comprises of normal and faulty signals. In another experiment using milling data set, short time Fourier transform of intrinsic mode functions representation performs better with 100% training accuracy, F1 score of 0.8933 which is better than that of using short time Fourier transform of raw signals whose training accuracy is 64% and F1 score of 0.7486. The numerical study shows that the empirical mode decomposition based spectrum representation delivers the highest accuracy in the learning model obviating the necessity for independent feature extraction, feature selection, and dimension reduction. The numerical experiment is extended using empirical mode decomposition based spectrums for multiple class classification problems in bearing dataset. The confusion matrix obtained for 10 classes, shows that validation accuracy is 100% for all classes. The performance comparison throws light on the merits of empirical mode decomposition spectrum method over other state of the art methods.
Dysarthria is a neurological speech disorder that can significantly impact affected individuals' communication abilities and overall quality of life. The accurate and objective classification of dysarthria and the determination of its severity are crucial for effective therapeutic intervention. While traditional assessments by speech-language pathologists (SLPs) are common, they are often subjective, time-consuming, and can vary between practitioners. Emerging machine learning-based models have shown the potential to provide a more objective dysarthria assessment, enhancing diagnostic accuracy and reliability. This systematic review aims to comprehensively analyze current methodologies for classifying dysarthria based on severity levels. Specifically, this review will focus on determining the most effective set and type of features that can be used for automatic patient classification and evaluating the best AI techniques for this purpose. We will systematically review the literature on the automatic classification of dysarthria severity levels. Sources of information will include electronic databases and grey literature. Selection criteria will be established based on relevance to the research questions. Data extraction will include methodologies used, the type of features extracted for classification, and AI techniques employed. The findings of this systematic review will contribute to the current understanding of dysarthria classification, inform future research, and support the development of improved diagnostic tools. The implications of these findings could be significant in advancing patient care and improving therapeutic outcomes for individuals affected by dysarthria.