
Humans can exhibit diverse emotional facial expressions, including happiness, sadness, anger, disgust, surprise, and fear. Distinct and distinctive traits characterise these emotional expressions. Integrating emotion-specific features into landmark-based face recognition systems by developing features based on emotional expressions can significantly enhance face recognition capability. Such landmark-based systems are essential for face recognition tasks in resource-constrained environments. By combining these emotion-specific features with the traditional features based on neutral faces, we can develop more accurate and robust face recognition systems that can better adapt to the dynamic nature of human emotional expressions, promising a new level of efficacy in face recognition.
Credit risk prediction is essential for financial institutions. Several studies have shown that oversampling and instance selection can improve credit risk prediction. However, the combination of these methods for credit risk prediction has been less studied. This paper conducts empirical research comparing the performance of the most widely used oversampling and instance selection methods and their combined use through successful supervised classifiers for credit risk prediction. The practical significance of our research is the assessment of the different options of application of oversampling and instance selection (including the combination of both that has not been studied before) in credit risk prediction under the same framework employing the most used standard-public data sets and one new dataset that allowed us to conclude which is the best option for credit risk prediction.
We propose a lightweight facial expression recognition (FER) model that integrates local binary pattern (LBP) and grey level co-occurrence matrix (GLCM) hybrid features with a channel attention mechanism based on the ECA.Net variant to effectively balance the accuracy-efficiency trade-offs. This model leverages texture features extracted via LBP and LBP+GLCM, combined with attention mechanisms to enhance the focus on salient facial cues. Built upon the VGG16 CNN architecture within the TensorFlow framework, the model was tested on the FER2013 and RAF-DB datasets for cross-validation, achieving a recognition accuracy of 79.89% and 86.77%, respectively. To validate its practical effectiveness, a deployable QT-based user interface system was developed, enabling real-time FER from images, videos, and live camera feeds. This approach aims to meet the increasing demand for lightweight, reliable solutions, particularly in scenarios with limited computational resources, while maintaining high recognition accuracy.
Accurate emotion detection from speech signals is essential for enhancing human-computer interaction (HCI) systems. However, existing SER methods often suffer from poor feature representation and limited dataset diversity, resulting in suboptimal performance. To address these challenges, this paper proposes an advanced deep bottleneck residual convolutional neural network (DBR-CNN) integrated with the SEResNeXt-101 feature extraction framework and optimised using the coati optimisation algorithm (COA). The model is trained and evaluated on four diverse benchmark datasets: URDU, EMO-DB, EMOVO, and SAVEE. In the pre-processing phase, speech signals undergo noise reduction and normalisation to enhance data quality. The SEResNeXt-101 extractor then captures high-level features with reduced complexity, which are subsequently processed by the DBR-CNN to classify emotions with greater accuracy. The COA fine-tunes the model to improve classification efficiency. Experimental results demonstrate that the proposed model significantly outperforms state-of-the-art (SOTA) methods. The optimised DBR-CNN framework provides robust, scalable, and highly accurate emotion recognition.
In this paper, we deal with document image binarisation, as an essential image pre-processing step, through presenting a comparison study of the different methods of the literature in terms of effectiveness. In addition, this sub-field of document image binarisation, including a lot of methods, techniques, and algorithms, proposed along the recent decades of years, is overviewed here through categorising the various methods into their different approaches as well as talking about the evaluation step via showing the metrics and the considered public datasets and benchmarks. Some issues caused by different types of noise, some frequent utilised filters, considered as pre-processing, as well as some other techniques of post-processing are also presented. To note that the results compared here are as given in the different references that are either regular papers or comparative studies in the form of competitions.
This paper presents a comprehensive analysis of 11 state-of-the-art deep convolutional neural network (CNN) models for COVID-19's X-ray image classification with the two configurations more studied in the literature: transfer learning with fine-tuning and training from scratch. All models were assessed under the same experimental framework. Unlike other works, we used a dataset compiled from several public datasets, increasing its variability to reduce the risk of overfitting. Our results show which deep convolutional neural networks performed the best in accuracy and F1-score when training from scratch and with transfer learning.
Electroencephalogram (EEG) is a complex, nonlinear signal which requires extensive training for detection of changes due to sleep disorder. Most of the traditional machine learning algorithms has been used in the past for detection of sleep disorder subjects. Recently deep learning has demonstrated a very promising approach for sensing EEG signals as it has excellent capacity of extracting features from raw signals. The proposed work aims to differentiate sleep disorder subjects from normal subjects using a deep learning-based model. To examine this, open-source EEG dataset from ten different electrodes of six sleep disorder subjects and six normal subjects is used here. Long short-term memory (LSTM) model, a class of recurrent neural network (RNN) is proposed for detection of sleep disorder subjects. Finally, in Table 3, accuracies are compared which are obtained in various models applied on same dataset. It is clearly predicted that the offered LSTM based technique gives classification performance of 70.75% accuracy as compared to other techniques in literature survey. Along with accuracy, recall of 88.34%, precision of 65.35% and specificity of 53.17% is evaluated for proposed LSTM model.
Face recognition biometric recognises human faces effectively where their performance is critically affected under deviating light effects. This work presents an efficient illumination invariant feature extraction technique using homomorphic filtering in integer wavelet transform (IWT) domain. The goal of this investigation is to subdue the low frequency components in small-scale extracted features with the simultaneous perpetuation of rugged texture components in face images. The technique exploits homomorphic filtering based illumination normalised (HFIN) images which are then utilised in analysing the low and high pass frequency coefficients. Furthermore, IWT-based multiscale features (MFIWT) over HFIN images are examined with orthogonal and biorthogonal wavelets. The HFIN-MFIWT features are hereafter mapped onto non-correlated lower dimensional subspace using eigenface mechanism. Significant facial features are then classified using K-nearest neighbour. The efficacy of HFIN-MFIWT approach is assessed on Yale, Yale B, CMU-PIE, and extended Yale B databases that evidently authenticate its effectiveness.
One of the critical pre-processing phases in a pattern recognition system is thinning. The process of finding a one-pixel-wide representation of a binary object while preserving its original shape is known as 'thinning'. Thinning makes recognition easier and more efficient due to the fact that the thin image of the object is less complex than the original one. It is also easy to find topological features like endpoints, junction points, lines, curves, etc. The general problem with the thinning method is that it often results in unwanted edges or hairs that deform the original shape of the object and ultimately decrease the system's accuracy. The current proposal presents a contour-based thinning algorithm to address this issue. The fundamental part of the suggested algorithm is to extract a skeleton from the initial binary object. The potential distribution of pixels in the junction region is estimated in the second phase using an auto-encoder trained on similar thinned images in order to provide more accurate thinned images. The third stage employs post-processing to produce an image that is one pixel wide. The proposed method was tested on a handwritten Gujarati numeral database with 99.89% accuracy.
Identifying and verifying a person based on scanned images of their handwriting is a needful biometric application in historical document analysis, behavioural biometrics study, forensic science, access control, graphology, and copyrights management. Writer identification and verification are still challenging in offline and online handwriting recognition. Since the performances of handwriting biometric identification and verification systems depend on both the quality and types of chosen features, this is one of the most critical phases. This article represents a literature survey on offline and online biometric features used in different scripts for writer verification and identification techniques. Several previous efficient works on online and offline writer authentication methods for biometrics using cutting-edge hand-craft features in different levels of handwriting analysis like documents, paragraphs, words, and characters are analysed systematically to date for the first time in detail.
A new artificial neural network (ANN)-based approach has been proposed in this paper to recognise handwritten digits. Handwritten digit recognition finds its applications in many areas of computer vision and artificial intelligence. The proposed ANN has a logical framework of five levels. Three hidden layers independently capture the features of a digit; then associative relationship among the features followed by the possible forms of a handwritten digit. The performance of the neural network is analysed by varying the number of nodes in these three layers. It is further suggested to pre-process the data to avoid the problem of overfitting in which case the noise is incorporated into the model instead of the signal. The data are pre-processed for removing white spaces outside the boundary of a digit's image, considering them as noise. In addition, the dropout strategy of Srivastava et al. (2014) has also been implemented, resulting in a better accuracy at a cost of about 18% of extra CPU time. Finally, the optimised size of the neural network with the proposed architecture is also determined to yield the best performance. The performance of the proposed architecture was found to be very close to that of Srivastava et al. (2014), but comparatively very small in size and requiring much less CPU time.
Mobile phone usage during driving is identified as one of the major causes of traffic accidents as it distracts the driver, mainly during driving the motorcycle. In this article authors are focused on detection of mobile phone usage during motorcycle driving. It has been observed that limited research work has been done in this domain due to the lack of ready datasets, occlusion of object (mobile phone), rotation and difficulty in extracting the features object. The authors collected the data in different Indian traffic conditions and applied convolutional neural network (CNN), deep learning-based YOLOv4 architecture with CSPDarknet-54 as the backbone of YOLOv4 algorithm. The results show the detection of mobile phone usage in traffic with a precision of 94%.
Gujarati script has a large number of characters with curvature shapes and complexities. The script can contain a variety of character sets like vowels, consonants, numerals, modifiers, conjunct characters, and other combinations of characters. Different forms of conjunct characters are possible but in this particular paper one of the forms of frequently used conjunct characters is considered for assessment point of view. This paper deals with the identification of typewritten and handwritten conjunct characters. There is no benchmark dataset available for Conjunct Gujarati characters; so, a train and test datasets are created using various features of characters such as a number of open edges, location of open edges in zone, pixels count on a constructed horizontal and vertical line. Artificial neural network is used for the classification of characters and a success rate of 99.4% and 94.1% is achieved in most of the typewritten and handwritten Gujarati conjunct characters respectively.
Parkinson's disease (PD) is a neurological disorder with more than six million people worldwide suffering from it. It is commonly diagnosed using clinical assessments and progression scale, which usually depends on the medical practitioner's expertise, and accuracy varies greatly between various examiners which also takes a long time to accurately diagnose. This paper proposes to develop a computer-aided diagnostic method to diagnose PD patients using MRI images of the brain, thus reducing cross-examiner variability and the time required to accurately differentiate between PD and control subjects. We have developed a CNN-based CAD system that classifies between PD and healthy patients, by utilising the differences between the Substantia Nigra region of the brain. The method proposed in this paper was successfully able to achieve an accuracy of 99.5%.
In this paper, an extensive study of diagnosis of breast cancer is made using support vector machine (SVM) technique. To build the cost-effective kernel machine for breast cancer diagnosis, the tools of principal component analysis (PCA) and k-fold cross-validation (CV) techniques are employed. The model is implemented on WDBC and WBC datasets to check the condition of the tumour for its malignancy. Classification accuracy and computation time are obtained and comparative experimental results are analysed under different conditions. For WBC dataset, 100% accuracy is obtained using polynomial kernel in just 0.03 second.
Due to non-invasive and easy to acquire procedure, electrocardiogram (ECG) is broadly adopted for extracting the correct heart health status of the subject (patient). In this paper, a method of analysing ECG signals (based on R-peak detection) recorded during different postures (i.e., sitting, standing and supine) has been presented. The results of this analysis are also compared with standard physioNet database (MIT-BIH arrhythmia database) for validation. Real time ECG signals in all three postures were acquired for 500 subjects, out of which signals of 17 subjects are used for analysis. For filtering these subjects' recordings, existing techniques need a higher order of analogue and digital filters. It increases the complexity of the system, which motivated us to use the combination of principal component analysis (PCA) and independent component analysis (ICA). This combination fulfils the need of higher order filters (HOFs). PCA is used for dimension reduction, whereas ICA is used for noise removal.
The actor-critic models are generally prone to overestimation of sub-optimal policies and Q-values. Our proposed approach is established on value-based deep reinforcement learning algorithm also known as twin delayed deep deterministic policy gradient algorithm or TD3. The suggested approach is used to solve complex reinforcement learning problem like half-humanoid robot, ant, and half-cheetah to cover a path. This problem can only be solved with an algorithm which can work on continuous-action spaces, without much delaying the result to propagate during the inference of model. The proposed model has been adapted to converge faster to optimal Q-values. The TD3 uses two deep neural networks for learning two Q-values, viz., Q1 and Q2; in the proposed approach the Q-values average is being taken as an input for final Q-value unlike the other reinforcement learning algorithm such as DDPG which is prone to overestimate the Q-values. The proposed approach has also made self-adjusting noise clipping function, which make it harder for the policy to exploit Q-function errors to further improve performance.
Increased digital transactions accentuate the need for secure communication over open channels. Confidentiality and safe distribution of shared key is a mandatory requirement in symmetric key-based systems. The work proposes a novel nature inspired optimisation technique for binding secret key with user traits extracted from iris biometrics. Validation of key binding is demonstrated by ensuring that successful decryption happens by authorised user alone. Nature inspired swami and population algorithms are used to extract optimal feature vectors from user trait. Chicken swami optimisation and deer hunting optimisation algorithms have been used for the first time with iris traits to achieve optimal key binding. Experiments for different shared key lengths have been carried out with IIT Delhi and Multimedia University iris datasets. Accuracy of the proposed model is 7% better than whale optimisation algorithm and 4% better than grey wolf optimisation.
In the last few decades, content-based image retrieval is considered as one of the most vivid research topics in the field of information retrieval. The limitation of current content-based image retrieval systems is that low-level features are highly ineffective to represent the semantic contents of the image. Most of the research work in content-based image retrieval is focused on bridging the semantic gap between the low-level features and high-level semantic concepts of image. This paper presents a thorough study of different techniques for the reduction of semantic gap. The existing techniques are broadly categorised as: 1) image annotation techniques to define the high-level concepts in image; 2) relevance feedback techniques to integrate user's perception; 3) machine learning and deep learning techniques to associate low-level features with high-level concepts. In addition, the general architecture of semantic-based image retrieval system has been discussed in this survey. This paper also highlights the current and future applications of content-based image retrieval. The paper concludes with promising future research directions.