Falls among the elderly pose a significant risk, often leading to serious injuries and a decline in overall well-being. This study employs an Heterogenous Hidden Markov Model (HHMM) that utilizes 3D vision-based body articulation data to propose an innovative method for fall detection and prediction. To ensure the precision and reliability of our model, we preprocessed the data to eliminate noise and extract pertinent features. This involved using a ZED camera to capture joint locations and body movements at a high frequency. The dataset was divided into 40% for testing and 60% for training the HHMM model, which comprised four states representing different body positions. The model achieved a 61% prediction rate with an accuracy of 81.51%. The Viterbi algorithm facilitated real-time recognition of body postures and fall predictions. The study suggests that HHMM can improve safety monitoring systems in healthcare and senior living facilities. To enhance prediction accuracy, future research could focus on incorporating additional data sources and expanding the dataset. In Conclusion, the study concludes that HHMM has the potential to effectively recognize and predict falls, thus contributing to fall prevention measures.
The elderly have a sensitive period of life in terms of physical and mental health and require close assistance or a caregiver. Medical assistance has the ability to recognize the emotional states of older adults through facial expressions and take care of them in real time. This paper provides a comprehensive Facial Emotion Recognition (FER) review, especially for the elderly. Several studies have been conducted on the facial emotion recognition of young and middle-aged adults. Very few studies have focused on automatic emotion recognition for the elderly. Aging comes with a decline in the ability to recognize emotions and impacts emotion perception in humans. Furthermore, older people are suffering from cognitive impairment worldwide, which leads to abnormal emotional patterns. This paper is a literature review of FER techniques in computer vision; FER approaches, and FER databases, and discusses the main challenge of facial expression recognition across age and lifespan.
Current speech recognition systems with fixed vocabularies have difficulties recognizing Out-of-Vocabulary words (OOVs) such as proper nouns and new words. This leads to misunderstandings or even failures in dialog systems. Ensuring effective speech recognition is crucial for the proper functioning of robot assistants. Non-native accents, new vocabulary, and aging voices can cause malfunctions in a speech recognition system. If this task is not executed correctly, the assistant robot will inevitably produce false or random responses. In this paper, we used a statistical approach based on distance algorithms to improve OOV correction. We developed a post-processing algorithm to be combined with a speech recognition model. In this sense, we compared two distance algorithms: Damerau–Levenshtein and Levenshtein distance. We validated the performance of the two distance algorithms in conjunction with five off-the-shelf speech recognition models. Damerau–Levenshtein, as compared to the Levenshtein distance algorithm, succeeded in minimizing the Word Error Rate (WER) when using the MoroccanFrench test set with five speech recognition systems, namely VOSK API, Google API, Wav2vec2.0, SpeechBrain, and Quartznet pre-trained models. Our post-processing method works regardless of the architecture of the speech recognizer, and its results on our MoroccanFrench test set outperformed the five chosen off-the-shelf speech recognizer systems.
The ever-increasing global population presents a looming threat to food production. To meet growing food demands while minimizing negative impacts on water and soil, agricultural practices must be altered. To make informed decisions, decision-makers require timely, accurate, and efficient crop maps. Remote sensing-based crop mapping faces numerous challenges. However, recent years have seen substantial advances in crop mapping through the use of big data, multi-sensor imagery, the democratization of remote sensing data, and the success of deep learning algorithms. This systematic literature review provides an overview of the history and evolution of crop mapping using remote sensing techniques. It also discusses the latest scientific advances in the field of crop mapping, which involve the use of machine and deep learning models. The review protocol involved the analysis of 386 peer-reviewed publications. The results of the analysis show that areas such as crop rotation mapping, double cropping, and early crop mapping require further exploration. The use of LiDAR as a tool for crop mapping also needs more attention, and hierarchical crop mapping is recommended. This review provides a comprehensive framework for future researchers interested in accurate large-scale crop mapping from multi-source image data and machine and deep learning techniques.
Recently, Morocco has started to invest in IoT systems to transform our cities into smart cities that will promote economic growth and make life easier for citizens. One of the most vital addition is intelligent transportation systems which represent the foundation of a smart city. However, the problem often faced in such systems is the recognition of entities, in our case, car and model makes. This paper proposes an approach that identifies makes and models for cars using transfer learning and a workflow that first enhances image quality and quantity by data augmentation and then feeds the newly generated data into a deep learning model with a scaling feature-that is, compound scaling. In addition, we developed a web interface using the FLASK API to make real-time predictions. The results obtained were 80% accuracy, fine-tuning it to an accuracy rate of 90% on unseen data. Our framework is trained on the commonly used Stanford Cars dataset.
Abstract We ask whether Adversariality in left-right stereo images can learn to estimate an optimal depth map through a consensus-based loss function and ego-motion. We describe the workflow as merging supervised learning (AS) and unsupervised learning (U) models. The supervised learning model optimizes a depth estimation network with prior knowledge of ground-truth depth maps. Inspired by the recent works on adversarial neural networks, we formulate the supervised model as an adversarial learning task. Thus we generate two depth maps from left-right stereo images, respectively. Based on their learning behavior–that is, the loss function, we take the then optimal depth map. On the contrary, the unsupervised learning module has no knowledge of the ground-truth depth yet optimizes the depth estimation network using 3D geometry. Our framework is trained and benchmarked on the KITTI driving dataset.
Artificial intelligence-based speech recognition systems are already available and capable of recognizing the French language. Still, it is quite time-consuming to compare which one will be effective for an assistant robot. The study aims to select the best French-language speech recognition system with the least error in a real environment. In this paper, we present related works on how an Automatic Speech Recognition (ASR) system works, the models used by each of its components, several open-source French datasets, and the frequently used evaluation techniques. Next, we compare deep learning-based speech recognition APIs and pre-trained models for French on two different datasets using the Word Error Rate (WER) metric. The experimental results reveal that Google's Speech-to-Text API outperforms the other systems, namely VOSK API, Wav2vec 2.0, QuartzNet, and Speech Brain's Convolutional, Recurrent, and Fully-connected Networks (CRDNN) model.
Remote sensing-based crop mapping has continued to grow in economic importance over the last two decades. Given the ever-increasing rate of population growth and the implications of multiplying global food production, the necessity for timely, accurate, and reliable agricultural data is of the utmost importance. When it comes to ensuring high accuracy in crop maps, spectral similarities between crops represent serious limiting factors. Crops that display similar spectral responses are notorious for being nearly impossible to discriminate using classical multi-spectral imagery analysis. Chief among these crops are soft wheat, durum wheat, oats, and barley. In this paper, we propose a unique multi-input deep learning approach for cereal crop mapping, called "CerealNet". Two time-series used as input, from the Sentinel-2 bands and NDVI (Normalized Difference Vegetation Index), were fed into separate branches of the LSTM-Conv1D (Long Short-Term Memory Convolutional Neural Networks) model to extract the temporal and spectral features necessary for the pixel-based crop mapping. The approach was evaluated using ground-truth data collected in the Gharb region (northwest of Morocco). We noted a categorical accuracy and an F1-score of 95% and 94%, respectively, with minimal confusion between the four cereal classes. CerealNet proved insensitive to sample size, as the least-represented crop, oats, had the highest F1-score. This model was compared with several state-of-the-art crop mapping classifiers and was found to outperform them. The modularity of CerealNet could possibly allow for injecting additional data such as Synthetic Aperture Radar (SAR) bands, especially when optical imagery is not available.
In this paper, we propose a robust real-time vehicle tracking and inter-vehicle distance estimation algorithm based on stereovision. Traffic images are captured by a stereoscopic system installed on the road, and then we detect moving vehicles with the YOLO V3 Deep Neural Network algorithm. Thus, the real-time video goes through an algorithm for stereoscopy-based measurement in order to estimate distance between detected vehicles. However, detecting the real-time objects have always been a challenging task because of occlusion, scale, illumination etc. Thus, many convolutional neural network models based on object detection were developed in recent years. But they cannot be used for real-time object analysis because of slow speed of recognition. The model which is performing excellent currently is the unified object detection model which is You Only Look Once (YOLO). But in our experiment, we have found that despite of having a very good detection precision, YOLO still has some limitations. YOLO processes every image separately even in a continuous video or frames. Because of this much important identification can be lost. So, after the vehicle detection and tracking, inter-vehicle distance estimation is done.
Falls are the most leading cause of accidental injury deaths worldwide, therefore, it poses a real challenge for the prevention of life-threatening conditions in geriatrics. The most damaged community is the ever-growing aging population. For this reason, there are considerable demands to distinguish a dangerous posture such as fall in real-time. Here we provide a literature review of conducted work on elderly fall detection and prediction mentioning the main methods to recognize human posture including computer vision-based and wearable sensor-based. We approached these perspectives: sensor fusion, Datasets, approaches proposition. In conclusion, our survey summarizes the progress achieved in the five past years to help the researchers in this field to spot areas where further effort would be beneficial and innovative.
Visual vehicle tracking is one of the most challenging research topics in computer vision. In this paper, we propose a novel and efficient approach based on the particle filter technique and deep learning for multiple vehicle tracking, where the main focus is to associate vehicles efficiently for online and real-time applications. Experimental results illustrate the effectiveness of the system we are proposing.
In this paper, we present a novel technique to estimate vehicle speed on highway using stereo images. First, traffic images are captured using calibrated and synchronized stereo cameras, then we detect moving vehicles on the left image by subtracting the background image. On each detected vehicle, we extract and match Speed Up Robust Features (SURF) in order to compute sparse depth maps. Finally, we get vehicle speed from vehicle depth variation using some geometric derivations. The experiments shows that the proposed algorithm has a satisfactory estimation of vehicle speed comparing to GPS ground truth with a speed error of 2 Km/h in the Moroccan environment.
In the last decades, Intelligent Transportation System become essential for urban traffic management all over the world. In thi s context, we are working on the first Moroccan intelligent transport system called MoVITS based on image processing and machine learning techniques. The development of an efficient system for Moroccan traffic faces many challenges such as bad driving behavior, the huge volume and the diversity of the traffic. In this paper, we will present a three-layers architecture of the proposed system de-scribing the main functionalities such as vehicle detection, tracking, speed estimation, color and type recognition, license plate recognition and road violation detection.
Several algorithms have already made to estimate vehicle’s speed using single camera. The main problem is the efficiency of these systems: processing is done on the whole image while moving objects occupy only a specific part. This work aims to present a technique of speed estimation based on one pixel’s width line processing. This line image can be extracted from a full size image or from a line camera in order to apply some processing like background subtraction, morphological operations, binarization and finally blob detection to allow tracking. This new approach enables implementation on low-cost platforms with low computing power. The size of data to process can almost be divided by the width’s image of a standard resolution camera, allowing a low data rate from acquisition.
In this paper, we propose an efficient algorithm for sharpness improvement of license plates based on image fusion, and the use of multiple images of the same vehicle in different positions of the road using information redundancy. This algorithm was used on the radar project led by MAScIR that enables to estimate the speed of vehicles, extract the license plates and improve their contrast through fusion. Unlike the flash radars, this system uses a video stream processing and can process more data. Different image processing techniques were used in this project, the canny edge detector and forms detection to localize and extract the plates, the morphological operations and filtering to help the process and all the geometrical transformations for an efficient fusion.
Machine vision algorithms require high-computing power. A high performance parallel system has been proposed in this paper by implementing a road traffic radar video processing chain in real-time on a new embedded architecture. The proposed machine consists of the Digital Signal Processor (DSP) 66AK2H12 from Texas Instruments (TI). The goal of this paper is the estimation of the vehicles number, speeds and classification through an optimal exploitation of the parallel architecture based on DSP and ARM cores and high speed buses used in the video acquisition and processing.
This work is part of developing a new type of radars which is based on stereoscopic effect obtained by using two cameras. The main work is to develop an algorithm for speed estimation. We begin by detecting motions and tracking vehicles in order to identify the vehicle in the next frame. Stereoscopic pictures allow us to calculate the distance from the cameras to the chosen object within the picture. The distance is calculated from differences between the pictures and by using intrinsic and extrinsic cameras' parameters. The object is selected on the left picture, while the same object on the right picture is automatically detected by calculating the cross-correlation's score between both pictures. The object's position can be calculated by doing some geometrical derivations. The speed is estimated by calculating the slope of the distances estimated in several frames. The accuracy of the position depends on picture resolution, optical distortions and distance between the cameras.