Multiple object tracking (MOT) consists of following the trajectories of different objects in a video with either fixed or moving background. In recent years, the use of deep learning for MOT in the videos recorded by unmanned aerial vehicles (UAVs) has introduced more challenges and hence has a lot of room for extensive research. For the tracking-by-detection method, the three main components, object detector, tracker and data associator, play an equally important role and each part should be tuned to the highest efficiency to increase the overall performance. In this paper, the parameter selection of the Kalman filter and Hungarian algorithm for sheep tracking in paddock videos is discussed. An experimental comparison is presented to show that if the detector is already providing good results, a small change in the system can degrade or improve the tracking capabilities of remaining components. The encouraging results provide an important step in an automated UAV-based sheep tracking system.
Farms in many countries use fenced pastures—known as paddocks—for keeping livestock. These farms usually spread over many hectares with paddocks of various shapes and sizes. Farm managers face the difficulty of counting their animals on regular basis and recently the use of an unmanned aerial vehicle (UAV) has been proposed for counting livestock. To get the data from only a respective paddock, the fence lines must also be detected correctly in the aerial images and videos. This will help to keep the data from each paddock separate. A unique system based on a combination of deep learning and machine learning is proposed to accurately detect the fence lines in the videos recorded by a UAV at an altitude of 80 m.
In the last decade, researchers have focused more on deep convolutional neural networks (CNNs) than other machine learning algorithms for object detection, localization, classification and segmentation. Such CNNs have achieved remarkable results in these fields and use the bounding boxes as the ground truth data. In this research article, we have used a fully connected network (FCN) for livestock detection in aerial images captured by an unmanned aerial vehicle (UAV), that used centroids as ground truth data. For performance evaluation and comparison, we have proposed a single-layered and a seven-layered CNN network in this article. These proposed networks are trained using state-of-the-art method, Region-based CNN. In addition, AlexNet, GoogLeNet, VGG16, VGG19 and ResNet50 were also fine-tuned for livestock detection. The results of the FCN and one of our proposed networks are then merged to improve the recall of the complete system from 90% to 98%.
Conventional livestock counting methods are tedious, time-consuming, and labor-intensive for farmers, which makes counting an irregular task. This inability to constantly monitor stock numbers gives farm rustlers sufficient time to steal farm animals and hence leads to a significant financial loss annually. To overcome this issue, research using unmanned aerial vehicles (UAVs) in pastoral farming is growing progressively since the last decade, but their use is still limited to some extent. This research article gives a detailed analysis of the existing hardware-based and software-based methods. The impact of shifting from current methods to a UAV-based system is discussed using the findings of research articles, and algorithms for object detection in images and tracking in videos are also analyzed. The article concludes that there are still unexplored practical uses of UAV in pastoral farming for monitoring, counting, and tracking farm animals, especially in countries like New Zealand, where pastoral products cover a major portion of export revenue.
In this work we consider the task of detecting sheep onboard an unmanned aerial vehicle (UAV) flying at an altitude of 80 m. At this height, the sheep are relatively small, only about 15 pixels across. Although deep learning strategies have gained enormous popularity in the last decade and are now extensively used for object detection in many fields, state-of-the-art detectors perform poorly in the case of smaller objects. We develop a novel dataset of UAV imagery of sheep and consider a variety of object detectors to determine which is the most suitable for our task in terms of both accuracy and speed. Our findings indicate that a UNet detector using the weighted Hausdorff distance as a loss function during training is an excellent option for detection of sheep onboard a UAV.
The memory requirements of digital signal processing and multimedia applications have grown steadily over the last several decades. From embedded systems to supercomputers, the design of computing platforms involves a balance between processing elements and memory sizes to avoid the memory wall. This paper presents an algorithm based on both dataflow and approximate computing approaches in order to find a good balance between the memory requirements of an application and the quality of the result. The designer of the computing system can use these evaluations early in the design process to make hardware and software design decisions. The proposed method does not require any modification in the algorithm's computations, but optimises how data are fetched from and written to memory. We show in this paper how the proposed algorithm saves 27.7% of memory for the full SKA SDP signal processing computing pipeline, and up to 68.75% for a wavelet transform in embedded systems.
Stixel calculations are commonly based on binocular vision; these calculations map millions of pixel disparities into a few hundred stixels. Depending on applied stereo vision, this binocular approach is sometimes incapable of dealing with low-textured road information or noisy data. The main objective of this work is to propose a more reliable approach to calculating stixels by incorporating laser scanners (i.e., LIDAR). We show that this supports more efficient and robust 3D point representations, even if only integrating monocular vision into the LIDAR-based approach for generating monocular stixels. Experimental results show a more accurate (by 15.4 %) stixel detection rate when the LIDAR-guided monocular configuration is used compared to a conventional binocular approach.
Counting livestock is generally done only during major events, such as drenching, shearing or loading, and thus farmers get stock numbers sporadically throughout the year. More accurate and timely stock information would enable farmers to manage their herds better. Additionally, prompt response to any stock in distress is extremely valuable, both in terms of animal welfare and the avoidance of financial loss. In this regard, the evolution of deep learning algorithms and Unmanned Aerial Vehicles (UAVs) is forging a new research area for remote monitoring and counting of different animal species under various climatic conditions. In this paper, we focus on detecting and counting sheep in a paddock from UAV video. Sheep are counted using a model based on Region-based Convolutional Neural Networks and the results are then compared with other techniques to evaluate their performance.
Compression of electroencephalogram (EEG) signals has long been a challenging research topic. Recently, compressed sensing (CS) was proposed for EEG acquisition and compression with the rapid development of wearable health-care monitoring systems. Building an optimal dictionary to achieve excellent reconstruction accuracy and low computational complexity is extremely desirable in this context. While most existing work has focused on static dictionaries such as Gabor, Fourier and wavelets, the dynamic nature of EEG signals motivates us to study learned dictionaries. In this paper, we study how to build K-SVD learned dictionaries and then adopt them in EEG compression with CS. For the convenience of comparison, the well-established database of scalp EEG signals from Physiobank is used in the numerical experiments. Results demonstrate that a K-SVD learned dictionary provides high reconstruction accuracy with a short computational time.
Recently, wireless acoustic sensor networks (WASNs) have received significant attention from the research community and a variety of methods have been proposed for numerous applications, such as location estimation and speech enhancement. The lack of publicly available datasets with signals recorded in WASNs, presents difficulties in obtaining consistent performance indicators across the different approaches. In this paper, we present and release a dataset of real recorded signals in an outdoor WASN comprised of four microphone arrays. Our dataset consists of several speakers recorded at various locations within the WASN and can be used for benchmarking purposes. We also present location estimation results using our real recorded dataset. Our results can serve as a baseline indicator of localization performance of single and multiple sources in a real environment.
The Square Kilometre Array (SKA) will push the boundaries of radio astronomy. As such, it will need enormous computing power to process the tremendous amount of data it will produce. Significant savings in computing could be made if some of this processing was done in single-precision. This paper presents our end-to-end modelling of the SKA Imaging Pipeline. We model the signals received by the Central Signal Processor (CSP) using our Sky Generator Model. We then process the sky signals with our CSP Correlator Model to generate the visibilities that are passed to the Science Data Processor (SDP). These visibilities are gridded Fourier transformed by our SDP Imaging Model to produce a dirty image. This dirty image is then deconvolved to produce an image of the sky. Through our models we investigate the error introduced with reduced numerical precision, and we perform various tests to explore the limits of single-precision processing.
Electroencephalogram (EEG) signals have been widely used to analyze brain activities so as to diagnose certain brain-related diseases. They are usually recorded for a fairly long interval with adequate resolution, consequently requiring a considerable amount of memory space for storage and transmission. Recently compressed sensing (CS) has been proposed in order to effectively compress EEG signals. However, its performance is closely dependent on how a compression dictionary is built. Through our study, we notice that building the best fit over-complete Gabor dictionary plays an important role in this task. In this paper, we evaluate the effect of different time and frequency step sizes in building Gabor atoms on the performance of EEG signal compression using CS with three common EEG databases used by the research community. Taking the Normalized Mean Square Error (NMSE) as a performance metric, we present a quantitative study with an attempt to provide more insight on how to adopt CS in EEG signal compression.
Aging populations are stretching existing healthcare systems to their limits in both developing and developed countries. Telemedicine is a promising solution to this challenging problem. Under the conventional data compression paradigm, long-time recording of electroencephalography (EEG) signals still generates excessive amount of data, which requires large data storage and long transmission time. While promoting mobile telemedicine with compressed sensing (CS) as a key system for EEG monitoring, this paper investigates the effect of epoch length on CS to compress EEG signals. Experimental results show that a longer epoch length leads to better signal compression at the expense of larger signal reconstruction time. At a sampling frequency of 256 Hz, a 4-s epoch length is suitable when using a general desktop computer to perform signal reconstruction.
With the fast development of wearable healthcare systems, compressed sensing (CS) has been proposed to be applied in electroencephalogram (EEG) acquisition. For CS, it is desired to build the best-fit dictionary in order to achieve good reconstruction accuracy. While most of existing works focused on static dictionaries such as Gabor, Fourier and wavelets. The dynamic nature of EEG signals motivates us to study learned dictionaries, which are supposed to provide better reconstruction accuracy and lower computation cost. In this paper, we provide the quantitative performance comparison of EEG CS using two different types of dictionaries, i.e., the well-known Gabor dictionaries versus K-SVD learned dictionaries. The performance comparison utilizes the well-established database of scalp EEG from Physiobank, which allows researchers in this field to compare their work with ours. In addition, it also attempts to inspire the systematic study of dictionary learning in EEG CS.
In this paper we present our end-to-end model of the imaging pipeline in the Square Kilometre Array. Our Sky Generator models the signals that are received by the Central Signal Processor (CSP), our CSP Correlator model then processes those signals to generate visibilities to pass to the Science Data Processor (SDP). Our SDP Imaging model then grids the visibilities and inverse Fourier transforms them to produce a dirty image of the sky. Our modelling allows us to investigate the error that is introduced due to reduced numerical precision, and we then propose techniques to mitigate this error, and thus reduce the required amount of computational hardware.
In this paper we present our end-to-end model of the imaging pipeline in the Square Kilometre Array. Our Sky Generator models the signals that are received by the Central Signal Processor (CSP), our CSP Correlator model then processes those signals to generate visibilities to pass to the Science Data Processor (SDP). Our SDP Imaging model then grids the visibilities and inverse Fourier transforms them to produce a dirty image of the sky. Our modelling allows us to investigate the error that is introduced due to reduced numerical precision, and we then propose techniques to mitigate this error, and thus reduce the required amount of computational hardware.
In this work we present a vision-based road user monitoring system for traffic intersections using a combination of Gaussian Mixture Model (GMM)-based deep learning approaches and geometric warping for further behaviour analysis. GMMs and detection of features in consecutive frames are used in bounding box prediction in order to track objects, while a Fast Regional Convolutional Neural Networks (R-CNN) approach classifies different road users. For better analysis, by knowing the actual coordinates of the intersection, we use geometric warping to map the 3D-plane to the 2D plane. We thus extract the real distance and the approximate real time speed of each road user. The results demonstrate that the proposed region proposal generator method for Fast R-CNN outperforms the other methods considered in terms of both accuracy and computation time.
Significant improvements in intelligibility of speech in noise can be obtained by modifying the speech signal in the time and/or frequency domains. However, most speech intelligibility enhancement algorithms are designed to use clean speech as an input, and their performance suffers once the input speech signal-to-noise ratio decreases, a common case in face-to-face communication environments such as restaurants or cafés. In this work we investigate whether a particularly successful speech intelligibility enhancement system-spectral shaping and dynamic range compression-and various front-end noise reduction methods might be suitable in such environments. Our evaluations suggest that such a complete system would provide an increase in speech intelligibility equivalent to a gain of 10 dB input signal-to-noise ratio in the more challenging face-to-face communication environments.
Desmond Taylor合作论文数Department of Electrical and Computer Engineering|University of Canterbury2
Yannick Deville合作论文数Signal, Image & Instrumentation (S2I)
Laboratoire d’Astrophysique de Toulouse-Tarbes1