The identification of biological echoes in radar data has revolutionized research into airborne migratory species. Deep learning applied to polarimetric weather radar observations can reveal signature patterns mass movement by bio-scatterers such as birds, bats, and insects. However, due to the difficulties in labelling bio-scatterers in these data, threshold approaches have been proposed in the literature. In this research, we used the depolarization ratio (DR) based on differential reflectivity (zDR) and the cross-correlation coefficient (pHV), along with citizen scientist-reported data, to label bio-scatterers for deep learning. This method labelling biological echoes in radar signatures is prone to noise, which impacts the accuracy of any model that relies on it. We introduce a novel semi-supervised co-training approach that uses a bootstrap ensemble with a confidence threshold. Our ensemble consists of the newly proposed STNet and two modified FNet models, which incorporate co-learning through bootstrap sampling for label correction. This innovative method significantly improves classification accuracy across all three multivariate numerical datasets compared baseline models that lack co-learning with bootstrap-based label correction.
This paper aims to explore the effectiveness of machine learning models for the classification of breast cancer images trained with data from a demographic region and tested or used in a different demographic region. Artificial intelligence has the potential to be integrated into the imaging process to reduce workload and broaden screening audiences. However, artificial intelligence has sometimes been shown to demonstrate bias in medical applications. Most available breast cancer images are collected in North American, European, or East Asian countries, and there is limited data available from other regions. Bias between demographics could lead to some groups being under-diagnosed, resulting in worsened prognoses. In this work, a high-performance breast cancer classification model with AUC of up to 0.7415 on ultrasound images and 0.8920 on mammogram images has been developed. The model is trained and tested on a variety of datasets, some specifically collected in Ghana to compare with publicly available datasets from the UK, Portugal, and Poland. Experiments were conducted to determined any bias between different demographic regions. A significant decrease in performance was found in five of the six experiments conducted.
Tracking wild animals through videos presents a non-intrusive and cost-effective way of gathering scientific information key for conservation. State-of-the-art research has shown convolutional neural networks to be highly accurate and generalisable to a plethora of problems. However, the application of this field on wild animal tracking has had relatively little interest. This is potentially due to the challenges of varying illumination, noisy backgrounds and camouflaged animals intrinsic to the problem. The aim of this work is to explore and apply state-of-the-art research to detect and track wild animals (specifically bears and primates, including their body parts) in video sequences in real-time. Due to obstructors such as foliage being prevalent in wild animal environments, body part tracking presents a solution to detecting animals when they are obstructed. Two deep convolutional neural networks are trained to detect and track animals in their natural habitat. By using the knowledge that an animal is composed of body parts, the score of weakly predicted bounding is boosted from the relative distance of related body parts. For tracking, the K-Means algorithm is used to locate the average position of each animal in frame. Using the temporality of the video, directional arrows are assigned to each position of the animals to show their relative movements in the frame. The solution presented here demonstrates an ability to detect animals in real-time (approximately 23 frames per second) with high precision. With the introduction of a body-part confidence boosting, the detection rate can be increased by approximately 2% for a weakly predicted class.
Breast cancer is the second most prevalent form of cancer and is the “leading cause of most cancer-related deaths in women”. Most women living in low- and middle-income countries (LMIC) have limited access to the existing poor health systems, restricted access to treatment facilities, and in general lack of breast cancer screening programmes. The likelihood of women living in LMIC attending a health facility with advanced-stage breast cancer is very high and the chances of them being able to afford treatment at that stage, even if the treatment is available, is very low. In this work, we evaluate the capabilities of deep learning as a classification tool with the aim of detecting cancerous ultrasound breast images. We aim to deploy a simple classifier on a mobile device with an inexpensive handheld ultrasound imaging system to pick up breast cancer cases that will need medical attention. We demonstrate in this work that with minimal ultrasound images, a de novo system trained from scratch can achieve accuracy of close to 64% and about 78% when the same model is pre-trained.
Despite tremendous advancement in computer vision, especially with deep learning, understanding scenes in the wild remains challenging. Even modern image classification models often misclassify when presented with out-of-distribution inputs despite having been trained on tens of millions of images or more. Moreover, training modern deep-learning classifiers requires a lot of energy due to the need to iterate many times over the training set, constantly updating billions of model parameters. Owing to problems with generalisability and robustness as well as efficiency, there is growing interest in computer vision to mimic biological vision (e.g., human vision) in the hope that doing so will require fewer resources for training both in terms of energy and in terms of data sets while increasing robustness and generalisability. This paper proposes a biologically plausible neuromorphic vision system that is based on a spiking neural network and is evaluated on the classification of hand-written digits from the MNIST dataset. The experimental outcome indicates improved robustness of the proposed approach over state-of-the-art considering non-digit detection.
Advances in machine learning coupled with the abundances of training data has facilitated the deep learning era, which has demonstrated its ability and effectiveness in solving complex detection and recognition problems. In general application areas with elements of machine learning have seen exponential growth with promising new and sophisticated solutions to complex learning problems. In computer vision, the challenge related to the detection of known objects in a scene is a thing of the past. With the tremendous increase in detection accuracies, some close to that of human detection, there are several areas still lagging in computer vision and machine learning where improvements may call for more architectural designs. In this paper, we propose a physiologically inspired model for scene understanding that encodes three key components: object location, size and category. Our aim is to develop an energy efficient artificial intelligent model for naturalistic scene understanding capable of deploying on a low power neuromorphic hardware. We have reviewed recent advances in deep learning architecture that have taken inspiration from human or primate learning systems and provided direct to future advancement on deep learning with inspiration from physiological experiments. Upon a review of areas that have benefitted from deep learning, we provide recommendations for enhancing those areas that might have stalled or grinded to a halt with little or no significant improvement.
Humans recall the past by replaying fragments of events temporally. Here, we demonstrate a similar effect in macaques. We trained six rhesus monkeys with a temporal-order judgement (TOJ) task and collected 5000 TOJ trials. In each trial, the monkeys watched a naturalistic video of about 10 s comprising two across-context clips, and after a 2 s delay, performed TOJ between two frames from the video. The data are suggestive of a non-linear, time-compressed forward memory replay mechanism in the macaque. In contrast with humans, such compression of replay is, however, not sophisticated enough to allow these monkeys to skip over irrelevant information by compressing the encoded video globally. We also reveal that the monkeys detect event contextual boundaries, and that such detection facilitates recall by increasing the rate of information accumulation. Demonstration of a time-compressed, forward replay-like pattern in the macaque provides insights into the evolution of episodic memory in our lineage.
This paper argues that energy consideration should be central to software development. It speculates that including the notion of energy awareness in programming language design for domain specific languages (DSLs) is a novel way in which energy-aware and energy-efficient applications can be developed. It outlines the design criteria and rationale for using a language-focused approach for energy-awareness. It proposes Lantern, a DSL for supporting energy awareness in Cyber-Physical Systems software development. Lantern allows the development of applications that better manage and reduce the carbon footprint of devices. The design of Lantern is aimed at supporting the general development of Cyber-Physical Systems. This paper focuses on the scenario of smart homes, using statically defined locations within a specified environment.
Humans can perform episodic memory replay dynamically by replaying fragments of fine-grained temporal patterns of events and skipping flexibly across subevents. Here, we demonstrate a similar effect in macaque monkeys. We report evidence that the monkeys apply a time-compressed, forward replay mechanism during their judgement the order of cinematic events. We trained six macaque monkeys with a temporal order judgement (TOJ) task and collected 5000 TOJ trials with naturalistic videos. In each trial, they watched a video of about 10 s comprising two across-context clips, and after a 2-s retention delay, performed a temporal order judgement between two frames extracted from the video. The results show that the monkeys adopt a self-terminating search of ordered frames in their mnemonic representation. The memory replay is forward and temporally compressed, paralleling evidence in humans. Such compression of replay is however not sophisticated enough to allow them to skip over irrelevant information by compressing the encoded video globally. Moreover, we also reveal that the monkeys can segment events using contextual boundaries like humans and such contextual segmentation facilitates their memory recall by an increased rate of information accumulation in a drift diffusion model framework. Memory replay is an elaborate mental process and our demonstration of a time-compressed, forward replay mechanism in the macaque monkeys provides insights into mapping the mechanisms and evolution of episodic memory in our lineage.
Wearable inertial measurement units incorporating accelerometers and gyroscopes are increasingly used for activity analysis and recognition. In this paper an activity classification algorithm is presented which includes a novel multi- step refinement with the aim of improving the classification accuracy of traditional approaches. To do so, after the classification takes place, information is extracted from the confusion matrix to focus the computational efforts on those activities with worse classification performance. It is argued that activities differ diversely from each other, therefore a specific set of features may be informative to classify a specific set of activities, but such informativeness should not necessarily be extended to a different activity set. This approach has shown promising results, achieving important classification accuracy improvements.
The adoption of high-accuracy speech recognition algorithms without an effective evaluation of their impact on the target computational resource is impractical for mobile and embedded systems. In this paper, techniques are adopted to minimise the required computational resource for an effective mobile-based speech recognition system. A Dynamic Multi-Layer Perceptron speech recognition technique, capable of running in real time on a state-of-the-art mobile device, has been introduced. Even though a conventional hidden Markov model when applied to the same dataset slightly outperformed our approach, its processing time is much higher. The Dynamic Multi-layer Perceptron presented here has an accuracy level of 96.94% and runs significantly faster than similar techniques.
The monitoring of bird populations can provide important information on the state of sensitive ecosystems; however, the manual collection of reliable population data is labour-intensive, time-consuming, and potentially error prone. Automated monitoring using computer vision is therefore an attractive proposition, which could facilitate the collection of detailed data on a much larger scale than is currently possible. A number of existing algorithms are able to classify bird species from individual high quality detailed images often using manual inputs (such as a priori parts labelling). However, deployment in the field necessitates fully automated in-flight classification, which remains an open challenge due to poor image quality, high and rapid variation in pose, and similar appearance of some species. We address this as a fine-grained classification problem, and have collected a video dataset of thirteen bird classes (ten species and another with three colour variants) for training and evaluation. We present our proposed algorithm, which selects effective features from a large pool of appearance and motion features. We compare our method to others which use appearance features only, including image classification using state-of-the-art Deep Convolutional Neural Networks (CNNs). Using our algorithm we achieved an 90% correct classification rate, and we also show that using effectively selected motion and appearance features together can produce results which outperform state-of-the-art single image classifiers. We also show that the most significant motion features improve correct classification rates by 7% compared to using appearance features alone.
An enduring puzzle in the neuroscience of memory is how the brain parsimoniously situates past events by their order in relation to time. By combining functional MRI, and representational similarity analysis, we reveal a multivoxel representation of time intervals separating pairs of episodic event-moments in the posterior medial memory system, especially when the events were experienced within a similar temporal context. We further show such multivoxel representations to be vulnerable to disruption through targeted repetitive transcranial magnetic stimulation and that perturbation to the mnemonic abstraction alters the neural—behavior relationship across the wider parietal memory network. Our findings establish a mnemonic “pattern-based” code of temporal distances in the human brain, a fundamental neural mechanism for supporting the temporal structure of past events, assigning the precuneus as a locus of flexibly effecting the manipulation of physical time during episodic memory retrieval.
Falls are one of the greatest risks for older adults living alone at home. This paper presents a novel visual-based fall detection approach to support independent living for older adults through analysing the motion and shape of the human body. The proposed approach employs a new set of features to detect a fall. Motion information of a segmented silhouette when extracted can provide a useful cue for classifying different behaviours, while variation in shape and the projection histogram can be used to describe human body postures and subsequent fall events. The proposed approach presented here extracts motion information using best-fit approximated ellipse and bounding box around the human body, produces projection histograms and determines the head position over time, to generate 10 features to identify falls. These features are fed into a multilayer perceptron neural network for fall classification. Experimental results show the reliability of the proposed approach with a high fall detection rate of 99.60% and a low false alarm rate of 2.62% when tested with the UR Fall Detection dataset. Comparisons with state of the art fall detection techniques show the robustness of the proposed approach.
Over the past two decades, the use of low power Field Programmable Gate Arrays (FPGA) for the acceleration of various vision systems mainly on embedded devices have become widespread. The reconfigurable and parallel nature of the FPGA opens up new opportunities to speed-up computationally intensive vision and neural algorithms on embedded and portable devices. This paper presents a comprehensive review of embedded vision algorithms and applications over the past decade. The review will discuss vision based systems and approaches, and how they have been implemented on embedded devices. Topics covered include image acquisition, preprocessing, object detection and tracking, recognition as well as high-level classification. This is followed by an outline of the advantages and disadvantages of the various embedded implementations. Finally, an overview of the challenges in the field and future research trends are presented. This review is expected to serve as a tutorial and reference source for embedded computer vision systems.
Falls are one of the greatest risks for the older adults living alone at home. This paper presents a novel visual-based fall detection approach to support independent living for older adults. The proposed approach employs three unique features; motion information, human shape variation and projection histogram to detect a fall. Motion information of a segmented silhouette, which when extracted can provide a useful cue for classifying different behaviours. Also, the projection histogram and variation in human shape can be used to describe human body postures and subsequently fall events. The proposed approach presented here extracts motion information, using best-fit approximated ellipse around the human body and in addition projection histogram features to further improve the accuracy of fall detection. Experimental results are presented and show high fall detection rate of 99.81% with partially occluded video data.
For a considerable time, it has been the goal of computational neuroscientists to understand biological nervous systems. However, the vast complexity of such systems has made it very difficult to fully understand even basic functions such as movement. Because of its small neuron count, the C. elegans nematode offers the opportunity to study a fully described connectome and attempt to link neural network activity to behaviour. In this paper a simulation of the neural network in C. elegans that responds to chemical stimulus is presented and a consequent realistic head movement demonstrated. An evolutionary algorithm (EA) has been utilised to search for estimates of the values of the synaptic conductances and also to determine whether each synapse is excitatory or inhibitory in nature. The chemotaxis neural network was designed and implemented, using the parameterisation obtained with the EA, on the Si elegans platform a state-of-the-art hardware emulation platform specially designed to emulate the C. elegans nematode.