Today, automotive technology is highly dependent on both internal and external communications, making it vulnerable to new attack vectors. To enhance communication security, researchers have proposed anomaly detection methods using advanced sequence modeling with transformers. Although these architectures achieve impressive accuracy in the upper 90 % range, their performance beyond limited testing scenarios is unclear. To overcome this obstacle, we propose to study the limits of the model using adversarial samples. The most effective method of generating adversarial samples uses gradient descent. Since transformers use discrete tokenized inputs, gradient descent cannot be applied directly. In order to overcome this challenge, we present the Gradient-based Adversarial Perturbation on CAN (GAP-CAN) framework, which leverages space relaxation techniques to apply gradient-based optimization. The GAPCAN framework is comprised of the following components: 1) gradient-based search for adversarial samples and 2) adversarial example generation for existing datasets. The benefit of this approach is the ability to generate multi-token adversarial perturbations in a tractable amount of time. We tested the GAPCAN framework on two common publicly available datasets. Our experimental results show that the GAP-CAN framework successfully finds an adversarial sample within the first 100 queries 63% of the time for a higher entropy dataset and 5% of the time for a lower entropy dataset.
As Deep Learning (DL) algorithms become more widely adopted in healthcare applications, there is a greater emphasis on understanding and addressing the potential privacy risks associated with these models. The purpose of this study is to investigate the privacy vulnerabilities of the Convolutional Neural Network (CNN) classifiers for Electroencephalogram (EEG) data in the Brain-Computer Interfaces (BCIs). Specifically, it focuses on the Membership Inference Attack (MIA), which seeks to determine if data from an individual were used in model training. The novelty of this work lies in its empirical analysis of MIA, by addressing two key challenges that are less common in other domains: 1) heterogeneous datasets and 2) spatio-temporal design choices. Motivated by these challenges, we investigate the susceptibility to MIA based on: 1) the specifics of the training data set (number of participants, demographics), and 2) specifics of the CNN (such as architecture, regularization). Our experiments revealed that an adversary with limited knowledge of the model and its training process can compromise the privacy of training participants, noting that the same attack is not effective against deep learning models trained on image and tabular datasets. Some of our findings are: 1) training on diverse participant datasets improves the privacy of most participants but increases risks of memorization and vulnerabilities for underrepresented groups; 2) regularization is less effective in defending against the MIA on EEG data CNN classifiers when compared to other types of input data; 3) the depth and width of the model architecture have no impact on the effectiveness of membership attack. We hope that the insights presented will help future researchers develop more privacy-aware deep learning-based BCI systems.
Group activity recognition in video is a complex task due to the need for a model to recognise the actions of all individuals in the video and their complex interactions. Recent studies propose that optimal performance is achieved by individually tracking each person and subsequently inputting the sequence of poses or cropped images/optical flow into a model. This helps the model to recognise what actions each person is performing before they are merged to arrive at the group action class. However, all previous models are highly reliant on high quality tracking and have only been evaluated using ground truth tracking information. In practice it is almost impossible to achieve highly reliable tracking information for all individuals in a group activity video. We introduce an innovative deep learning-based group activity recognition approach called Rendered Pose based Group Activity Recognition System (RePGARS) which is designed to be tolerant of unreliable tracking and pose information. Experimental results confirm that RePGARS outperforms all existing group activity recognition algorithms tested which do not use ground truth detection and tracking information.
With the many advancements in automobile technology, there has been a sharp increase in the number of sensors and systems present in vehicles. This has enabled a rapid increase in the features and capabilities available in modern automobiles but has also vastly increased their vulnerability surface. Attackers are now able to remotely attack and control some facets of modern automobiles, which creates dangerous situations for drivers and passengers. In order to address this, this paper proposes a self-attention bottleneck network utilizing an encoder-decoder architecture. This is used as an anomaly detection system (ADS) that can detect anomalous behavior present within CAN bus communications. We evaluated this approach using a publicly available CAN bus car hacking dataset and show that our architecture is able to achieve an accuracy of over 99% for detecting anomalies present in CAN bus data.
Numerous state-of-the-art solutions for neural speech decoding and synthesis incorporate deep learning into the processing pipeline. These models are typically opaque and can require significant computational resources for training and execution. A deep learning architecture is presented that learns input bandpass filters that capture task-relevant spectral features directly from data. Incorporating such explainable feature extraction into the model furthers the goal of creating end-to-end architectures that enable automated subject-specific parameter tuning while yielding an interpretable result. The model is implemented using intracranial brain data collected during a speech task. Using raw, unprocessed timesamples, the model detects the presence of speech at every timesample in a causal manner, suitable for online application. Model performance is comparable or superior to existing approaches that require substantial signal preprocessing and the learned frequency bands were found to converge to ranges that are supported by previous studies.
Neuroprosthetics have demonstrated the potential to decode speech from intracranial brain signals, and hold promise for one day returning the ability to speak to those who have lost it. However, data in this domain is scarce, highly variable, and costly to label for supervised modeling. In order to address these constraints, we present brain2vec, a transformer-based approach for learning feature representations from intracranial electroencephalogram data. Brain2vec combines a self-supervised learning methodology, neuroanatomical positional embeddings, and the contextual representations of transformers to achieve three novelties: (1) learning from unlabeled intracranial brain signals, (2) learning from multiple participants simultaneously, all while (3) utilizing only raw unprocessed data. To assess our approach, we use a leave-one-participant-out validation procedure to separate brain2vec’s feature learning from the holdout participant’s speech-related supervised classification tasks. With only two linear layers, we achieve 90% accuracy on a canonical speech detection task, 42% accuracy on a more challenging 4-class speech-related behavior recognition, and 53% accuracy when applied to a 10-class, few-shot word classification task. Combined with the visualizations of unsupervised class separation in the learned features, our results evidence brain2vec’s ability to learn highly generalized representations of neural activity without the need for labels or consistent sensor location.
We introduce a novel deep learning based group activity recognition approach called the Pose Only Group Activity Recognition System (POGARS), designed to use only tracked poses of people to predict the performed group activity. In contrast to existing approaches for group activity recognition, POGARS uses 1D CNNs to learn spatiotemporal dynamics of individuals involved in a group activity and forgo learning features from pixel data. The proposed model uses a spatial and temporal attention mechanism to infer person-wise importance and multi-task learning for simultaneously performing group and individual action classification. Experimental results confirm that POGARS achieves highly competitive results compared to state-of-the-art methods on a widely used public volleyball dataset despite only using tracked pose as input. Further our experiments show by using pose only as input, POGARS has better generalization capabilities compared to methods that use RGB as input.
We present a perspective of the national transplant program based on organizational theory and complexity theory, framing the system’s allocation of donor organs as an interorganizational directed multiplex of agents with diverse belief formation in a cooperative-competitive environment. Simulation and analysis of this macroscale complexity may help explain known behavioural variations across member organizations. However, the transplant community still relies on system-scale simulations since effective macroscale methodologies are not well established. Therefore, we offer this perspective of the national transplant program as a means to stimulate new methods that capture macroscale impacts of policy development for deceased donor organ allocation.
The practice of deceased-donor solid-organ transplantation is unique among clinical therapies because it substantially benefits recipients, The allocation process is a complex System-of-Systems (SoS) with multiple stakeholders, components, rules, and policies. While we model and simulate such SoS, they are challenging to verify and validate thoroughly due to a large number of use cases and the sensitivity to initial conditions. Furthermore, subject matter experts within communities of practice do not trust such models because they are difficult to explain. This paper proposes a trust-centric approach to validation focusing on experimentally exploring the simulation's behavior space and demonstrating its utility to a community of experts. We describe a three-step validation process to show how modelling and simulation professionals can foster trust in simulations of a complex SoS. We apply the framework to verify and validate a simulation model of the Kidney-Pancreas (KP) allocation process within the United States.
This study aimed to investigate the feasibility of using wearable tracking data to train a player detection and moving camera calibration model for Australian Football. Player tracking data were collected from multiple matches of professional Australian Football using wearable local positioning system (LPS) devices sampling at 10 Hz. Each match was also filmed from two angles with moving high-definition cameras. Image registration using the Speeded-Up Robust Features (SURF) keypoint detector and descriptor was performed to calibrate the moving camera footage and a pre-trained object detector was used to detect players in each frame. These methods required no manual annotation of the video footage and resulted in an incomplete player tracking data set. This initial data set was cross referenced against LPS player locations to identify the subset of successfully calibrated frames. This subset was used to train a deep learning based landmark detection model to track known markings on the field, and to fine tune the object detection model, by using the LPS tracking data to annotate the frames. The pre-trained object detection model had low precision and recall (<50
Recent advances in deep learning approaches to computer vision problems have led to renewed interest in the task of predicting 3D human joint locations from raw image data, with application areas including sports analysis, human-computer interaction, and physical rehabilitation. Although supervised learning of deep neural networks has proven to be effective for pose estimation, it requires a wealth of varied data to generalise well to previously unseen examples. Consequently, progress in 3D human pose estimation has been slowed by the fact that 3D keypoint annotations are notoriously difficult to obtain, traditionally requiring a large array of cameras and/or the use of wearable markers/sensors. In this paper we describe a methodology for obtaining 3D human pose annotations using only three video cameras and without any wearables. We apply this methodology to construct ASPset-510 (Australian Sports Pose Dataset), a large collection of natural sports-related video with 3D pose annotations. Using ASPset-510 as an additional source of training examples, we found that we could improve pose model generalisation on the established MPI-INF-3DHP benchmark. We make ASPset-510 publicly available, and provide strong baseline results for future work to compare against.
Classification of human activities from wearable sensor data is challenged by inter-subject variance and resource-constrained platforms.We address these issues with SincEMG, a deep neural network that exploits digital signal processing concepts and transfer learning to reduce model size for activity recognition on raw sensor data.The model's first layer decomposes signals into frequency bands using finite impulse response filters optimized directly from the data.The subsequent convolutional layers downsample across time and aggregate the first layer's band data.Batch normalization and dropout help to regularize intermediate layer outputs.This approach reduces compute requirements by decreasing the number of learned parameters and eliminating any significant data pre-processing.In addition to these improvements, the model's first layer learns a set of bandpass filters, which provide insight into predictive regions of the source spectrum.We evaluate SincEMG using two publicly available surface electromyography datasets.Our model uses far fewer parameters and achieves state-of-the-art results with 98.53% accuracy for 7-classes and 68.45% accuracy for 18-classes.
It is very important for swimming coaches to analyse a swimmer's performance at the end of each race, since the analysis can then be used to change strategies for the next round. Coaches rely heavily on statistics, such as stroke length and instantaneous velocity, when analysing performance. These statistics are usually derived from time-consuming manual video annotations. To automatically obtain the required statistics from swimming videos, we need to solve the following four challenging computer vision tasks: swimmer head detection; tracking; stroke detection; and camera calibration. We collectively solve these problems using a two-phased deep learning approach, we call Deep Detector for Actions and Swimmer Heads (DeepDASH). DeepDASH achieves a 20.8% higher F1 score for swimmer head detection and operates 6 times faster than the popular Faster R-CNN object detector. We also propose a hierarchical tracking algorithm based on the existing SORT algorithm which we call HISORT. HISORT produces significantly longer tracks than SORT by preserving swimmer identities for longer periods of time. Finally, DeepDASH achieves an overall F1 score of 97.5% for stroke detection across all four swimming stroke styles.
Multiple object tracking is an important but challenging computer vision problem. The complex motion of objects makes tracking difficult during long periods of object occlusion, and as a result occlusions frequently cause fragmented tracks with gaps. Previous works use linear interpolation to fill in such gaps, a technique which is only able to model simple motion. As a result, tracked bounding box locations can be quite poor in these situations. In this paper, we propose a 1D CNN based solution to filling gaps which models complex motion in a data-driven way. Our proposed solution uses only bounding box coordinates as input, and as such does not incur the computational cost of processing image features directly. We show that our model significantly outperforms linear interpolation on dynamic sports datasets in terms of mean intersection over union between predicted and ground truth bounding boxes.
Automatically determining three-dimensional human pose from monocular RGB image data is a challenging problem. The two-dimensional nature of the input results in intrinsic ambiguities which make inferring depth particularly difficult. Recently, researchers have demonstrated that the flexible statistical modelling capabilities of deep neural networks are sufficient to make such inferences with reasonable accuracy. However, many of these models use coordinate output techniques which are memory-intensive, not differentiable, and/or do not spatially generalise well. We propose improvements to 3D coordinate prediction which avoid the aforementioned undesirable traits by predicting 2D marginal heatmaps under an augmented soft-argmax scheme. Our resulting model, MargiPose, produces visually coherent heatmaps whilst maintaining differentiability. We are also able to achieve state-of-the-art accuracy on publicly available 3D human pose estimation data.
In netball, analysis of the movement of players and the ball across different court locations can provide information about trends otherwise hidden. This study aimed to develop a method to discover latent passing patterns in women’s netball. Data for both pass location and playing position were collected from centre passes during selected games in the 2016 Trans-Tasman Netball Championship season and 2017 Australian National Netball League. A motif analysis was used to characterise passing-sequence observations. This revealed that the most frequent, sequential passing style from a centre pass was the “ABCD” motif in an alphabetical system, or in a positional system “Centre–Goal Attack–Wing Attack–Goal Shooter” and rarely was the ball passed back to the player it was received from. An association rule mining was used to identify frequent ball movement sequences from a centre pass play. The most confident rule flowed down the right-hand side of the court, however seven of the ten most confident rules demonstrated a preference for ball movement down the left-hand side of the court. These results can offer objective insight into passing sequences, and potentially inform team strategy and tactics. This method can also be generalised to other invasion sports.
Deep brain stimulation (DBS) is well recognized as an effective treatment for symptoms of movement disorders such as Parkinson's disease (PD), Essential Tremor, and dystonia. The selection of the appropriate contact on the DBS lead for optimal clinical efficacy can be challenging, particularly when considering directional leads. Electroencephalograms (EEG) and electrocorticography has been utilized to better understand the pathophysiology of PD but a methodology to provide an objective biomarker of effective stimulation has yet to be developed. Using machine learning techniques for feature extraction and classification, we contrast high resolution EEG captured during DBS against its resting state counterpart with the DBS off. We demonstrate, using 16 patients under DBS treatment for movement disorders, EEG's informative capacity to detect both effective DBS and the region undergoing stimulation.
We study deep learning approaches to inferring numerical coordinates for points of interest in an input image. Existing convolutional neural network-based solutions to this problem either take a heatmap matching approach or regress to coordinates with a fully connected output layer. Neither of these approaches is ideal, since the former is not entirely differentiable, and the latter lacks inherent spatial generalization. We propose our differentiable spatial to numerical transform (DSNT) to fill this gap. The DSNT layer adds no trainable parameters, is fully differentiable, and exhibits good spatial generalization. Unlike heatmap matching, DSNT works well with low heatmap resolutions, so it can be dropped in as an output layer for a wide range of existing fully convolutional architectures. Consequently, DSNT offers a better trade-off between inference speed and prediction accuracy compared to existing techniques. When used to replace the popular heatmap matching approach used in almost all state-of-the-art methods for pose estimation, DSNT gives better prediction accuracy for all model architectures tested.
Camera calibration is a preliminary step in sports analytics which enables us to transform player positions to standard playing area coordinates. While many camera calibration systems work well when the visual content contains sufficient clues, such as a key frame, calibrating without such information, such as may be needed when processing footage captured by a coach from the sidelines or stands, is challenging. In this paper an innovative automatic camera calibration system, which does not make use of any key frames, is presented for sports analytics. The proposed system consists of three components: a robust linear panorama module, a playing area estimation module, and a homography estimation module. It can eliminate distortion and calibrate the camera in each frame simultaneously, using correspondences between pairs of consecutive frames. Experiments on real data evaluate the performance and demonstrate the robustness of the system.
Basketball strategy is often focused on how to use space on the court. However, very little research has investigated performance from a spatial perspective beyond the now ubiquitous shooting heat maps. The aim of this study was to quantify how effectively teams move the ball across the basketball court and identify the most commonly occurring sequences of ball movement in international women's basketball. The results of the spatial analysis characterised trends in team play from the women's 2016 Olympic basketball competition and demonstrated that overall, the right-hand side under the basket and the top-right 3-point area were the most-effective areas on the court. In general terms, the right-hand side of the court was more effective than the left, and the middle of the court was more effective than the wings. Of the teams included in the study, the United States of America demonstrated the greatest overall effectiveness. Finally, the most commonly occurring ball movement sequences were identified with five of the seven teams demonstrating the same pattern. The quantification of spatial effectiveness in the current study provides insight into the specific tendencies of different teams and the areas that lead to the most effective outcomes. Coaches can apply this information to devise game plans aimed at counteracting the specific tendencies of opposing teams.
Zhen He合作论文数Department of Computer Science and Computer Engineering
La Trobe University9
Dietmar Saupe合作论文数Department of Computer and Information Science, University of Konstanz2
Tanja Schultz合作论文数Cognitive Systems Lab, University of Bremen;Language Technologies Institute, School of Computer Science, Carnegie Mellon University2
John Zeleznikow合作论文数Laboratory for Decision Support and Dispute Management, School of Management and Information Systems2