For skeleton-based action recognition, since there usually exists many nuances between different datasets, including viewpoints, the number of available joints for a skele-ton, the type of actions, etc, it hinders to apply and leverage the knowledge of a pretrained model for one dataset to an-other except retraining a new model for the target dataset. To address this issue, we propose a cross-domain knowledge transfer module based on gradient reversal layer along with adaptive graph convolutional network to effectively transfer the knowledge from one domain to another. The adaptive graph convolution module allows the proposed method to adaptively learn the topological relation between joints and is very useful for the scenarios when the numbers of skele-ton joints for the two domains are different and the topo-logical correspondences of joints are not clearly specified. With extensive experiments from NTU-RGB+D 60 to the PKU, CITI3D, and NW datasets, the proposed approach achieves significantly better results than other state-of-the-art spatio-temporal graph convolutional network methods which are trained on the target dataset only, and this also demonstrates the effectiveness of the proposed approach.
In this paper we provide an auditory processing system with higher biological plausibility than previous studies to solve the problem of acoustic anomaly detection in household environment. First, in the proposed system the log filter bank is adopted for extracting audio features, simulating the function of the peripheral auditory system (outer, middle and inner ears to auditory nerves). Next, we use multi-layer neural networks to imitate auditory cortex in human brains, in order to extract abstract semantic contents. Then, a semantic pointer architecture model is used to imitate prefrontal cortex, basal ganglia, and thalamus, in which the anomaly is detected using symbol-like rules. Compared with other anomaly detection methods with different biological plausibility in performance, our proposed method gets the best result on the testing set, with 0.956 AUC. Meanwhile, it takes less computational time to detect the anomaly. Hence, it is suitable for detecting acoustic anomalies in real-world cases.
In this paper, we propose an end-to-end key-player-based group activity recognition network specially applied to the identification of basketball offensive tactics in limited data scenarios. Our previous studies show that basketball tactics can be better recognized via key player detection with multiple instance learning (MIL) using the support vector machine (SVM). However, the SVM in that work is required to extract features depending on basketball- and tactic-specific knowledge for good performance. Thus, in this study, we develop an end-to-end trainable neural network without prior knowledge and integrate MIL into it. As long as a tactic label is given, MIL can train the network to identify tactic's key players. For testing, our network can recognize the key players in a video clip and provide a tag of the tactic related to them. Like other neural network models, our network requires a large annotated dataset. At the same time, we could collect only a few labeled data, which is common in dealing with group activity recognition. To overcome such a limitation, we propose a novel data augmentation framework, the tactical-based conditional generative adversarial network (GAN), for generating new labeled trajectories. The experimental results show that our method significantly improves 9.13 % in tactic recognition and 4.965 % in key player detection.
Apart from the coherence and fluency of responses, an empathetic chatbot emphasizes more on people's feelings. By considering altruistic behaviors between human interaction, empathetic chatbots enable people to get a better interactive and supportive experience. This study presents a framework whereby several empathetic chatbots are based on understanding users' implied feelings and replying empathetically for multiple dialogue turns. We call these chatbots CheerBots. CheerBots can be retrieval-based or generative-based and were finetuned by deep reinforcement learning. To respond in an empathetic way, we develop a simulating agent, a Conceptual Human Model, as aids for CheerBots in training with considerations on changes in user's emotional states in the future to arouse sympathy. Finally, automatic metrics and human rating results demonstrate that CheerBots outperform other baseline chatbots and achieves reciprocal altruism. The code and the pre-trained models will be made available.
Video re-localization aims to localize a sub-sequence, called target segment, in an untrimmed reference video that is similar to a given query video. In this work, we propose an attention-based model to accomplish this task in a weakly supervised setting. Namely, we derive our CNN-based model without using the annotated locations of the target segments in reference videos. Our model contains three modules. First, it employs a pre-trained C3D network for feature extraction. Second, we design an attention mechanism to extract multiscale temporal features, which are then used to estimate the similarity between the query video and a reference video. Third, a localization layer detects where the target segment is in the reference video by determining whether each frame in the reference video is consistent with the query video. The resultant CNN model is derived based on the proposed co-attention loss which discriminatively separates the target segment from the reference video. This loss maximizes the similarity between the query video and the target segment while minimizing the similarity between the target segment and the rest of the reference video. Our model can be modified to fully supervised re-localization. Our method is evaluated on a public dataset and achieves the state-of-the-art performance under both weakly supervised and fully supervised settings.
Phonotactic constraints can be employed to distinguish languages by representing a speech utterance as a multinomial distribution or phone events. In the present study, we propose a new learning mechanism based on subspace-based representation, which can extract concealed phonotactic structures from utterances, for language verification and dialect/accent identification. The framework mainly involves two successive parts. The first part involves subspace construction. Specifically, it decodes each utterance into a sequence of vectors filled with phone-posteriors and transforms the vector sequence into a linear orthogonal subspace based on low-rank matrix factorization or dynamic linear modeling. The second part involves subspace learning based on kernel machines, such as support vector machines and the newly developed subspace-based neural networks (SNNs). The input layer of SNNs is specifically designed for the sample represented by subspaces. The topology ensures that the same output can be derived from identical subspaces by modifying the conventional feed-forward pass to fit the mathematical definition of subspace similarity. Evaluated on the “General LR” test of NIST LRE 2007, the proposed method achieved up to 52%, 46%, 56%, and 27% relative reductions in equal error rates over the sequence-based PPR-LM, PPR-VSM, and PPR-IVEC methods and the lattice-based PPR-LM method, respectively. Furthermore, on the dialect/accent identification task of NIST LRE 2009, the SNN-based system performed better than the aforementioned four baseline methods.
Empirical mode decomposition (EMD) is an extensively utilized tool in a time-frequency analysis. However, disturbances, such as impulse noise, can result in a mode-splitting effect, in which one physically meaningful component is split into two or more intrinsic mode functions (IMFs). In this paper, we propose a novel method, minimum arclength EMD (MA-EMD), to robustly decompose time series data with impulse-like noises. The idea is to apply a minimum arclength criterion to adjust the knot positions of impulses during the sifting process in EMD. In this way, the impulse-like artifact is extracted with the first IMF, and the mode splitting effect of the latter decomposition is alleviated. Furthermore, when the first IMF contains the desired information, we separate the spikes and the first IMF by adding a pair of masking signals. For using this masking-aided MA-EMD (MAMA-EMD) method, we also mathematically derived the appropriate ranges of the frequency and the amplitude of the masking signal. The MAMA-EMD is utilized to deal with the simulated Duffing wave and four real-world data, including electrical current, vibration signals, the cyclic alternating pattern in sleep EEG (electroencephalography), and circadian of core body temperature. The results show that the MA-EMD and MAMA-EMD have a sound improvement when encountering impulse noises.
Music videos are one of the most popular types of video streaming services, and instrument playing is among the most common scenes in such videos. In order to understand the instrument-playing scenes in the videos, it is important to know what instruments are played, when they are played, and where the playing actions occur in the scene. While audio-based recognition of instruments has been widely studied, the visual aspect of music instrument playing remains largely unaddressed in the literature. One of the main obstacles is the difficulty in collecting annotated data of the action locations for training-based methods. To address this issue, we propose a weakly supervised framework to find when and where the instruments are played in the videos. We propose using two auxiliary models: 1) a sound model and 2) an object model to provide supervision for training the instrument-playing action model. The sound model provides temporal supervisions, while the object model provides spatial supervisions. They together can simultaneously provide temporal and spatial supervisions. The resulting model only needs to analyze the visual part of a music video to deduce which, when, and where instruments are played. We found that the proposed method significantly improves localization accuracy. We evaluate the result of the proposed method temporally and spatially on a small dataset (a total of 5400 frames) that we manually annotated.
We address offensive tactic recognition in broadcast basketball videos. As a crucial component towards basketball video content understanding, tactic recognition is quite challenging because it involves multiple independent players, each of which has respective spatial and temporal variations. Motivated by the observation that most intra-class variations are caused by non-key players, we present an approach that integrates key player detection into tactic recognition. To save the annotation cost, our approach can work on training data with only video-level tactic annotation, instead of key players labeling. Specifically, this task is formulated as an MIL (multiple instance learning) problem where a video is treated as a bag with its instances corresponding to subsets of the five players. We also propose a representation to encode the spatio-temporal interaction among multiple players. It turns out that our approach not only effectively recognizes the tactics but also precisely detects the key players.
Recent years have witnessed an increased interest in the application of persistent homology, a topological tool for data analysis, to machine learning problems. Persistent homology is known for its ability to numerically characterize the shapes of spaces induced by features or functions. On the other hand, deep neural networks have been shown effective in various tasks. To our best knowledge, however, existing neural network models seldom exploit shape information. In this paper, we investigate a way to use persistent homology in the framework of deep neural networks. Specifically, we propose to embed the so-called "persistence landscape," a rather new topological summary for data, into a convolutional neural network (CNN) for dealing with audio signals. Our evaluation on automatic music tagging, a multi-label classification task, shows that the resulting persistent convolutional neural network (PCNN) model can perform significantly better than state-of-the-art models in prediction accuracy. We also discuss the intuition behind the design of the proposed model, and offer insights into the features that it learns.
Unlike conventional omnidirectional antennas, switched beam antennas exploit antenna arrays and signal processing techniques to focus energy in a specific beam-width and orientation. Recent research has shown that access points of a WLAN can exploit such switched beam antennas to increase the overall network capacity. The achievable sum rate of a WLAN with switch beam antennas is however mainly determined by how each AP selects its beam, including the orientation and width, and how each client associates with a proper AP. The goal of this paper is to solve the Joint Beam configuration and Client association (JBC) problem such that the sum rate of all clients in the network can be maximized. We formulate the JBC problem as a mixed integer linear programming model, and propose a 2-approximation algorithm to solve it. Our proposed algorithm has two distinctive properties: 1) it can be realized as a distributed protocol that allows the APs to configure their beams without the help of a central coordinator, and 2) it can be generally applied both in specific scenarios, where exact client locations are known, and in uncertain scenarios, where only geographic client distribution is given. Finally, we adjust the sum-rate maximization algorithm to the throughput maximization algorithm, which further takes medium sharing among clients into account. The simulation results show that the proposed algorithm outperforms both WLANs using omnidirectional antennas and other heuristics using switched beam antennas.
Player segmentation in team sports videos is challenging but crucial to video semantic understanding, such as player interaction identification and tactic analysis. We leverage the appearance similarity among players of the same team, and cast this task as a co-segmentation problem. In this way, the extra knowledge shared across players significantly reduces unfavorable uncertainty in segmenting individual players. We are also aware that the performance of co-segmentation highly depends on the used features, and further propose a contrast-based approach to estimate the discriminant power of each feature in an unsupervised manner. It turns out that our approach can properly fuse features by assigning higher weights to discriminant ones, and result in remarkable performance gains. The promising results on segmenting basketball players manifest the effectiveness of our approach.
This study addresses the task of aligning lyrics with accompanied singing recordings. With a vowel-only representation of lyric syllables, our approach evaluates likelihood scores of vowel types with glottal pulse shapes and formant frequencies extracted from a small set of singing examples. The proposed vowel likelihood model is used in conjunction with a prior model of frame-wise syllable sequence in determining an optimal evolution of syllabic position. In lyrics alignment experiments, we optimized numerical parameters on two independent development sets and then tested the optimized system on two other datasets. New objective performance measures are introduced in the evaluation to provide further insight into the quality of alignment. Use of glottal pulse shapes and formant frequencies is shown by a controlled experiment to account for a 0.07 difference in average normalized alignment error. Another controlled experiment demonstrates that, with a difference of 0.03, F0-invariant glottal pulse shape gives a lower average normalized alignment error than does F0-invariant spectrum envelope, the latter being assumed by MFCC-based timbre models.
To estimate the degree of sincerity conveyed by a speech utterance and received by listeners, we propose an instance-based learning framework with shallow neural networks. The framework plays as not only a regressor that intends to fit the predicted value to the actual value but also a ranker that preserves the relative target magnitude between each pair of utterances, in an attempt to derive a higher Spearman’s rank correlation coefficient. In addition to describing how to simultaneously minimize regression and ranking losses, the issue of how utterance pairs work in the training and evaluation phases is also addressed by two kinds of realizations. The intuitive one is related to random sampling while the other seeks for representative utterances, named anchors, to form non-stochastic pairs. Our system outperforms the baseline by more than 25% relative improvement in the development set.
A system for automatically evaluating singing enthusiasm is proposed in this study. The definition of singing enthusiasm is how much enthusiasm is perceived in a song being evaluated. This system evaluates the singing enthusiasm on the basis of pitch accuracy, vibrato, diminuendo, roughness, and the correlation between pitch and loudness. A support vector regression (SVR) machine is used for the evaluation. This system can deal with songs having multiple phrases without any reference information such as the pitch ground truth or phrase location. To the authors' knowledge, only one such system has previously been proposed which could only handle a single phrase of about 5-second long. To evaluate this system, a singing corpus with 342 song clips sung by nine participants was recorded and ground-truth enthusiasm evaluation scores were obtained by an online questionnaire. The experimental results obtained from a leave-one-singer-out test revealed that the enthusiasm scores evaluated by the proposed system had a significant positive correlation coefficient of 0.51 with the human-labeled ground truth.
Modeling the association between music and emotion has been considered important for music information retrieval and affective human computer interaction. This paper presents a novel generative model called acoustic emotion Gaussians (AEG) for computational modeling of emotion. Instead of assigning a music excerpt with a deterministic (hard) emotion label, AEG treats the affective content of music as a (soft) probability distribution in the valence-arousal space and parameterizes it with a Gaussian mixture model (GMM). In this way, the subjective nature of emotion perception is explicitly modeled. Specifically, AEG employs two GMMs to characterize the audio and emotion data. The fitting algorithm of the GMM parameters makes the model learning process transparent and interpretable. Based on AEG, a probabilistic graphical structure for predicting the emotion distribution from music audio data is also developed. A comprehensive performance study over two emotion-labeled datasets demonstrates that AEG offers new insights into the relationship between music and emotion (e.g., to assess the “affective diversity” of a corpus) and represents an effective means of emotion modeling. Readers can easily implement AEG via the publicly available codes. As the AEG model is generic, it holds the promise of analyzing any signal that carries affective or other highly subjective information.
The problem of finding critical features for different crime types has been the focus in the field of environmental criminology because crime would lead to bad zoning in urban areas. However, conventional analysis ignores social dynamics of human beings. With the increasing growth of location-based social networks, the fine-grained data associated with social connections and the geographical information of users are available for representing the spatio-social dynamics of people. In this work, we devise a series of features to characterize an urban climate by data obtained from Foursquare and Gowalla in San Francisco. As for crime, we take use of crime data provided by the authorities. The features we mined are based on two general signals: geographical features that capture the distribution of various types of venues in urban areas, and social features that model the topological interactions between people in a region. We use these features to analyze and detect urban areas with high crime activities. The experimental results show the effectiveness of the proposed features on five different crime types and encourage future advanced criminal analysis using location-based social network data.
This paper presents a novel approach to extraction of vocal melodies from accompanied singing recordings. Central to our approach is a model of vocal fundamental frequency (F0) likelihood that integrates acoustic-phonetic knowledge and real-world data. This model consists of a timbral fitness score and a loudness measure of each F0 candidate. Timbral fitness is measured for the partial amplitudes of an F0 candidate, with respect to a small set of vocal timbre examples. This F0-specific measurement of timbral fitness depends on an acoustic-phonetic F0 modification of each timbre example. In the loudness part of the likelihood model, sinusoids are detected, tracked, and pruned to give loudness values that minimize interference from the accompaniment. A final F0 estimate is determined by a prior model of F0 sequence in addition to the likelihood model. Melody extraction is completed by detecting voiced time positions according to the singing voice loudness variations given by the estimated F0 sequence. The numerical parameters involved in our approach were optimized on three development sets from different sources before the system was evaluated on ten test sets separate from these development sets. Controlled experiments show that use of the timbral fitness score accounts for a 13% difference in overall accuracy.
In this paper, we first reformulate the derivation of the conventional i-vector scheme, which is the state-of-the-art utterance representation for speaker verification, as a modeling of universal background model (UBM)-based mixtures of factor analyzers (UMFA), and then propose a clustering-based UMFA method called CMFA. In UMFA, each analyzer is characterized by a subspace, and the same projection coordinate of an utterance into individual subspaces is called the i-vector. We relax this assumption by grouping the mixture components of the UBM into clusters according to their acoustic traits. Therefore, in CMFA, each utterance is represented by multiple i-vectors, each of which generated by similar subspaces associated with a same cluster. We also investigate two strategies for merging these i-vectors into a single one to be applied in the classifier of the conventional i-vector framework. The results of experiments conducted on the male portion of the core task in the NIST 2005 Speaker Recognition Evaluation (SRE) in terms of normalized decision cost function (minDCF) and equal error rate (EER) demonstrate the merits of the new i-vector method over the conventional i-vector method.