YouTube represents one of the largest scale and most sophisticated industrial recommendation systems in existence. In this paper, we describe the system at a high level and focus on the dramatic performance improvements brought by deep learning. The paper is split according to the classic two-stage information retrieval dichotomy: first, we detail a deep candidate generation model and then describe a separate deep ranking model. We also provide practical lessons and insights derived from designing, iterating and maintaining a massive recommendation system with enormous user-facing impact.
Multimodal speech and speaker modeling and recognition are widely accepted as vital aspects of state of the art human-machine interaction systems. While correlations between speech and lip motion as well as speech and facial expressions are widely studied, relatively little work has been done to investigate the correlations between speech and gesture. Detection and modeling of head, hand and arm gestures of a speaker have been studied extensively and these gestures were shown to carry linguistic information. A typical example is the head gesture while saying "yes/no". In this study, correlation between gestures and speech is investigated. In speech signal analysis, keyword spotting and prosodic accent event detection has been performed. In gesture analysis, hand positions and parameters of global head motion are used as features. The detection of gestures is based on discrete pre-designated symbol sets, which are manually labeled during the training phase. The gesture-speech correlation is modeled by examining the co-occurring speech and gesture patterns. This correlation can be used to fuse gesture and speech modalities for edutainment applications (i.e. video games, 3-D animations) where natural gestures of talking avatars are animated from speech. A speech driven gesture animation example has been implemented for demonstration
Empowered by advances in information technology, such as social media network, digital library and mobile computing, there emerges an ever-increasing amounts of multimedia data. As the key technology to address the problem of information overload, multimedia recommendation system has been received a lot of attentions from both industry and academia. This course aims to 1) provide a series of detailed review of state-of-the-art in multimedia recommendation; 2) analyze key technical challenges in developing and evaluating next generation multimedia recommendation systems from different perspectives and 3) give some predictions about the road lies ahead of us.
Our paper presents a novel high dimensional probability density estimation technique using any dimensionality reduction method. Our method first performs subspace reduction using any matrix factorization algorithm and estimates the density in the low-dimensional space using sample-point variable bandwidth kernel density estimation. Subsequently, the high dimensional density is approximated from the low dimensional density parameters. The reconstruction error due to dimensionality reduction process is also modeled in a principled and efficient manner to obtain the high dimensional density estimate. We show the effectiveness of our technique by using two popular dimensionality reduction tools, principal component analysis and non-negative matrix factorization. This technique is applied to AT&T, Yale, Pointing'04 and CMU-PIE face recognition datasets and improved performance compared to other dimensionality reduction and density estimation algorithms is obtained.
Automatic Language Identification (LID) in music has received significantly less attention than LID in speech. Here, we study the problem of LID in music videos uploaded on YouTube. We use a "bag-of-words" approach based on state-of-the-art content based audio-visual features and linear SVM classifiers for automatic LID. Our system obtains 48% accuracy for a corpus of 25000 music videos and 25 different languages.
Tau is a multiply phosphorylated protein that is essential for the development and maintenance of the nervous system. Errors in Tau action are associated with Alzheimer disease and related dementias. A huge literature has led to the widely held notion that aberrant Tau hyperphosphorylation is central to these disorders. Unfortunately, our mechanistic understanding of the functional effects of combinatorial Tau phosphorylation remains minimal. Here, we generated four singly pseudophosphorylated Tau proteins (at Thr(231), Ser(262), Ser(396), and Ser(404)) and four doubly pseudophosphorylated Tau proteins using the same sites. Each Tau preparation was assayed for its abilities to promote microtubule assembly and to regulate microtubule dynamic instability in vitro. All four singly pseudophosphorylated Tau proteins exhibited loss-of-function effects. In marked contrast to the expectation that doubly pseudophosphorylated Tau would be less functional than either of its corresponding singly pseudophosphorylated forms, all of the doubly pseudophosphorylated Tau proteins possessed enhanced microtubule assembly activity and were more potent at regulating dynamic instability than their compromised singly pseudophosphorylated counterparts. Thus, the effects of multiple pseudophosphorylations were not simply the sum of the effects of the constituent single pseudophosphorylations; rather, they were generally opposite to the effects of singly pseudophosphorylated Tau. Further, despite being pseudophosphorylated at different sites, the four singly pseduophosphorylated Tau proteins often functioned similarly, as did the four doubly pseudophosphorylated proteins. These data lead us to reassess the conventional view of combinatorial phosphorylation in normal and pathological Tau action. They may also be relevant to the issue of combinatorial phosphorylation as a general regulatory mechanism.
We consider the problem of large-scale video classification. Our attention is focused on online video services since they can provide rich cross-video signals derived from user behavior. These signals help us to extract correlated information across videos which are co-browsed, co-uploaded, co commented, co-queried, etc. Majority of the video classification methods omit this rich information and focus solely on a single test instance. In this paper, we propose a video classification system that exploits various cross-video signals offered by large-scale video databases. In our experiments, we show up to 4.5% absolute equal error rate (17% relative) improvement over the baseline on four video classification problems.
In this paper, we propose a novel approach for detecting highlights in sports videos. The videos are temporally decomposed into a series of events based on an unsupervised event discovery and detection framework. The framework solely depends on easy-to-extract low-level visual features such as color histogram (CH) or histogram of oriented gradients (HOG), which can potentially be generalized to different sports. The unigram and bigram statistics of the detected events are then used to provide a compact representation of the video. The effectiveness of the proposed representation is demonstrated on cricket video classification: Highlight vs. Non-Highlight for individual video clips (7000 training and 7000 test instances). We achieve a low equal error rate of 12.1% using event statistics based on CH and HOG features.
We have utilized tau‐assembled and tau‐stabilized microtubules (MTs), in the absence of taxol, to investigate the effects of tau isoforms with three and four MT binding repeats upon kinesin‐driven MT gliding. MTs were assembled in the presence of either 3‐repeat tau (3R tau) or 4‐repeat tau (4R tau) at tau:tubulin dimer molar ratios that approximate those found in neurons. MTs assembled with 3R tau glided at 31.1 μm/min versus 25.8 μm/min for 4R tau, a statistically significant 17% difference. Importantly, the gliding rates for either isoform did not change over a fourfold range of tau concentrations. Further, tau‐assembled MTs underwent minimal dynamic instability behavior while gliding and moved with linear trajectories. In contrast, MTs assembled with taxol in the absence of tau displayed curved gliding trajectories. Interestingly, addition of 4R tau to taxol‐stabilized MTs restored linear gliding, while addition of 3R tau did not. The data are consistent with the ideas that (i) 3R and 4R tau‐assembled MTs possess at least some isoform‐specific features that impact upon kinesin translocation, (ii) tau‐assembled MTs possess different structural features than do taxol‐assembled MTs, and (iii) some features of tau‐assembled MTs can be masked by prior assembly by taxol. The differences in kinesin‐driven gliding between 3R and 4R tau suggest important features of tau function related to the normal shift in tau isoform composition that occurs during neural development as well as in neurodegeneration caused by altered expression ratios of otherwise normal tau isoforms. © 2010 Wiley‐Liss, Inc.
Simultaneous registration and segmentation (SRS) provides a powerful framework for tracking an object of interest in an image sequence. The state-of-the-art SRS-based tracking methods assume that the illumination is maintained constant across consecutive frames. However, this assumption does not hold in many natural image sequences due to dynamic light source and shadows. We propose a generalized model for SRS-based tracking in this paper to account for non-uniform additive illumination changes. More specifically, we introduce two new terms in the SRS energy functional which address the above mentioned problem. The first term couples the shape-based cue and intensity-based cue to establish a correspondence between them. The second term compensates for the illumination change which is complementary to the first term. We demonstrate that the proposed SRS energy functional yields superior performance over the state-of-the-art SRS-based methods for various indoor and outdoor image sequences.
Tracking of curvilinear structures is a task of fundamental importance in the quantitative analysis of biological structures such as neurons, blood vessels, retinal interconnects, microtubules, etc. The state of the art HMM-based contour tracking scheme for tracking microtubules, while performing well in most scenarios, can miss the track if, during its growth, it intersects another microtubule in its neighbourhood. In this paper we present a graphical model-based tracking algorithm which propagates across frames information about the dynamics of all the microtubules. This allows the algorithm to faithfully differentiate the contour of interest from others that contribute to the clutter, and maintain tracking accuracy. We present results of experiments on real microtubule images captured using fluorescence microscopy, and show that our proposed scheme outperforms the existing HMM-based scheme.
This paper focuses on contour tracking, an important problem in computer vision, and specifically on open contours that often directly represent a curvilinear object. Compelling applications are found in the field of bioimage analysis where blood vessels, dendrites, and various other biological structures are tracked over time. General open contour tracking, and biological images in particular, pose major challenges including scene clutter with similar structures (e.g., in the cell), and time varying contour length due to natural growth and shortening phenomena, which have not been adequately answered by earlier approaches based on closed and fixed end-point contours. We propose a model-based estimation algorithm to track open contours of time-varying length, which is robust to neighborhood clutter with similar structures. The method employs a deformable trellis in conjunction with a probabilistic (hidden Markov) model to estimate contour position, deformation, growth and shortening. It generates a maximum a posteriori estimate given observations in the current frame and prior contour information from previous frames. Experimental results on synthetic and real-world data demonstrate the effectiveness and performance gains of the proposed algorithm.
We present a method for object tracking over time sequence imagery. The image plane is represented with a 4-connected planar graph where vertices are associated with pixels. On each image, the outer contour of the object is localized by finding the optimal cycle in the graph such that a cost function based on temporal, appearance and shape priors is minimized. Our contribution is the particle filtering-based framework to integrate the shape cue with the temporal and appearance cues. We demonstrate that incorporating the shape prior yields promising performance improvement over temporal and appearance priors on various object tracking scenarios.
In this paper, we present an algorithm for occlusion boundary detection. The main contribution is a probabilistic detection framework defined on spatio-temporal lattices, which enables joint analysis of image frames. For this purpose, we introduce two complementary cost functions for creating the spatio-temporal lattice and for performing global inference of the occlusion boundaries, respectively. In addition, a novel combination of low-level occlusion features is discriminatively learnt in the detection framework. Simulations on the CMU Motion Dataset provide ample evidence that proposed algorithm outperforms the leading existing methods.
We introduce a dynamical model for simultaneous registration and segmentation in a variational framework for image sequences, where the dynamics is incorporated using a Bayesian formulation. A linear stochastic equation relating the tracked object (or a region of interest) is first derived under the assumption that the successive images in the sequence are related by a dense and possibly non-linear displacement field. This derivation allows for the use of a computationally efficient and recursive implementation of the Bayesian formulation in this framework. The contour of the tracked object returned by the dynamical model is not only close to the previously detected shape but is also consistent with the temporal statistics of the tracked object. The performance of the proposed approach is evaluated on real image sequences. It is shown that, with respect to a variety of error metrics such as F-measure, mean absolute deviation and Hausdorff distance, the proposed approach outperforms the state-of-the art approach without the dynamical model.
E. Erzin合作论文数College of Engineering, Koc University4
Vivek Kwatra合作论文数Computer Science Department at UNC Chapel Hill4