Visual Emotion Analysis (VEA) is attracting increasing attention. One of the biggest challenges of VEA is to bridge the affective gap between visual clues in a picture and the emotion expressed by the picture. As the granularity of emotions increases, the affective gap increases as well. Existing deep approaches try to bridge the gap by directly learning discrimination among emotions globally in one shot. They ignore the hierarchical relationship among emotions at different affective levels, and the variation in the affective level of emotions to be classified. In this paper, we present the multi-level dependent attention network (MDAN) with two branches to leverage the emotion hierarchy and the correlation between different affective levels and semantic levels. The bottom-up branch directly learns emotions at the highest affective level and largely prevents hierarchy violation by explicitly following the emotion hierarchy while predicting emotions at lower affective levels. In contrast, the top-down branch aims to disentangle the affective gap by one-to-one mapping between semantic levels and affective levels, namely, Affective Semantic Mapping. A local classifier is appended at each semantic level to learn discrimination among emotions at the corresponding affective level. Then, we integrate global learning and local learning into a unified deep framework and optimize it simultaneously. Moreover, to properly model channel dependencies and spatial attention while disentangling the affective gap, we carefully designed two attention modules: the Multi-head Cross Channel Attention module and the Level-dependent Class Activation Map module. Finally, the proposed deep framework obtains new state-of-the-art performance on six VEA benchmarks, where it outperforms existing state-of-the-art methods by a large margin, e.g., +3.85% on the WEBEmo dataset at 25 classes classification accuracy.
LiteSing proposed in this paper is a high-quality singing voice synthesis (SVS) system, which is fast, lightweight and expressive. This model mainly stacks several non-autoregressive WaveNet blocks in the encoder and decoder under a generative adversarial architecture, predicts full conditions from the musical score, and generates acoustic features from these conditions. The full conditions in this paper consist of dynamic spectrogram energy, voiced/unvoiced (V/UV) decision and dynamic pitch curve, which are proven related to the expressiveness. We predict the pitch and the timbre features separately, avoiding the interdependence between these two features. Instead of neural network vocoders, a parametric WORLD vocoder is employed for the pitch curve consistency. Experiment results show that LiteSing outperforms the baseline model using feed-forward Transformer by 1.386 times faster on inference speed, 15 times smaller on training parameters number, and achieves a similar MOS on sound quality. Through an A/B test, LiteSing achieves 67.3% preference rate over baseline in pitch curve and dynamic spectrogram energy prediction. which demonstrates the advantage of LiteSing over the other compared models.
Timbre evokes emotion in music, as do loudness, pitch, rhythm, and other music qualities. Recent research has confirmed the correlation between timbral features with emotion in isolated musical instrument tones. Attack time, spectral centroid and its deviation have significant correlation with emotional characteristics in tones of sustaining (e.g., trumpet) and non-sustaining instruments (e.g., harp). Previous work has considered instrument tones of one to two seconds duration. This paper presents a similar experiment with pairwise comparison of very short (250 ms) isolated tones of nonsustaining instruments, including plucked string, pitched percussion, and keyboard instruments. The tones were investigated for the emotion categories Happy, Sad, Heroic, Scary, Comic, Shy, Joyful, and Depressed. The results agreed with those for one-second tones, with plucked string tones evoking negative emotional categories and pitched percussion tones evoking positive categories. Surprisingly, the emotional characteristics were even more clearly distinguished for 250 ms tones. We also found that the density of significant harmonics, a feature that has particular relevance to non-sustaining tones, was observed to achieve especially strong correlation for all tested emotion categories in these tones with early release.
Recent research has shown the connection between timbre and music emotion and has focused on sustaining instruments. This paper undertakes listening tests to investigate the connection for nonsustaining instruments including plucked string, pitched percussion, and keyboard sounds compared pairwise by listening test subjects over eight emotional categories. Different durations were studied to investigate the effect of early release on the emotional characteristics. The experiments showed that the harpsichord, marimba, vibraphone, and xylophone were highly rated for positive emotional characteristics; the guitar, harp, and plucked violin were highly rated for negative emotional characteristics; and the piano was rated emotionally neutral. There were differences in emotional characteristics due to early release, such as the harpsichord decreasing in ranking for positive emotional characteristics on sounds with early releases. By correlating timbral features with the listener rankings, we found that decay slope and density of significant harmonics were significant timbral features for many emotional characteristics.
Though previous research has shown the effects of reverberation on clarity, spaciousness, and other perceptual aspects of music, it is still largely unknown to what extent reverberation influences the emotional characteristics of musical instrument sounds.This paper investigates the effect of simple parametric reverberation on music emotion, in particular, the effect of reverberation length and amount.We conducted a listening test to compare the effect of reverberation on the emotional characteristics of eight instrument sounds representing the wind and bowed string families.We compared these sounds over eight emotional categories.We found that reverberation length and amount had a strongly significant effect on the emotional characteristics Romantic and Mysterious and a medium effect on Sad, Scary, and Heroic for the samples we tested.Interestingly, for Comic, reverberation length and amount had the opposite effect; that is, anechoic tones were judged most Comic.Reverb had a mild effect on Happy and relatively little effect on Shy.These results give audio engineers and musicians an interesting perspective on simple parametric artificial reverberation.
Music conveys emotions by means of pitch, rhythm, loudness, and many other musical qualities. It was recently confirmed that timbre also has direct association with emotion, for example, that a horn is perceived as sad and a trumpet heroic in even isolated instrument tones. As previous work has mainly focused on sustaining instruments such as bowed strings and winds, this paper presents an experiment with non-sustaining instruments, using a similar approach with pairwise comparisons of tones for emotion categories. Plucked string, mallet percussion, and keyboard instrument tones were investigated for eight emotions: Happy, Sad, Heroic, Scary, Comic, Shy, Joyful, and Depressed. We found that plucked string tones tended to be Sad and Depressed, while harpsichord and mallet percussion tones induced positive emotions such as Happy and Heroic. The piano was emotionally neutral. Beyond spectral centroid and its deviation, which are important features in sustaining tones, decay slope was also significantly correlated with emotion in non-sustaining tones.
为了克服传统的以单幅图像作为信息来源的水平集模型分割复杂背景图像的局限性,结合区域生长法和水平集方法各自的特点,提出了一种新的由多幅图像信息构建的水平集分割算法模型.在运用水平集方法分割人体腹腔图像前,首先运用本文提出的一种有效的区域生长法在腹腔图像中得到肝脏的粗略分割结果作为先验形状图像.通过先验形状图像在Chan-Vese模型下控制水平集的演化,使活动轮廓的先验形状信息融合到水平集分割算法模型中,同时,利用Li模型在人体腹腔图像中进一步获取肝脏的边缘信息.这种融合多幅图像信息的复合水平集分割算法模型能够充分利用图像信息,有效地描述水平集方法中活动轮廓与目标区域肝脏的关系.通过实验验证,提出的算法模型能够很好地从人体腹腔图像中提取出肝脏区域.
Spectral centroid is a primary feature of timbre. Nevertheless, no previous work has considered the discrimination and identification of spectrally tilted instrument tones. Our study shows spectral centroid to be so dominant that spectral tilts are equally detectable regardless of instrument. Negative tilts (decreases in spectral centroid) result in higher discrimination and greater loss of instrument identity than positive tilts.
Music is one of the strongest triggers of emotions. Melody, rhythm, and harmony are important triggers, but what about timbre? Do musical instruments have an emotional predisposition? For example, is the melancholy sound of the English horn due to its timbre or how it is used? Though music emotion recognition has received a lot of attention; researchers have only recently begun considering the relationship between emotion and timbre. Therefore, we designed listening tests to compare sounds from eight wind and bowed string instruments. We wanted to know if some sounds were consistently perceived as being happier or sadder in pairwise comparisons. Eight emotions were tested. The results showed strong emotional predispositions for each instrument. The emotions Happy, Joyful, Heroic, and Comic were strongly correlated with one another, and the violin, trumpet, and clarinet best evoked these emotions. Sad and Depressed were also strongly correlated and were best evoked by the horn and flute. Scary was an emotional outlier, and the oboe had an emotionally neutral disposition. We also correlated emotion and several spectral features and found that emotion correlated significantly with average spectral centroid and spectral centroid deviation for nearly all emotions. To determine which spectral features were most important aside from spectral centroid, we conducted follow-up listening tests of centroid-equalized sounds as well as static sounds. The results showed that the even/odd harmonic ratio significantly correlated with most emotions. This suggests that the even/odd harmonic ratio is perhaps the most salient timbral feature after attack time and brightness. These results provide new insights for orchestration.
Timbre and emotion are two of the most important aspects of musical sounds. Both are complex and multidimensional, and strongly interrelated. Previous research has identified many different timbral attributes, and shown that spectral centroid and attack time are the two most important dimensions of timbre. However, a consensus has not emerged about other dimensions. This study will attempt to identify the most perceptually relevant timbral attributes after spectral centroid and attack time. To do this, we will consider various sustained musical instrument tones where spectral centroid and attack time have been equalized. While most previous timbre studies have used discrimination and dissimilarity tests to understand timbre, researchers have begun using emotion tests recently. Previous studies have shown that attack and spectral centroid play an essential role in emotion perception, and they can be so strong that listeners do not notice other spectral features very much. Therefore, in this paper, to isolate the third most important timbre feature, we designed a subjective listening test using emotion responses for tones equalized in attack, decay, and spectral centroid. The results showed that the even/odd harmonic ratio is the most salient timbral feature after attack time and spectral centroid.
In this paper, we describe the Beatsens Team solution of Emotion in Music task in MediaEval benchmarking campaign 2014. We extracted and designed several sets of features and used continuous conditional random eld(CCRF) for dynamic emotion characterization task. The best runs for Pearson correlation are 0:23 0:56 and 0:12 0:55 of valence and arousal respectively, for RMSE are 0:12 0:06 and 0:09 0:05.
Music emotion recognition, which aims to automatically recognize the affective content of a piece of music, has become one of the key components of music searching, exploring, and social networking applications. Although researchers have given more and more attention to music emotion recognition studies, the recognition performance has come to a bottleneck in recent years. One major reason is that experts' labels for music emotion are mostly song-level, while music emotion usually varies within a song. Traditional methods have considered each song as a single instance and have built models based on song-level features. However, they ignored the dynamics of music emotion and failed to capture accurate emotion-feature correlations. In this paper, we model music emotion recognition as a novel multi-label multi-layer multi-instance multi-view learning problem: music is formulated as a hierarchical multi-instance structure (e.g., song-segment-sentence) where multiple emotion labels correspond to at least one of the instances with multiple views of each layer. We propose a Hierarchical Music Emotion Recognition model (HMER) -- a novel hierarchical Bayesian model using sentence-level music and lyrics features. It captures music emotion dynamics with a song-segment-sentence hierarchical structure. HMER also considers emotion correlations between both music segments and sentences. Experimental results show that HMER outperforms several state-of-the-art methods in terms of $F_1$ score and mean average precision.
Music is one of the strongest triggers of emotions. Recent studies have shown strong emotional predispositions for musical instrument timbres. They have also shown significant correlations between spectral centroid and many emotions. Our recent study on spectral centroid-equalized tones further suggested that the even/odd harmonic ratio is a salient timbral feature after attack time and brightness. The emergence of the even/odd harmonic ratio motivated us to go a step further: to see whether the spectral shape of musical instruments alone can have a strong emotional predisposition. To address this issue, we conducted followup listening tests of static tones. The results showed that the even/odd harmonic ratio again significantly correlated with most emotions, consistent with the theory that static spectral shapes have a strong emotional predisposition.
Time-sync video tagging aims to automatically generate tags for each video shot. It can improve the user's experience in previewing a video's timeline structure compared to traditional schemes that tag an entire video clip. In this paper, we propose a new application which extracts time-sync video tags by automatically exploiting crowdsourced comments from video websites such as Nico Nico Douga, where videos are commented on by online crowd users in a time-sync manner. The challenge of the proposed application is that users with bias interact with one another frequently and bring noise into the data, while the comments are too sparse to compensate for the noise. Previous techniques are unable to handle this task well as they consider video semantics independently, which may overfit the sparse comments in each shot and thus fail to provide accurate modeling. To resolve these issues, we propose a novel temporal and personalized topic model that jointly considers temporal dependencies between video semantics, users' interaction in commenting, and users' preferences as prior knowledge. Our proposed model shares knowledge across video shots via users to enrich the short comments, and peels off user interaction and user bias to solve the noisy-comment problem. Log-likelihood analyses and user studies on large datasets show that the proposed model outperforms several state-of-the-art baselines in video tagging quality. Case studies also demonstrate our model's capability of extracting tags from the crowdsourced short and noisy comments.
Previous studies related to MP3 compression have investigated the discrimination of compressed instrument tones. However, these studies have not considered the effect of MP3 compression on the timbre space. In the current study, in a triadic listening test subjects were asked to rate the dissimilarity of all pairs of eight original instrument tones from various instrument families. The same process was repeated on MP3-compressed tones using various bit rates (32, 64, and 128 Kbps). The results showed strong correlations between the dissimilarity scores of the original and compressed tones, indicating relatively subtle perceptual changes overall. The 2-D multidimensional scaling solutions for tones compressed with bit rates of 64 and 128 Kbps were very similar to the original but the coordinates changed more dramatically for a bit rate of 32 Kbps (especially in the saxophone), indicating a change in the underlying timbre space for low bit rates.
Music is one of the strongest inducers of emotion in humans. Melody, rhythm, and harmony provide the primary triggers, but what about timbre? Do the musical instruments have underlying emotional characters? For example, is the well-known melancholy sound of the English horn due to its timbre or to how composers use it? Though music emotion recognition has received a lot of attention, researchers have only recently begun considering the relationship between emotion and timbre. To this end, we devised a listening test to compare representative tones from eight different wind and string instruments. The goal was to determine if some tones were consistently perceived as being happier or sadder in pairwise comparisons. A total of eight emotions were tested in the study. The results showed strong underlying emotional characters for each instrument. The emotions Happy, Joyful, Heroic, and Comic were strongly correlated with one another. The violin, trumpet, and clarinet best represented these emotions. Sad and Depressed were also strongly correlated. These two emotions were best represented by the horn and flute. Scary was the emotional outlier of the group, while the oboe had the most emotionally neutral timbre. Also, we found that emotional judgment correlates significantly with average spectral centroid for the more distinctive emotions, including Happy, Joyful, Sad, Depressed, and Shy. These results can provide insights in orchestration, and lay the groundwork for future studies on emotion and timbre.
Previous chapter Next chapter Full AccessProceedings Proceedings of the 2013 SIAM International Conference on Data Mining (SDM)SMART: Semi-Supervised Music Emotion Recognition with Social TaggingBin Wu, Erheng Zhong, Derek Hao Hu, Andrew Horner, and Qiang YangBin Wu, Erheng Zhong, Derek Hao Hu, Andrew Horner, and Qiang Yangpp.279 - 287Chapter DOI:https://doi.org/10.1137/1.9781611972832.31PDFBibTexSections ToolsAdd to favoritesExport CitationTrack CitationsEmail SectionsAboutAbstract Music emotion recognition (MER) aims to recognize the affective content of a piece of music, which is important for applications such as automatic soundtrack generation and music recommendation. MER is commonly formulated as a supervised learning problem. In practice, except for Pop music, there is little labeled data in most genres. In addition, emotion is genre specific in music and thus the labeled data of Pop music cannot be used for other genres. In this paper, we aim to solve the genre-specific MER problem by exploiting two kinds of auxiliary data: unlabeled songs and social tags. However, using these two kinds of data effectively is a non-trivial task, e.g. tags are noisy and therefore cannot be treated as fully trustworthy. To build an accurate model with the help from the unlabeled songs and noisy tags, we present SMART, which stands for Semi-Supervised Music Affective Emotion Recognition with Social Tagging, combining of a graph-based semi-supervised learning algorithm with a novel tag refinement method. Experiments on the Million Song Dataset show that our proposed approach, trained with only 10 labeled instances, is as accurate as Support Vector Regression trained with 750 labeled songs. Previous chapter Next chapter RelatedDetails Published:2013ISBN:978-1-61197-262-7eISBN:978-1-61197-283-2 https://doi.org/10.1137/1.9781611972832Book Series Name:ProceedingsBook Code:PRDT13Book Pages:1-804
Horror scene detection is a research problem that has much practical use. The supervised method requires the training data to be labeled manually, which can be tedious and onerous. In this paper, a more challenging setting of the problems without complete information on data labels is investigated. In particular, as the horror scene is characterized by multiple features, this problem is formulated as a special multiple instance learning (MIL) problem – Multiple Grouped Instance Learning (MGIL), which requires partial labeled training. To solve the MGIL problem, a learning method is proposed – Multiple Distance- Expectation Maximization Diversity Density (MD-EMDD).Additionally, a survey is conducted to collect people’s opinions based on the definition of horror scenes. Combined with the survey results, Labeled with Ranking – MD – EMDD is proposed and demonstrated better results when compared to the traditional MIL algorithm and close to performance achieved by supervised method.