This study investigates how listeners reduce polyphonic drum patterns into simplified monophonic rhythm contours through tapping behavior. We examine how these contours relate to metrical structure, onset density, syncopation, and frequency-weighted onset density (FWOD), and whether these features can support rhythm retrieval tasks. Across three experiments, participants tapped rhythmic contours using a force-sensitive interface while listening to polyphonic drum patterns. The results show that tap force is shaped primarily by metrical structure, while FWOD captures an additional bottom-up component related to timbral and density-based salience. Among the analytical contour representations tested, FWOD also provided the best performance in rhythm retrieval tasks within a two-dimensional rhythm-space embedding. Together, these findings suggest that listeners simplify polyphonic rhythmic information through a combination of metrical expectations and spectrally weighted onset density. The study further supports FWOD as a compact and behaviorally relevant representation for rhythm analysis, retrieval, and interactive music systems, with potential applications in rhythm-based interaction, music production workflows, and generative musical interfaces.
Complex networks have emerged as a powerful framework for understanding and analyzing musical compositions, revealing underlying structures and dynamics that may not be immediately apparent. This article explores the application of complex network representations to the study of symbolic drum sequences, a topic that has received limited attention in the literature. The proposed methodology involves encoding drum rhythms as directed, weighted complex networks, where nodes represent drum events, and edges capture the temporal succession of these events. This network-based representation allows for the analysis of similarities between different drumming styles, as well as the generation of novel drum patterns. Through a series of experiments, we demonstrate the effectiveness of this approach. First, we show that the complex network representation can accurately classify drum patterns into their respective musical styles, even with a limited number of training samples. Second, we present a generative model based on Markov chains operating on the network structure, which is able to produce new drum patterns that retain the essential features of the training data. Finally, we validate the perceptual relevance of the generated patterns through listening tests, where participants are unable to distinguish the generated patterns from the original ones, suggesting that the network-based representation effectively captures the underlying characteristics of different drumming styles. The findings of this study have significant implications for music research, genre classification, and generative music applications, highlighting the potential of complex networks to provide a transparent and elegant approach to the analysis and synthesis of rhythmic structures in music.
This work investigates how personalised Music Emotion Recognition (MER) systems may lead to sensitive profiling when applied to musically induced emotions in politically charged contexts. We focus on traditional Colombian music with explicit political content, including (1) vallenatos and social songs aligned with the left-wing guerrilla Fuerzas Armadas Revolucionarias de Colombia (FARC), and (2) corridos linked to sympathisers of the right-wing paramilitary group Autodefensas Unidas de Colombia (AUC). Using data from 49 participants with diverse political leanings, we train personalised machine learning models to predict induced emotional responses – particularly negative emotions. Our findings reveal that political identity plays a significant role in shaping emotional experiences of music with explicit political content, and that emotion recognition models can capture this variation to a certain extent. These results raise critical concerns about the potential misuse of emotion recognition technologies. What is often framed as a tool for wellbeing and emotional regulation could, in politically sensitive contexts, be repurposed for user profiling. This work highlights the ethical risks of deploying AI-driven emotion analysis without safeguards, particularly among populations that are politically or socially vulnerable. We argue that subjective emotional responses may constitute sensitive personal data, and that failing to account for their sociopolitical context could amplify harm and exclusion.
This article details a correction to: Gómez-Marín, D., Ospina-Caicedo, R., Díaz-Cely, J., Paz, J., Jordà, S. and Herrera, P. (2024) ‘Salsa, a Dataset for Beat Estimation in Salsa Music’, Transactions of the International Society for Music Information Retrieval, 7(1), p. 264–273. Available at: https://doi.org/10.5334/tismir.183.
A pesar de su papel fundamental en la creación musical de las últimas décadas, el sintetizador no goza de un reconocimiento generalizado en las enseñanzas musicales de Cataluña, España. Este estudio examinó las prácticas educativas relacionadas con el sintetizador en las Enseñanzas Artísticas Superiores de Música de dicha región. Para ello, se llevaron a cabo entrevistas semiestructuradas a ocho docentes especializados en este instrumento. El análisis de datos se realizó siguiendo las directrices del Análisis Temático, lo que permitió identificar tres temas generales: (i) la enseñanza del sintetizador se orienta hacia objetivos creativos, (ii) el sintetizador se enseña de manera diferente dependiendo de la especialidad y el repertorio (iii) el sintetizador no goza de una modalidad instrumental propia. El artículo concluye con una discusión sobre la importancia de incorporar el sintetizador a las enseñanzas musicales para fomentar una educación más contextualizada y significativa.
This paper introduces a dataset of salsa music, a genre deeply rooted in Latin American culture that is known for its intricate yet captivating rhythms, which pose challenges even for seasoned dancers. Our work involved creating a comprehensive track selection and meticulously annotating beat occurrences. The dataset comprises 124 expert-analyzed salsa songs, offering a rich resource for further beat estimation and related studies within the salsa music domain. We detail the dataset, outline the methodology carried out for compiling and validating beat annotations, and finally test two contemporary beat prediction models on the dataset. Our contributions include the establishment of a labeled dataset for beat estimation research in salsa music and a robust methodology for identifying beat occurrences. Through this work, we aspire to enrich contemporary and future studies on Latin American culture, particularly the integral aspect of salsa music, fostering rhythm analysis and other musical properties that can derive from it.
In audiovisual contexts, different conventions determine the level at which background music is mixed into the final program, and sometimes, the mix renders the music to be practically or totally inaudible. From a perceptual point of view, the audibility of music is subject to auditory masking by other aural stimuli such as voice or additional sounds (e.g., applause, laughter, horns), and is also influenced by the visual content that accompanies the soundtrack, and by attentional and motivational factors. This situation is relevant to the music industry because, according to some copyright regulations, the non-audible background music must not generate any distribution rights, and the marginally audible background music must generate half of the standard value of audible music. In this study, we conduct two psychoacoustic experiments to identify several factors that influence background music perception, and their contribution to its variable audibility. Our experiments are based on auditory detection and chronometric tasks involving keyboard interactions with original TV content. From the collected data, we estimated a sound-to-music ratio range to define the audibility threshold limits of the barely audible class. In addition, results show that perception is affected by loudness level, listening condition, music sensitivity, and type of television content.
We present a platform and a dataset to help research on Music Emotion Recognition (MER). We developed the Music Enthusiasts platform aiming to improve the gathering and analysis of the so-called “ground truth” needed as input to MER systems. Firstly, our platform involves engaging participants using citizen science strategies and generate music emotion annotations – the platform presents didactic information and musical recommendations as incentivization, and collects data regarding demographics, mood, and language from each participant. Participants annotated each music excerpt with single free-text emotion words (in native language), distinct forced-choice emotion categories, preference, and familiarity. Additionally, participants stated the reasons for each annotation – including those distinctive of emotion perception and emotion induction. Secondly, our dataset was created for personalized MER and contains information from 181 participants, 4721 annotations, and 1161 music excerpts. To showcase the use of the dataset, we present a methodology for personalization of MER models based on active learning. The experiments show evidence that using the judgment of the crowd as prior knowledge for active learning allows for more effective personalization of MER systems for this particular dataset. Our dataset is publicly available and we invite researchers to use it for testing MER systems.
Previous research in music emotion recognition (MER) has tackled the inherent problem of subjectivity through the use of personalized models – models which predict the emotions that a particular user would perceive from music. Personalized models are trained in a supervised manner, and are tested exclusively with the annotations provided by a specific user. While past research has focused on model adaptation or reducing the amount of annotations required from a given user, we propose a methodology based on uncertainty sampling and query-by-committee, adopting prior knowledge from the agreement of human annotations as an oracle for active learning (AL). We assume that our disagreements define our personal opinions and should be considered for personalization. We use the DEAM dataset, the current benchmark dataset for MER, to pre-train our models. We then use the AMG1608 dataset, the largest MER dataset containing multiple annotations per musical excerpt, to re-train diverse machine learning models using AL and evaluate personalization. Our results suggest that our methodology can be beneficial to produce personalized classification models that exhibit different results depending on the algorithms’ complexity.
Our previous research showed promising results when transferring features learned from speech to train emotion recognition models for music. In this context, we implemented a denoising autoencoder as a pretraining approach to extract features from speech in two languages (English and Mandarin). From that, we performed transfer and multi-task learning to predict classes from the arousal-valence space of music emotion. We tested and analyzed intra-linguistic and cross-linguistic settings, depending on the language of speech and lyrics of the music. This paper presents additional investigation on our approach, which reveals that: (1) performing pretraining with speech in a mixture of languages yields similar results than for specific languages - the pretraining phase appears not to exploit particular language features, (2) the music in Mandarin dataset consistently results in poor classification performance - we found low agreement in annotations, and (3) novel methodologies for representation learning (Contrastive Predictive Coding) may exploit features from both languages (i.e., pretraining on a mixture of languages) and improve classification of music emotions in both languages. From this study we conclude that more research is still needed to understand what is actually being transferred in these type of contexts.
In this study, we address emotion recognition using unsupervised feature learning from speech data, and test its transferability to music. Our approach is to pre-train models using speech in English and Mandarin, and then fine-tune them with excerpts of music labeled with categories of emotion. Our initial hypothesis is that features automatically learned from speech should be transferable to music. Namely, we expect the intra-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in English) should result in improved performance over the cross-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in Mandarin). Our results confirm previous research on cross-domain transferability, and encourage research towards language-sensitive Music Emotion Recognition (MER) models.
Emotion is one of the main reasons why people engage and interact with music [1]. Songs can express our inner feelings, produce goosebumps, bring us to tears, share an emotional state with a composer or performer, or trigger specific memories. Interest in a deeper understanding of the relationship between music and emotion has motivated researchers from various areas of knowledge for decades [2], including computational researchers. Imagine an algorithm capable of predicting the emotions that a listener perceives in a musical piece, or one that dynamically generates music that adapts to the mood of a conversation in a film—a particularly fascinating and provocative idea. These algorithms typify music emotion recognition (MER), a computational task that attempts to automatically recognize either the emotional content in music or the emotions induced by music to the listener [3]. To do so, emotionally relevant features are extracted from music. The features are processed, evaluated, and then associated with certain emotions. MER is one of the most challenging high-level music description problems in music information retrieval (MIR), an interdisciplinary research field that focuses on the development of computational systems to help humans better understand music collections. MIR integrates concepts and methodologies from several disciplines, including music theory, music psychology, neuroscience, signal processing, and machine learning.
This work presents an initial proof of concept of how Music Emotion Recognition (MER) systems could be intentionally biased with respect to annotations of musically induced emotions in a political context. In specific, we analyze traditional Colombian music containing politically charged lyrics of two types: (1) vallenatos and social songs from the "left-wing" guerrilla Fuerzas Armadas Revolucionarias de Colombia (FARC) and (2) corridos from the "right-wing" paramilitaries Autodefensas Unidas de Colombia (AUC). We train personalized machine learning models to predict induced emotions for three users with diverse political views - we aim at identifying the songs that may induce negative emotions for a particular user, such as anger and fear. To this extent, a user's emotion judgements could be interpreted as problematizing data - subjective emotional judgments could in turn be used to influence the user in a human-centered machine learning environment. In short, highly desired "emotion regulation" applications could potentially deviate to "emotion manipulation" - the recent discredit of emotion recognition technologies might transcend ethical issues of diversity and inclusion.
Comunicacio presentada a: International Society for Music Information Retrieval Conference celebrat de l'11 al 16 d'octubre de 2020 de manera virtual.
Tagging a musical excerpt with an emotion label may result in a vague and ambivalent exercise. This subjectivity entangles several high-level music description tasks when the computational models built to address them produce predictions on the basis of a ground truth. In this study, we investigate the relationship between emotions perceived in pop and rock music (mainly in Euro-American styles) and personal characteristics from the listener, using language as a key feature. Our goal is to understand the influence of lyrics comprehension on music emotion perception and use this knowledge to improve Music Emotion Recognition (MER) models. We systematically analyze over 30K annotations of 22 musical fragments to assess the impact of individual differences on agreement, as defined by Krippendorff's coefficient. We employ personal characteristics to form group-based annotations by assembling ratings with respect to listeners' familiarity, preference, lyrics comprehension, and music sophistication. Finally, we study our group-based annotations in a two-fold approach: (1) assessing the similarity within annotations using manifold learning algorithms and unsupervised clustering, and (2) analyzing their performance by training classification models with diverse ground truths. Our results suggest that a) applying a broader categorization of taxonomies and b) using multi-label, group-based annotations based on language, can be beneficial for MER models.
This paper reports on the design and evaluation of drum rhythm spaces as interactive bi-dimensional maps used for the visualisation, retrieval and generation of drum patterns. We carry out two experiments exploring human processing of polyphonic drum patterns concluding with a list of descriptors that significantly influence similarity sensations. These features are used to build spaces based on drum pattern collections, where patterns are organised by similarity, modelled according to human perception. A drum-interpolation algorithm is introduced (and evaluated) to enhance rhythm space functionality by means of patterns that bound it, converting a discrete space to a continuous generative one.
Comunicacio presentada a: Workshop-Symposium on Research Methods in Music and Emotion celebrat el 14 de setembre de 2019 a Durham, Regne Unit.
In the present study, we address the relationship between the emotions perceived in pop and rock music (mainly in Euro-American styles with English lyrics) and the language spoken by the listener. Our goal is to understand the influence of lyrics comprehension on the perception of emotions and use this information to improve Music Emotion Recognition (MER) models. Two main research questions are addressed: 1. Are there differences and similarities between the emotions perceived in pop/rock music by listeners raised with different mother tongues? 2. Do personal characteristics have an influence on the perceived emotions for listeners of a given language? Personal characteristics include the listeners' general demographics, familiarity and preference for the fragments, and music sophistication. Our hypothesis is that inter-rater agreement (as defined by Krippendorff's alpha coefficient) from subjects is directly influenced by the comprehension of lyrics.