This paper addresses the general problem of extracting music-theoretic and psychological features from symbolically encoded melodies. We review existing melodic feature extraction toolboxes, enumerate their features, and organise them into a common taxonomy. We then describe a new software library that provides implementations of all of these features in a straightforward Python package. We then demonstrate the combined feature set on the Essen Folksong Collection, using the dataset to produce a series of style classification models. These models help us answer key questions about the interpretability and dimensionality of the feature set. Our results show excellent classification accuracy using the full feature set, and promising performance for an eight-dimensional factor-analytic solution that improves the interpretability of the classifier. We distribute our new toolbox as an open-source Python package, 𝚖𝚎𝚕𝚘𝚍𝚢-𝚏𝚎𝚊𝚝𝚞𝚛𝚎𝚜, which can easily be used in various applications within music analysis, music psychology, and music information retrieval.
Artists are often recognizable through collections of distinctive patterns (‘fingerprints’) in their work. Identifying such traits has important applications in authorship attribution, education, cultural heritage research and historical analysis. Here we focus on music, a domain with a rich tradition of theoretical and mathematical analysis. We train a variety of supervised learning models to identify 20 iconic jazz musicians from a curated dataset of 84 h of recordings. In particular, we introduce a multi-input architecture that represents four musical domains separately: melody, harmony, rhythm and dynamics. This design allows us to accurately identify individual performers (our best model obtains 94% accuracy across 20 classes) and to examine which musical elements most strongly distinguish between individual artists. We release open-source implementations of our models and an accompanying web application for exploring our results. Cheston et al. develop a machine learning pipeline that identifies 20 iconic jazz pianists from audio recordings with up to 94% accuracy, revealing how melody, harmony, rhythm and dynamics shape each performer’s individual musical fingerprint.
A foundational question in empirical music aesthetics concerns how certain note combinations (chords) are perceived to be more pleasant than others. While normative patterns of chord pleasantness judgments are well studied, relatively little is known about how listeners vary in these judgments. We address this question using a set of variance decomposition techniques: structural equation modeling to disentangle preference variance from response noise, and Q-type principal component analysis to identify latent dimensions of participant variation. We apply these techniques to a new dataset where 106 online participants rated a set of 68 chords for pleasantness. Structural equation modeling shows that the chords vary in both preference variance and response noise; while response noise varies relatively unsystematically, preference variance increases with familiarity and decreases with spectral interference (i.e., beating resulting from interactions between neighboring partials). Q-type principal component analysis meanwhile shows that listener preferences vary on two primary dimensions: one corresponding to spectral interference, the other to cardinality (the number of notes in the chord). The former component is associated with musical expertise (more expertise means more disliking of interference), whereas the latter component does not correlate with any measured background variables. These findings confirm the previously established effect of musical expertise on chord preferences but contrast with previous studies' claims that responses to interference are undifferentiated across participants. Our study provides a potential model for using variance decomposition techniques to study individual differences in other aesthetic domains.
Artistic style has been studied for centuries, and recent advances in machine learning create new possibilities for understanding it computationally. However, ensuring that machine-learning models produce insights aligned with the interests of practitioners and critics remains a significant challenge. Here, we focus on musical style, which benefits from a rich theoretical and mathematical analysis tradition. We train a variety of supervised-learning models to identify 20 iconic jazz musicians across a carefully curated dataset of 84 hours of recordings, and interpret their decision-making processes. Our models include a novel multi-input architecture that enables four musical domains (melody, harmony, rhythm, and dynamics) to be analysed separately. These models enable us to address fundamental questions in music theory and also advance the state-of-the-art in music performer identification (94% accuracy across 20 classes). We release open-source implementations of our models and an accompanying web application for exploring musical styles.
In many musical styles, performing a piece of music means to produce an 'interpretation' of a score. This interpretation involves performers manipulating musical parameters such as timing, dynamics, timbre, and pitch to communicate their artistic conception of the piece, often to an audience. Much previous research into musical interpretations has examined aspects of expressive performance strategies. However, these studies have largely focused solely on the sounds produced in the performance, investigating players' manipulation of musical parameters but little of the performance's broader context and impact. Multimodal datasets, which contain multiple diverse data types offering distinct perspectives on the musical performance (e.g. audio, Musical Instrument Digital Interface, video, motion capture, physiological data), can support more holistic cross-and interdisciplinary study of performers' interpretative decision-making and its effects on audiences. We propose a taxonomy of modalities relevant to the study of musicians' interpretations of musical scores. These modalities are distinct facets of the performance or its context through which the performance and musical interpretation can be analysed (e.g. 'venue acoustics', 'performer movements', 'performance sound'). We use this taxonomy to systematically review relevant open-access multimodal datasets and the modalities they support. Underrepresented modalities are then highlighted, along with practical suggestions for including data that support these modalities in future datasets. We next examine key challenges of reporting and working with multimodal datasets, emphasising the need for standardisation of data reporting and reliable options for data storage and access. Finally, we summarise the broader interdisciplinary applications of these datasets in artificial intelligence and performance research.
The influence of room acoustic parameters on musical emotion has to a degree been studied musicologically and empirically. However, there remain large gaps related to limitations in emotion measures and aspects of acoustic setting, with various iterations of digital acoustic reproduction represented in research. This psychological study explores the ways in which systematic alterations to reverberation time (RT) may influence the emotional experience of music listening over headphones. A quantitative approach was adopted, whereby musical stimuli with parametrically altered RTs were heard over user headphones. These were compared for domain-specific musical emotions on the Geneva Emotional Music Scale (GEMS). The main findings showed that the RTs and related acoustic features did not have a strong effect on ‘Unease’ or ‘Vitality’ components of the GEMS, but rather longer RTs had a significant positive effect on aspects of ‘Sublimity’ (i.e., ‘Nostalgia’, ‘Transcendence’, ‘Wonder’). These results suggest that subjective percepts beyond pleasantness or emotional impact are affected by reverberation-based manipulations to room acoustic sound. The study outcomes have particular relevance to recorded music with artificial reverberation, and create scope for complex interactions between reverberation time and emotion more broadly.
Previous research has found that instrument-like timbres (hereafter, timbres) can affect the Goodness-of-Fit (GoF) evaluations of cadences (Vuvan & Hughes, 2019). Here, we expand these findings by testing more timbres and chord sequences and analyzing a wide array of chordal and timbral variables. One hundred and thirty-five participants with varying levels of music training provided GoF ratings for 15 C-C-X chord sequences constructed with piano, clean electric guitar, and choir timbres. The third chord was a major, minor, major-minor seventh, or minor seventh chord. The ratings of choir stimuli were higher and their range narrower than the ratings for the other two timbres, regardless of music training. Additionally, the profile formed by the GoF ratings of the 15 choir stimuli was different from the profiles of the other two timbres. Further analyses provided detailed information about the effect of timbre as well as harmonic and melodic factors on the ratings. Attack was identified as a likely contributor to GoF ratings of choir stimuli being higher than the ratings of the other stimuli. Tonal contextuality (Leman, 2000), affected by diffuse partials, and the importance given to the soprano are discussed as two plausible explanations for the narrow range and other idiosyncrasies of the GoF ratings of the choir stimuli.
A foundational question in empirical music aesthetics concerns how certain note combinations (chords) are perceived to be more pleasant than others. While normative patterns of chord pleasantness judgments are well-studied, relatively little is known about how listeners vary in these judgments. We address this question using a set of variance decomposition techniques: structural equation modelling (SEM) to disentangle preference variance from response noise, and Q-type principal component analysis (Q-PCA) to identify latent dimensions of participant variation. We apply these techniques to a new dataset where 106 online participants rated a set of 68 chords for pleasantness. SEM shows that the chords vary in both preference variance and response noise; while response noise varies relatively unsystematically, preference variance increases with familiarity and decreases with spectral interference (i.e., beating resulting from interactions between neighbouring partials). Q-PCA meanwhile shows that listener preferences vary on two primary dimensions: one corresponding to spectral interference, the other to cardinality (the number of notes in the chord). The former component is associated with musical expertise (more expertise means more disliking of interference), whereas the latter component does not correlate with any measured background variables. These findings confirm the previously established effect of musical expertise on chord preferences, but contrast with previous studies’ claims that responses to interference are undifferentiated across participants. Our study provides a potential model for using variance decomposition techniques to study individual differences in other aesthetic domains.
Aesthetic preference is intricately linked to learning and creativity. Previous studies have largely examined the perception of novelty in terms of pleasantness and the generation of novelty via creativity separately. The current study examines the connection between perception and generation of novelty in music; specifically, we investigated how pleasantness judgements and brain responses to musical notes of varying probability (estimated by a computational model of auditory expectation) are linked to learning and creativity. To facilitate learning de novo, 40 non-musicians were trained on an unfamiliar artificial music grammar. After learning, participants evaluated the pleasantness of the final notes of melodies, which varied in probability, while their EEG was recorded. They also composed their own musical pieces using the learned grammar which were subsequently assessed by experts. As expected, there was an inverted U-shaped relationship between liking and probability: participants were more likely to rate the notes with intermediate probabilities as pleasant. Further, intermediate probability notes elicited larger N100 and P200 at posterior and frontal sites, respectively, associated with prediction error processing. Crucially, individuals who produced less creative compositions preferred higher probability notes, whereas individuals who composed more creative pieces preferred notes with intermediate probability. Finally, evoked brain responses to note probability were relatively independent of learning and creativity, suggesting that these higher-level processes are not mediated by brain responses related to performance monitoring. Overall, our findings shed light on the relationship between perception and generation of novelty, offering new insights into aesthetic preference and its neural correlates.
This study investigates the effect of chord duration on the relative salience of chord-type and voicing changes. Participants ( N = 111) with varying levels of musical training were presented with sequences of five block chords on the piano and asked to indicate which chord sounded most different. Each sequence consisted of three identical chords and two oddballs, one with a voicing change and one with a chord-type change. All possible chord-type pairings between standard and oddball formed of major, minor, dominant seventh, major seventh, and minor seventh chords were tested. In addition, each sequence of five chords was tested using three chord duration conditions (500, 1,000, and 1,500 ms), and the durations were pseudo-randomized throughout the experiment. Chord-type changes became more salient with longer durations and this effect could be seen for all participants regardless of their levels of musical training. However, with higher level of musical training, chord-type changes became more salient across all duration conditions. Leman’s model of tonal contextuality suggests that the effect of duration in our experiment could be explained by sensory mechanisms related to echoic memory. The potential contribution of other factors to the effect of duration is discussed.
Recent research has confirmed that musical consonance is not only determined by the frequency ratios between tones, but also by the frequency spectra of the underlying tones (Marjieh et al., 2024). However, this prior research was limited to artificial tones, specifically tones built from a small number of pure tones, producing sounds that do not match the acoustic complexity of real musical instruments. Here we therefore investigate tones recorded from a real musical instrument, the Westerkerk Carillon, conducting a ‘dense rating’ experiment where participants (N = 113) rated musical intervals drawn from the continuous range 0-15 semitones. We show that the traditional consonances of the major third and the minor sixth become dissonances in the carillon, and we show that small intervals (in particular 0.5-2.5 semitones) also become particularly dissonant in the carillon. Through computational modelling we show that these effects are primarily caused by interference between partials (e.g. beating), but we also show that preference for harmonicity is also necessary to produce an accurate overall account of participants’ preferences. The results support musicians’ writings about the carillon and contribute to ongoing debates about the psychological mechanisms underpinning consonance perception.
Recent advances in automatic music transcription have facilitated the creation of large databases of symbolic transcriptions of improvised music forms including jazz, where traditional notated scores are not normally available. In conjunction with music source separation models that enable audio to be “demixed” into separate signals for multiple instrument classes, these algorithms can also be applied to generate annotations for every musician in a performance. This can enable the analysis of interesting performer-level and ensemble-level features that have often been difficult to explore. To this end, we introduce Jazz Trio Database (JTD), a dataset of 44.5 h of jazz piano solos accompanied by bass and drums, with automatically generated annotations for each performer. These annotations consist of onset, beat, and downbeat timestamps, alongside MIDI for the piano soloist. Suitable recordings, broadly representative of the “straight-ahead” jazz style, were identified by scraping user-based listening and discographic data; source separation models were applied to isolate audio for each performer in the trio; annotations were generated by applying appropriate algorithms to both the separated and the mixed audio sources. Onset annotations generated by the pipeline achieved a mean F-measure of 0.94 when compared with ground truth annotations. We conduct several analyses of JTD, including with relation to swing and inter-performer synchronization. We anticipate that JTD will be useful in a variety of music information–retrieval tasks, including artist identification and expressive performance modeling. We have made JTD, including the annotations and associated source code, available at https://github.com/HuwCheston/Jazz-Trio-Database
Previous psychological studies have shown that musical consonance is not only determined by the frequency ratios between tones, but also by the frequency spectra of those tones. However, these prior studies used artificial tones, specifically tones built from a small number of pure tones, which do not match the acoustic complexity of real musical instruments. The present experiment therefore investigates tones recorded from a real musical instrument, the Westerkerk Carillon, conducting a “dense rating” experiment where participants (N = 113) rated musical intervals drawn from the continuous range 0–15 semitones. Results show that the traditional consonances of the major third and the minor sixth become dissonances in the carillon and that small intervals (in particular 0.5–2.5 semitones) also become particularly dissonant. Computational modelling shows that these effects are primarily caused by interference between partials (e.g., beating), but that preference for harmonicity is also necessary to produce an accurate overall account of participants' preferences. The results support musicians' writings about the carillon and contribute to ongoing debates about the psychological mechanisms underpinning consonance perception, in particular disputing the recent claim that interference is largely irrelevant to consonance perception.
Music is present in every known society but varies from place to place. What, if anything, is universal to music cognition? We measured a signature of mental representations of rhythm in 39 participant groups in 15 countries, spanning urban societies and Indigenous populations. Listeners reproduced random ‘seed’ rhythms; their reproductions were fed back as the stimulus (as in the game of ‘telephone’), such that their biases (the prior) could be estimated from the distribution of reproductions. Every tested group showed a sparse prior with peaks at integer-ratio rhythms. However, the importance of different integer ratios varied across groups, often reflecting local musical practices. Our results suggest a common feature of music cognition: discrete rhythm ‘categories’ at small-integer ratios. These discrete representations plausibly stabilize musical systems in the face of cultural transmission but interact with culture-specific traditions to yield the diversity that is evident when mental representations are probed across many cultures.
The phenomenon of musical consonance is an essential feature in diverse musical styles. The traditional belief, supported by centuries of Western music theory and psychological studies, is that consonance derives from simple (harmonic) frequency ratios between tones and is insensitive to timbre. Here we show through five large-scale behavioral studies, comprising 235,440 human judgments from US and South Korean populations, that harmonic consonance preferences can be reshaped by timbral manipulations, even as far as to induce preferences for inharmonic intervals. We show how such effects may suggest perceptual origins for diverse scale systems ranging from the gamelan’s slendro scale to the tuning of Western mean-tone and equal-tempered scales. Through computational modeling we show that these timbral manipulations dissociate competing psychoacoustic mechanisms underlying consonance, and we derive an updated computational model combining liking of harmonicity, disliking of fast beats (roughness), and liking of slow beats. Altogether, this work showcases how large-scale behavioral experiments can inform classical questions in auditory perception.
Coordination between participants is a necessary foundation for successful human interaction. This is especially true in group musical performances, where action must often be temporally coordinated between the members of an ensemble for their performance to be effective. Networked mediation can disrupt this coordination process by introducing a delay between when a musical sound is produced and when it is received. This can result in significant deteriorations in synchrony and stability between performers. Here we show that five duos of professional jazz musicians adopt diverse strategies when confronted by the difficulties of coordinating performances over a network—difficulties that are not exclusive to networked performance but are also present in other situations (such as when coordinating performances over large physical spaces). What appear to be two alternatives involve: 1) one musician being led by the other, tracking the timings of the leader’s performance; or 2) both musicians accommodating to each other, mutually adapting their timing. During networked performance, these two strategies favor different sides of the trade-off between, respectively, tempo synchrony and stability; in the absence of delay, both achieve similar outcomes. Our research highlights how remoteness presents new complexities and challenges to successful interaction.
Great musicians have a unique style and, with training, humans can learn to distinguish between these styles. What differences between performers enable us to make such judgements? We investigate this question by building a machine learning model that predicts performer identity from data extracted automatically from an audio recording. Such a model could be trained on all kinds of musical features, but here we focus specifically on rhythm, which (unlike harmony, melody and timbre) is relevant for any musical instrument. We demonstrate that a supervised learning model trained solely on rhythmic features extracted from 300 recordings of 10 jazz pianists correctly identified the performer in 59% of cases, six times better than chance. The most important features related to a performer’s ‘feel’ (ensemble synchronization) and ‘complexity’ (information density). Further analysis revealed two clusters of performers, with those in the same cluster sharing similar rhythmic traits, and that the rhythmic style of each musician changed relatively little over the duration of their career. Our findings highlight the possibility that artificial intelligence can perform performer identification tasks normally reserved for experts. Links to each recording and the corresponding predictions are available on an interactive map to support future work in stylometry.
Expectation is crucial for our enjoyment of music, yet the underlying generative mechanisms remain unclear. While sensory models derive predictions based on local acoustic information in the auditory signal, cognitive models assume abstract knowledge of music structure acquired over the long term. To evaluate these two contrasting mechanisms, we compared simulations from four computational models of musical expectancy against subjective expectancy and pleasantness ratings of over 1000 chords sampled from 739 US Billboard pop songs. Bayesian model comparison revealed that listeners' expectancy and pleasantness ratings were predicted by the independent, non-overlapping, contributions of cognitive and sensory expectations. Furthermore, cognitive expectations explained over twice the variance in listeners' perceived surprise compared to sensory expectations, suggesting a larger relative importance of long-term representations of music structure over short-term sensory-acoustic information in musical expectancy. Our results thus emphasize the distinct, albeit complementary, roles of cognitive and sensory expectations in shaping musical pleasure, and suggest that this expectancy-driven mechanism depends on musical information represented at different levels of abstraction along the neural hierarchy. This article is part of the theme issue 'Art, aesthetics and predictive processing: theoretical and empirical perspectives'.
Music is often regarded as the 'language of the emotions' (Cooke, 1995), and since the early 20th century, empirical research into musical emotions has been conducted to explore the mystery of how they are evoked by music. In this context, Frank et al. presented an experimental study that examined the scarcely researched field of historical listening. Their study aimed to investigate the question, "Do modern listeners hear the emotional content in Baroque music that the composer intended to portray?". The results indicated that modern listeners placed the three modern excerpts in the expected quadrants of the valence-arousal space. However, there were significantly different valence and arousal ratings among the Baroque examples and the modern excerpts, and significant differences between paired examples (where Baroque and modern examples were expected to fall into the same quadrants) occurred. This commentary summarizes Frank et al.'s experimental study, discusses methodological considerations, and suggests possible refinements for future (experimental) studies on historical listening.