Hand tracking is a critical component of natural user interactions in extended reality (XR) environments, including extended reality musical instruments (XRMIs). However, self-occlusion remains a significant challenge for vision-based hand tracking systems, leading to inaccurate results and degraded user experiences. In this paper, we propose a multimodal hand tracking system that combines vision-based hand tracking with surface electromyography (sEMG) data for finger joint angle estimation. We validate the effectiveness of our system through a series of hand pose tasks designed to cover a wide range of gestures, including those prone to self-occlusion. By comparing the performance of our multimodal system to a baseline vision-based tracking method, we demonstrate that our multimodal approach significantly improves tracking accuracy for several finger joints prone to self-occlusion. These findings suggest that our system has the potential to enhance XR experiences by providing more accurate and robust hand tracking, especially in the presence of self-occlusion.
In recent years, the guitar has received increased attention from the music information retrieval (MIR) community driven by the challenges posed by its diverse playing techniques and sonic characteristics. Mainly fueled by deep learning approaches, progress has been limited by the scarcity and limited annotations of datasets. To address this, we present the Guitar On Audio and Tablatures (GOAT) dataset, comprising 5.9 hours of unique high-quality direct input audio recordings of electric guitars from a variety of different guitars and players. We also present an effective data augmentation strategy using guitar amplifiers which delivers near-unlimited tonal variety, of which we provide a starting 29.5 hours of audio. Each recording is annotated using guitar tablatures, a guitar-specific symbolic format supporting string and fret numbers, as well as numerous playing techniques. For this we utilise both the Guitar Pro format, a software for tablature playback and editing, and a text-like token encoding. Furthermore, we present competitive results using GOAT for MIDI transcription and preliminary results for a novel approach to automatic guitar tablature transcription. We hope that GOAT opens up the possibilities to train novel models on a wide variety of guitar-related MIR tasks, from synthesis to transcription to playing technique detection.
Music is the art of arranging sounds in time so as to produce a continuous, unified, and evocative composition. Electronic dance music (EDM) is a collection of musical sub-genres produced using computers and electronic instruments and often presented through the medium of DJing, where tracks are curated and mixed sequentially into a continuous stream of music to offer unique listening and dancing experiences over time periods ranging from several minutes to several hours. A DJ's actions and decisions occur at several levels of temporal granularity, from real-time audio manipulation (e.g. of tempo) for smooth inter-track transitions to long-term planning of track selection and sequencing for mix content and flow. While human DJs can instinctively operate across these different temporal resolutions, replicating this capability in an end-to-end automated DJing system presents significant challenges. In this paper, we analyse existing works in DJ mix information retrieval and generation from this temporal perspective. We first explain the close link between DJing and the temporal notion of musical rhythm, then describe a framework for categorising DJing actions by temporal granularity. Using this framework, we summarise and contrast potential approaches for automating and augmenting sequential DJ decision making, and discuss the unique characteristics of DJ mix track selection as a sequential recommendation task. In doing so, we hope to facilitate the implementation of more robust and complete automated DJing systems in future research. 2012 ACM Subject Classification Applied computing -> Sound and music computing; Computing methodologies -> Control methods; Computing methodologies -> Planning under uncertainty
Music generation in the audio domain using artificial intelligence (AI) has witnessed steady progress in recent years. However for some instruments, particularly the guitar, controllable instrument synthesis remains limited in expressivity. We introduce GuitarFlow, a model designed specifically for electric guitar synthesis. The generative process is guided using tablatures, an ubiquitous and intuitive guitar-specific symbolic format. The tablature format easily represents guitar-specific playing techniques (e.g. bends, muted strings and legatos), which are more difficult to represent in other common music notation formats such as MIDI. Our model relies on an intermediary step of first rendering the tablature to audio using a simple sample-based virtual instrument, then performing style transfer using Flow Matching in order to transform the virtual instrument audio into more realistic sounding examples. This results in a model that is quick to train and to perform inference, requiring less than 6 hours of training data. We present the results of objective evaluation metrics, together with a listening test, in which we show significant improvement in the realism of the generated guitar audio from tablatures.
This paper explores changes in a novel graph-structured corpus of British folk music across time, to discover whether any evidence for evolution can be found. Feature-based approaches are compared to graph-neural network (GNN) models. Firstly, a large dataset of over 13,000 dated British folk tunes is collected and pitch and rhythm vectors are extracted. This dataset is made publicly available for future research (Dataset: https://zenodo.org/records/14692025 ). 1000 tunes are considered henceforth in two datasets of even class distribution, with tunes grouped into 50-year or 25-year time periods respectively. Exploratory analysis is undertaken with the K-means algorithm on pitch and rhythm musicological descriptors extracted from each tune, revealing ill-defined clusters with significant overlap, implying that any differences across time periods are more high-dimensional than can be represented by simple features. However, clusters do seem to indicate that there are broad differences across 100-year periods. A graph based upon similarity between tunes is then constructed by creating edges using Euclidean distance between tune-vectors. Louvain community detection illustrates ill-defined communities in terms of time-period, with no clear evolutionary trends. Graph Convolutional Neural Network (GCN) and GraphSAGE models are trained on the two datasets and are found to perform above chance for detecting the time period of tunes. Peak performance reaches 57
The Musical Instrument Digital Interface (MIDI), introduced in 1983, revolutionized music production by allowing computers and instruments to communicate efficiently. MIDI files encode musical instructions compactly, facilitating convenient music sharing. They benefit music information retrieval (MIR), aiding in research on music understanding, computational musicology, and generative music. The GigaMIDI dataset contains over 1.4 million unique MIDI files, encompassing 1.8 billion MIDI note events and over 5.3 million MIDI tracks. GigaMIDI is currently the largest collection of symbolic music in MIDI format available for research purposes under fair dealing. Distinguishing between non‑expressive and expressive MIDI tracks is challenging, as MIDI files do not inherently make this distinction. To address this issue, we introduce a set of innovative heuristics for detecting expressive music performance. These include the distinctive note velocity ratio (DNVR) heuristic, which analyzes MIDI note velocity; the distinctive note onset deviation ratio (DNODR) heuristic, which examines deviations in note onset times; and the note onset median metric level (NOMML) heuristic, which evaluates onset positions relative to metric levels. Our evaluation demonstrates these heuristics effectively differentiate between non‑expressive and expressive MIDI tracks. Furthermore, after evaluation, we create the most substantial expressive MIDI dataset, employing our heuristic NOMML. This curated iteration of GigaMIDI encompasses expressively performed instrument tracks detected by NOMML, containing all General MIDI instruments, constituting 31% of the GigaMIDI dataset, totaling 1,655,649 tracks.
Loopable music generation systems enable diverse applications, but they often lack controllability and customization capabilities. We argue that enhancing controllability can enrich these models, with emotional expression being a crucial aspect for both creators and listeners. Hence, building upon LooperGP, a loopable tablature generation model, this paper explores endowing systems with control over conveyed emotions. To enable such conditional generation, we propose integrating musical knowledge by utilizing multi-granular semantic and musical features during model training and inference. Specifically, we incorporate song-level features (Emotion Labels, Tempo, and Mode) and bar-level features (Tonal Tension) together to guide emotional expression. Through algorithmic and human evaluations, we demonstrate the approach's effectiveness in producing music conveying two contrasting target emotions, happiness and sadness. An ablation study is also conducted to clarify the contributing factors behind our approach's results.
Generative AI models have recently blossomed, significantly impacting artistic and musical traditions. Research investigating how humans interact with and deem these models is therefore crucial. Through a listening and reflection study, we explore participants' perspectives on AI- vs human-generated progressive metal, in symbolic format, using rock music as a control group. AI-generated examples were produced by ProgGP, a Transformer-based model. We propose a mixed methods approach to assess the effects of generation type (human vs. AI), genre (progressive metal vs. rock), and curation process (random vs. cherry-picked). This combines quantitative feedback on genre congruence, preference, creativity, consistency, playability, humanness, and repeatability, and qualitative feedback to provide insights into listeners' experiences. A total of 32 progressive metal fans completed the study. Our findings validate the use of fine-tuning to achieve genre-specific specialization in AI music generation, as listeners could distinguish between AI-generated rock and progressive metal. Despite some AI-generated excerpts receiving similar ratings to human music, listeners exhibited a preference for human compositions. Thematic analysis identified key features for genre and AI vs. human distinctions. Finally, we consider the ethical implications of our work in promoting musical data diversity within MIR research by focusing on an under-explored genre.
This work is part of a larger project exploring how affective computing can support the design of player-adaptive video games. We investigate how controlling some of the game mechanics using biofeedback affects physiological reactions, performance, and the experience of the player. More specifically, we assess how different game speeds affect player physiological responses and game performance. We developed a game prototype with Unity1 which includes a biofeedback loop system based on the level of physiological activation through skin resistance (SKR) measured with a smart wristband. In two conditions, the player moving speed was driven by SKR, to increase (respectively decrease) speed when the player is less activated (SKR decreases). A control condition was also used where player speed is not affected by SKR. We collected and synchronized biosignals (heart rate [HR], skin temperature [SKT] and SKR), and game information, such as the total time to complete a level, the number of ennemy collisions, and their timestamps. Additionally, emotional profiling (TIPI, I-Panas-SF), measured using a Likert scale in a post-task questionnaire, and semi-open questions about the game experience were used. The results show that SKR was significantly higher in the speed down condition, and game performance improved in the speed up condition. Study collected data involved 13 participants (10 males, 3 females) aged from 18 to 50 (M = 24.30, SD = 9.00). Most of the participants felt engaged with the game (M = 6.46, SD = 0.96) and their level of immersion was not affected by wearing the prototype smartband. Thematic analysis (TA) revealed that the game speed impacted the participants stress levels such as high speed was more stressful than hypothesized; many participants described game level-specific effects in which they felt that their speed of movement reflected their level of stress or relaxation. Slowing down the participants indeed increased the participant stress levels, but counter intuitively, more stress was detected in high speed situations.
A broad variety of audio content is available online through an increasing number of repositories and platforms. Resources such as music tracks, recorded sounds or instrument samples may be accessed by users for tasks ranging from customised music listening and exploration, to music making and sound design using existing sounds and samples. However, each online repository offers its own API and represents information through its own data model, making it difficult for applications to exploit the plurality of online audio and music content on the web. A crucial step toward integrating audio repositories in a flexible manner is a shared basis for modelling the data therein. This paper describes and extends the Audio Commons Ontology, a common data model designed to integrate existing repositories in the audio media domain. The ontology is designed with the involvement of users through surveys and requirements analyses, and evaluated in-use, by demonstrating how it supports the integration of four relevant repositories with heterogeneous APIs and data models. While this work proves the concept in the audio domain, our proposed methodology may transfer across a broad range of media integration tasks.
This paper describes the design and evaluation of Netz, a novel mixed reality musical instrument that leverages artificial intelligence for reducing errors in gesture interpretation by the system. We followed a participatory design approach over three months through regular sessions with a professional musician. We explain our design process and discuss technological sensing errors in mixed reality devices, which emerged during the design sessions. We investigate the use of interactive machine learning techniques to mitigate such errors. Results from statistical analyses indicate that a deep learning model based on interactive machine learning can significantly reduce the number of technological errors in a set of musical performance tasks with the mixed reality musical instrument. Based on our findings, we argue that the application of interactive machine learning techniques can be beneficial for embodied, hand-controlled musical instruments in the mixed reality domain.
Hand tracking is a critical component of natural user interactions in extended reality (XR) environments, including extended reality musical instruments (XRMIs). However, self-occlusion remains a significant challenge for vision-based hand tracking systems, leading to inaccurate results and degraded user experiences. In this paper, we propose a multimodal hand tracking system that combines vision-based hand tracking with surface electromyography (sEMG) data for finger joint angle estimation. We validate the effectiveness of our system through a series of hand pose tasks designed to cover a wide range of gestures, including those prone to self-occlusion. By comparing the performance of our multimodal system to a baseline vision-based tracking method, we demonstrate that our multimodal approach significantly improves tracking accuracy for several finger joints prone to self-occlusion. These findings suggest that our system has the potential to enhance XR experiences by providing more accurate and robust hand tracking, even in the presence of self-occlusion.
VR could transform creative engagement with spatial audio, given affordances for spatial visualisation and embodied interaction. But, issues exist addressing how to support collaboration for spatial audio production (SAP). Exploring this problem, we made a VR voice-based trajectory sketching tool, named Invoke, that allows two users to shape sonic ideas together. In this paper, thematic analysis is used to review two areas of a formative evaluation with expert users: (i) video analysis of VR interactions; and (ii) analysis of open questions about using the tool. Implications present new opportunities to explore co-creative VR tools for SAP.
Recent work in the field of symbolic music generation has shown value in using a tokenization based on the GuitarPro format, a symbolic representation supporting guitar expressive attributes, as an input and output representation. We extend this work by fine-tuning a pre-trained Transformer model on ProgGP, a custom dataset of 173 progressive metal songs, for the purposes of creating compositions from that genre through a human-AI partnership. Our model is able to generate multiple guitar, bass guitar, drums, piano and orchestral parts. We examine the validity of the generated music using a mixed methods approach by combining quantitative analyses following a computational musicology paradigm and qualitative analyses following a practice-based research paradigm. Finally, we demonstrate the value of the model by using it as a tool to create a progressive metal song, fully produced and mixed by a human metal producer based on AI-generated music.
Musical professionals who produce material for non-musical stakeholders often face communication challenges in the early ideation stage. Expressing musical ideas can be difficult, especially when domain-specific vocabulary is lacking. This position paper proposes the use of artificial intelligence to facilitate communication between stakeholders and accelerate the consensus-building process. Rather than fully or partially automating the creative process, the aim is to give more time for creativity by reducing time spent on defining the expected outcome. To demonstrate this point, the paper discusses two application scenarios for interactive music systems that are based on the authors' research into gesture-to-sound mapping.
Current music emotion recognition (MER) systems rely on emotion data averaged across listeners and over time to infer the emotion expressed by a musical piece, often neglecting time- and listener-dependent factors. These limitations can restrict the efficacy of MER systems and cause misjudgements. We present two exploratory studies on music emotion perception. First, in a live music concert setting, fifteen audience members annotated perceived emotion in the valence-arousal space over time using a mobile application. Analyses of inter-rater reliability yielded widely varying levels of agreement in the perceived emotions. A follow-up lab-based study to uncover the reasons for such variability was conducted, where twenty-one participants annotated their perceived emotions whilst viewing and listening to a video recording of the original performance and offered open-ended explanations. Thematic analysis revealed salient features and interpretations that help describe the cognitive processes underlying music emotion perception. Some of the results confirm known findings of music perception and MER studies. Novel findings highlight the importance of less frequently discussed musical attributes, such as musical structure, performer expression, and stage setting, as perceived across audio and visual modalities. Musicians are found to attribute emotion change to musical harmony, structure, and performance technique more than non-musicians. We suggest that accounting for such listener-informed music features can benefit MER in helping to address variability in emotion perception by providing reasons for listener similarities and idiosyncrasies.
Artificial Intelligence (AI) technologies such as deep learning are evolving very quickly bringing many changes to our everyday lives. To explore the future impact and potential of AI in the field of music and sound technologies a doctoral day was held between Queen Mary University of London (QMUL, UK) and Sciences et Technologies de la Musique et du Son (STMS, France). Prompt questions about current trends in AI and music were generated by academics from QMUL and STMS. Students from the two institutions then debated these questions. This report presents a summary of the student debates on the topics of: Data, Impact, and the Environment; Responsible Innovation and Creative Practice; Creativity and Bias; and From Tools to the Singularity. The students represent the future generation of AI and music researchers. The academics represent the incumbent establishment. The student debates reported here capture visions, dreams, concerns, uncertainties, and contentious issues for the future of AI and music as the establishment is rightfully challenged by the next generation.
Recently, symbolic music generation with deep learning techniques has witnessed steady improvements. Most works on this topic focus on MIDI representations, but less attention has been paid to symbolic music generation using guitar tablatures (tabs) which can be used to encode multiple instruments. Tabs include information on expressive techniques and fingerings for fretted string instruments in addition to rhythm and pitch. In this work, we use the DadaGP dataset for guitar tab music generation, a corpus of over 26k songs in GuitarPro and token formats. We introduce methods to condition a Transformer-XL deep learning model to generate guitar tabs (GTR-CTRL) based on desired instrumentation (inst-CTRL) and genre (genre-CTRL). Special control tokens are appended at the beginning of each song in the training corpus. We assess the performance of the model with and without conditioning. We propose instrument presence metrics to assess the inst-CTRL model's response to a given instrumentation prompt. We trained a BERT model for downstream genre classification and used it to assess the results obtained with the genre-CTRL model. Statistical analyses evidence significant differences between the conditioned and unconditioned models. Overall, results indicate that the GTR-CTRL methods provide more flexibility and control for guitar-focused symbolic music generation than an unconditioned model.
This paper presents an application of affective conditional modifiers (ACMs) in adaptive video game music – a technique whereby the emotional intent of background music is adapted, based on biofeedback, to enforce a target emotion state in the player, thus providing a more immersive experience. The proposed methods are explored in a bespoke horror game titled "The Hidden", which uses ACMs to enforce states of calmness in stressed players, and states of stress in calm players, through the procedural adaptation of background music timbre and instrumentation. These two conditions, along with a control condition, are investigated through an experimental study. Due to the low number of participants, the results of the user study provide limited insight into the effectiveness of the proposed ACMs. Nevertheless, the experiment design and user feedback highlight a number of important considerations and potential directions for future work. Namely, the need for consideration of the individual affective profile of the player, the audio-visual and narrative cues that may reduce the impact of affective audio, the effects of game familiarity on affective responses, and the need for ACM thresholds that are well-suited to the context and narrative of the game.
Real-time music information retrieval (RT-MIR) has much potential to augment the capabilities of traditional acoustic instruments. We develop RT-MIR techniques aimed at augmenting percussive fingerstyle, which blends acoustic guitar playing with guitar body percussion. We formulate several design objectives for RT-MIR systems for augmented instrument performance: (i) causal constraint, (ii) perceptually negligible action-to-sound latency, (iii) control intimacy support, (iv) synthesis control support. We present and evaluate real-time guitar body percussion recognition and embedding learning techniques based on convolutional neural networks (CNNs) and CNNs jointly trained with variational autoencoders (VAEs). We introduce a taxonomy of guitar body percussion based on hand part and location. We follow a cross-dataset evaluation approach by collecting three datasets labelled according to the taxonomy. The embedding quality of the models is assessed using KL-Divergence across distributions corresponding to different taxonomic classes. Results indicate that the networks are strong classifiers especially in a simplified 2-class recognition task, and the VAEs yield improved class separation compared to CNNs as evidenced by increased KL-Divergence across distributions. We argue that the VAE embedding quality could support control intimacy and rich interaction when the latent space's parameters are used to control an external synthesis engine. Further design challenges around generalisation to different datasets have been identified.
Elaine Chew合作论文数Daniel J. Epstein Department of Industrial and Systems Engineering;Integrated Media Systems Center;Ming Hsieh Department of Electrical Engineering (joint appointmt);USC Andrew and Erna Viterbi School of Engineering6
Philippe Depalle合作论文数Sound Processing and Control Lab, Department of Music Research, Schulich School of Music, McGill University3