Singing voice transcription is a key task in Music Information Retrieval (MIR) that focuses on identifying sung notes within a music audio segment. Advancing state-of-the-art methods in this area relies heavily on high-quality data, yet annotating such data is resource-intensive and requires musical expertise. In genres like pop music, data sharing is further complicated by copyright and distribution limitations. In this paper, we refine a recently proposed data augmentation technique that leverages AI-generated music audio to address these data-related challenges. Specifically, we create musical accompaniments for vocals with known target notes, enabling the generation of new mixes that retain the original piece’s harmony while introducing substantial audio variation. Our cross-dataset experiments reveal that using harmony-matched mixes improves generalization, though performance remains below that achieved by training with additional real data.
Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature of text descriptions to guide music similarity modeling, addressing the limitations of traditional uni-modal approaches in capturing complex musical relationships. To overcome the scarcity of high-quality text-music paired data, this paper introduces a dual-source data acquisition approach combining online scraping and LLM-based prompting, where carefully designed prompts leverage LLMs' comprehensive music knowledge to generate contextually rich descriptions. Exten1sive experiments demonstrate that the proposed framework achieves significant performance improvements over existing benchmarks through objective metrics, subjective evaluations, and real-world A/B testing on the Huawei Music streaming platform.
Multi-modal deep learning techniques for matching free-form text with music have shown promising results in the field of Music Information Retrieval (MIR). Prior work is often based on large proprietary data while publicly available datasets are few and small in size. In this study, we present WikiMuTe, a new and open dataset containing rich semantic descriptions of music. The data is sourced from Wikipedia's rich catalogue of articles covering musical works. Using a dedicated text-mining pipeline, we extract both long and short-form descriptions covering a wide range of topics related to music content such as genre, style, mood, instrumentation, and tempo. To show the use of this data, we train a model that jointly learns text and audio representations. The model is evaluated on two tasks: tag-based music retrieval and music auto-tagging. The results show that while our approach has state-of-the-art performance on multiple tasks, we still observe a difference in performance depending on the data used for training.
Singing voice transcription is a very popular task in MIR which consists of obtaining the notes being sung in a given excerpt of music audio. Data is essential for state-of-the-art methods, but annotating it is labor-intensive, requiring musical expertise. Moreover, in cases such as pop music, sharing such data is even more challenging due to copyright and data distribution rights. We present in this paper a data augmentation technique leveraging AI-generated music audio to alleviate the aforementioned data-related difficulties. Specifically, we generate the music accompanying the vocals for which target notes are already known. In this way, we create new mixes that maintain the harmony of the original piece while providing large variations in the audio content. We conducted a set of cross-dataset experiments and discovered that our proposed approach with harmonic-matching mixes yields the best results compared to other augmentation strategies. The employed tools for this augmentation are open-source and ready to be used by the research community.
Multi-instrument music transcription aims to convert polyphonic music recordings into musical scores assigned to each instrument. This task is challenging for modeling as it requires simultaneously identifying multiple instruments and transcribing their pitch and precise timing, and the lack of fully annotated data adds to the training difficulties. This paper introduces Y ourMT3+, a suite of models for enhanced multi-instrument music transcription based on the recent language token decoding approach of MT3. We enhance its encoder by adopting a hierarchical attention transformer in the time-frequency domain and integrating a mixture of experts. To address data limitations, we introduce a new multi-channel decoding method for training with incomplete annotations and propose intra- and cross-stem augmentation for dataset mixing. Our experiments demonstrate direct vocal transcription capabilities, eliminating the need for voice separation pre-processors. Benchmarks across ten public datasets show our models' competitiveness with, or superiority to, existing transcription models. Further testing on pop music recordings highlights the limitations of current models. Fully reproducible code and datasets are available with demos at https://github.com/mimbres/YourMT3.
We present an analysis of large-scale pretrained deep learning models used for cross-modal (text-to-audio) retrieval. We use embeddings extracted by these models in a metric learning framework to connect matching pairs of audio and text. Shallow neural networks map the embeddings to a common dimensionality. Our system, which is an extension of our submission to the Language-based Audio Retrieval Task of the DCASE Challenge 2022, employs the RoBERTa foundation model as the text embedding extractor. A pretrained PANNs model extracts the audio embeddings. To improve the generalisation of our model, we investigate how pretraining with audio and associated noisy text collected from the online platform Freesound improves the performance of our method. Furthermore, our ablation study reveals that the proper choice of the loss function and fine-tuning the pretrained models are essential in training a competitive retrieval system.
In this paper we study the relative phase offsets between partials in the sustained part of harmonic sounds and investigate their suitability for complex matrix decomposition of spectrograms. We formally introduce this property in a sinusoidal model and visualise the phase relations of a musical instrument. A model of complex matrix decomposition in the time-frequency domain is derived and equations for the estimation of the model parameters are provided in the mono-phonic case. We illustrate the model with the analysis of a mono-phonic saxophone signal. The results suggest that the phase offset is able to capture inherent time-invariant phase properties of harmonic sounds and outline its potential use for complex matrix decomposition.
For a user-assisted music transcription system in which the user is asked to label some notes for each instrument in the recording, we investigate ways to limit the amount of information the user has to provide. Different methods are proposed and experimentally compared that enable the estimation of template spectra at pitch positions that have not been annotated by the user, in order to derive a full set of instrument templates that can be used within a non-negative matrix factorisation framework. A set of error metrics is presented that enables the evaluation of the NMF gain matrix. The results show that purely data-driven methods outperform more refined instrument models when the user annotates notes at many different pitches for each instrument. When notes are labelled at a smaller number of different pitches, the highest accuracies are obtained using pre-stored instrument templates that are adapted to the instruments in the mixture.
Automatic music transcription is considered by many to be a key enabling technology in music signal processing. However, the performance of transcription systems is still significantly below that of a human expert, and accuracies reported in recent years seem to have reached a limit, although the field is still very active. In this paper we analyse limitations of current methods and identify promising directions for future research. Current transcription methods use general purpose models which are unable to capture the rich diversity found in music signals. One way to overcome the limited performance of transcription systems is to tailor algorithms to specific use-cases. Semi-automatic approaches are another way of achieving a more reliable transcription. Also, the wealth of musical scores and corresponding audio data now available are a rich potential source of training data, via forced alignment of audio to scores, but large scale utilisation of such data has yet to be attempted. Other promising approaches include the integration of information from multiple algorithms and different musical aspects.
. We present an algorithm for tracking individual instruments in polyphonic music recordings. The algorithm takes as input the instrument identities of the recording and uses non-negative matrix factorisation to compute an instrument-independent pitch activation function. The Viterbi algorithm is applied to find the most likely path through a number of candidate instrument and pitch combinations in each time frame. The transition probability of the Viterbi algorithm includes three different criteria: the frame-wise reconstruction error of the instrument combination, a pitch continuity measure that favours similar pitches in consecutive frames, and the activity status of each instrument. The method was evaluated on mixtures of 2 to 5 instruments and outper-formed other state-of-the-art multi-instrument tracking methods.
In this paper, we address the task of semi-automatic music transcription in which the user provides prior information about the polyphonic mixture under analysis. We propose a non-negative matrix deconvolution framework for this task that allows instruments to be represented by a different basis function for each fundamental frequency (“shift variance”). Two different types of user input are studied: information about the types of instruments, which enables the use of basis functions from an instrument database, and a manual transcription of a number of notes which enables the template estimation from the data under analysis itself. Experiments are performed on a data set of mixtures of acoustical instruments up to a polyphony of five. The results confirm a significant loss in accuracy when database templates are used and show the superiority of the Kullback-Leibler divergence over the least squares error cost function.
Automatic music transcription is considered by many to be the Holy Grail in the field of music signal analysis. However, the performance of transcription systems is still significantly below that of a human expert, and accuracies reported in recent years seem to have reached a limit, although the field is still very active. In this paper we analyse limitations of current methods and identify promising directions for future research. Current transcription methods use general purpose models which are unable to capture the rich diversity found in music signals. In order to overcome the limited performance of transcription systems, algorithms have to be tailored to specific use-cases. Semi-automatic approaches are another way of achieving a more reliable transcription. Also, the wealth of musical scores and corresponding audio data now available are a rich potential source of training data, via forced alignment of audio to scores, but large scale utilisation of such data has yet to be attempted. Other promising approaches include the integration of information across different methods and musical aspects.
Audio-to-audio alignment is the task of synchronizing two audio sequences with similar musical content in time. We investigated a large set of audio features for this task. The features were chosen to represent four different content-dependent similarity categories: the envelope, the timbre, note-onsets and the pitch. The features were subjected to two processing stages. First, a feature subset was selected by evaluating the alignment performance of each individual feature. Second, the selected features were combined and subjected to an automatic weighting algorithm. A new method for the objective evaluation of audio-to-audio alignment systems is proposed that enables the use of arbitrary kinds of music as ground truth data. We evaluated our algorithm by this method as well as on a data set of real recordings of solo piano music. The results showed that the feature weighting algorithm could improve the alignment accuracies compared to the results of the individual features.