Onset detection is the process of identifying the start points of musical note events within an audio recording. While the detection of percussive onsets is often considered a solved problem, soft onsets-as found in string instrument recordings-still pose a significant challenge for state-of-the-art algorithms. The problem is further exacerbated by a paucity of data containing expert annotations and research related to best practices for curating soft onset annotations for string instruments. To this end, we investigate inter-annotator agreement between 24 participants, extend an algorithm for determining the most consistent annotator, and compare the performance of human annotators and state-of-the-art onset detection algorithms. Experimental results reveal a positive trend between musical experience and both inter-annotator agreement and performance in comparison with automated systems. Additionally, onsets produced by changes in fingering as well as those from the cello were found to be particularly challenging for both human annotators and automatic approaches. To promote research in best practices for annotation of soft onsets, we have made all experimental data associated with this study publicly available. In addition, we publish the ARME Virtuoso Strings dataset, consisting of over 144 recordings of professional performances of an excerpt from Haydn's string quartet Op. 74 No. 1 Finale, each with corresponding individual instrumental onset annotations.
This meeting report gives an overview of the DAFx 2019 conference held in September 2019 at Birmingham City University, Birmingham, UK. The conference had the same theme as this special issue: digital audio effects. In total, 51 papers were presented at DAFx 2019 either in oral or in poster sessions. The conference had 157 delegates, almost half from industry and the rest from universities around the world. As the number of submissions and participants remains sufficiently high, it is planned that the DAFx conference series will be continued every autumn.
Intelligent Music Production presents the state of the art in approaches, methodologies and systems from the emerging field of automation in music mixing and mastering. This book collects the relevant works in the domain of innovation in music production, and orders them in a way that outlines the way forward: first, covering our knowledge of the music production processes; then by reviewing the methodologies in classification, data collection and perceptual evaluation; and finally by presenting recent advances on introducing intelligence in audio effects, sound engineering processes and music production interfaces. Intelligent Music Production is a comprehensive guide, providing an introductory read for beginners, as well as a crucial reference point for experienced researchers, producers, engineers and developers.
Accurate estimation of note onset timing is important for music ensemble performance analysis and synthesis. In this study, we present a method for the detection of onsets from polyphonic mixtures, using score information. First, a MIDI score is aligned to the audio signal using dynamic time warping, and pitches of performed notes are refined using a multi-pitch estimation technique. Notes in a signal are then isolated using a spectral masking method, based on the average harmonic structure learned from each source. Onset timing is finally estimated by maximizing the time derivative of the energy curve of the note within an observation window. We show that this method significantly improves the onset timing estimation accuracy, measured by both the align rate and onset time deviation, and outperforms a state-of-art reference method.
The majority of state-of-the-art methods for music infor-mation retrieval (MIR) tasks now utilise deep learningmethods reliant on minimisation of loss functions such ascross entropy. For tasks that include framewise binaryclassification (e.g., onset detection, music transcription)classes are derived from output activation functions byidentifying points of local maxima, or peaks. However, theoperating principles behind peak picking are different tothat of the cross entropy loss function, which minimises theabsolute difference between the output and target valuesfor a single frame. To generate activation functions moresuited to peak-picking, we propose two versions of a newloss function that incorporates information from multipletime-steps: 1)multi-individual, which uses multiple indi-vidual time-step cross entropies; and 2)multi-difference,which directly compares the difference between sequentialtime-step outputs. We evaluate the newly proposed lossfunctions alongside standard cross entropy in the popularMIR tasks of onset detection and automatic drum tran-scription. The results highlight the effectiveness of theseloss functions in the improvement of overall system ac-curacies for both MIR tasks. Additionally, directly com-paring the output from sequential time-steps in the multi-difference approach achieves the highest performance.
In this study we investigate ways in which data sonification can improve standard data analysis techniques currently employed in the analysis of stem-cells using Fourier Transform Infrared (FTIR) Spectroscopy. Four different sonification methods have been evaluated and their effectiveness has been evaluated through listening tests, designed to assess the discriminating capability of the auditory technique. We identify FM synthesis driven by feature extraction as the most perceptually relevant technique for the auditory classification of FTIR data. Whilst this technique is not commonly used in sonification research, it allows us to utilise the most salient characteristics of the absorption spectra, leading to an improved classification accuracy with a clear timbral differences between differentiated and non-differentiated cell-types.