
This article presents a systematic review of technologically expanded acoustic instruments, focusing on systems that expand traditional Western musical instruments without compromising their acoustic identity. Drawing from 40 peer-reviewed publications (2015-2025), the study analyses the technological, performative and evaluative strategies used to mediate instrumental interaction. A comprehensive categorisation framework is proposed, covering types of expansions, interaction modalities, technologies, software, performance contexts and evaluation methods. The findings highlight a predominance of gesture-based multimodal systems, real-time audio processing with Max/MSP and a strong emphasis on expressivity and performer-system co-agency. While diverse in technical implementation, these systems share common goals: enhancing expressivity, redefining performer-instrument relationships and expanding the boundaries of acoustic performance through interactive media. However, the review also reveals a lack of unified evaluation methodologies and a conceptual fragmentation that limits cross-study synthesis. The article contributes to the epistemological structuring of the field by proposing an operational framework for comparing, documenting and evaluating technologically expanded acoustic instruments and positions its findings as a foundation for future practice-based and theoretical research. This study contributes to the consolidation of a rapidly evolving domain at the intersection of performance, digital lutherie and interaction design.
This article presents a research platform for musicians' interaction on room stages. It is intended for stage acoustical quality and psychophysical experiments on music performance including considerations for hearing one's own and others' instruments. The platform is a real-time auralisation system; hence, its quality depends on the system's calibration and latency, among other factors. Therefore, the presented system focuses on the acoustic considerations for laboratory implementations. The calibration is implemented as a set of filters accounting for the microphone-instrument distances and the directivity factors, as well as the transducers' frequency responses. Moreover, sources of errors are characterised using both state-of-the-art information and derivations from the mathematical definition of the calibration filter. In order to compensate for hardware latency without cropping parts of the simulated impulse responses, the virtual direct sound of musicians hearing themselves is skipped from the simulation and addressed by letting the actual direct sound reach the listener through open headphones. The required latency compensation meets both the minimum distance requirement between musicians for the interactive part (i.e. hearing others) and the minimum ray path requirement for the floor reflection for the individual part (i.e. hearing oneself), which is 2 m for the implemented system. Finally, a proof of concept is provided that includes objective and subjective experiments, which give support to the feasibility of the proposed setup.
Music performance is one of the most refined forms of skilled human behaviour, combining high-level motor and cognitive demands with expressive intentions. We examined how musical features and movement expression affect postural sway in expert saxophone players. Twenty participants (nine female) performed excerpts varying in tempo, rhythmical density, articulation, and technical demands, in standing position, under movement-restricted and expressive-movement conditions. Generalised linear mixed models were used to assess the effects of these factors on centre-of-mass measurements. Results showed that participants swayed faster and travelled longer distances in music with faster tempo and increased rhythmical density, but slower and with reduced mediolateral range when performing tonguing technique (staccato). When limited to technical motion, participants still showed increased mediolateral sway during staccato passages, suggesting compensatory postural adjustments. Sway frequency was unaffected by movement condition, possibly reflecting an unconscious, task-related motor response. Performances were longer when movement was constrained, highlighting the role of body motion in temporal regulation. Our findings help understand how saxophone players accommodate technical and expressive goals, offering new insights into motor control and multisensory integration during performance.
Traditional sound synthesis methods offer distinct advantages but often face challenges in either computational efficiency or realism. Nonlinear dynamical systems present a compelling middle ground, especially with recent advances in the numerical integration of ordinary differential equations, allowing efficient implementation while capturing the essential dynamics of acoustic instruments. Here, we explore the application of nonlinear dynamical systems, particularly the theory of normal forms, to sound synthesis, offering a novel perspective for creating virtual instruments. The theory of normal forms provides a rigorous mathematical framework for understanding how oscillations can be generated, destroyed, or how they interact in nonlinear systems, making it a powerful tool in sound design. Despite its potential, this approach remains underexplored in the context of sound synthesis. We present an overview of relevant concepts in dynamical systems such as phase space, fixed points, and bifurcations. Three specific dynamical systems are then examined to analyze their suitability for sound synthesis. Finally, we present a package for the Max programming environment for real-time synthesis and show that it is possible to develop virtual instruments based on the examples presented before.
Theories of accent production lack empirical data supporting how notes are marked as accented in performance. The present study used percussion performance as a lens to investigate how movements of the performer's mallets, hands, and wrists contribute to marking notes as accented and how such movements may be refined via training. Experienced percussionists completed a single experimental session where they practiced an excerpt scored for multiple drums. During training, an instructor provided theoretically informed coaching prompts designed to improve performance quality and effectiveness of accent production between successive performances of the excerpt. Motion capture technology measured movements of the mallets, hands, and wrists along specific phases of the accent trajectory: the preparatory upstroke, accent downstroke, post-accent upstroke, and following-note downstroke. Analyses revealed that at post-training, the position of both mallets during the accent downstroke were higher than pre-training. Mallet velocity was also greater in post-training vs. pre-training for the accent downstroke. The average hand position was higher in post-training vs. pre-training for the preparatory upstroke. These changes in movement kinematics coincided with increased effectiveness of accent production in post- vs. pre-training performances, as evaluated by trained judges. The results were interpreted with regards to unique mechanisms that give rise to accent production in relation to mallet and upper-limb movements, and how such mechanisms can be applied towards translating theory into practice for improving accent production in percussion.
Discovering the emotional-semantic dimensions underlying music description is central to music psychology and widely applied in sonic branding practice. Academic work typically relies on dimension reduction approaches such as principal components analysis (PCA), which (i) require subjective reinterpretation of latent components, (ii) often yield uneven component importances when a fixed number of dimensions is imposed, and (iii) offer limited guidance for selecting practically manageable subsets of descriptors. Addressing this gap, we evaluate whether an existing feature-selection algorithm - Diversity-Induced Self-Representation (D-ISR; Liu et al. [2017]) - can serve as an objective and scalable alternative for identifying concise yet representative sets of emotional-semantic attributes. Using a large real-world dataset (NParticipants = 55,593; NResponses = 5,820,188; NAudioTracks = 251), we compare D-ISR and PCA within a unified experimental framework. D-ISR selects 14 core attributes from an industry-scale pool of 212 attributes and reconstructs the original 212-dimensional space with good accuracy. Direct comparison with PCA demonstrates how D-ISR provides a more balanced trade-off between interpretability, reconstruction fidelity, and the need for a practically small set of descriptors. Our findings (i) document and analyse a large-scale emotional-semantic music dataset rarely accessible in the public domain, (ii) demonstrate a principled framework for comparing feature-selection and component-extraction methods for music-descriptor research, and (iii) illustrate how large attribute sets can be reduced to flexible, task-appropriate representations of the emotional-semantic music space. This contributes a clear methodological foundation for both scientific studies of musical meaning and applied work such as sonic branding.
Musical stimuli provide a valuable approach to understanding and eliciting emotions. This research aimed to evaluate the validity and reliability of a musical stimuli set, composed to elicit happiness, fear/anger, sadness, and serenity, developing a novel tool for non-verbal emotion assessment. Furthermore, a comparative analysis explored variations in assessing perceived emotions, valence, and arousal dimensions among musicians and non-musicians. A sample of 200 individuals, aged 18 to 43 years (M = 27.27; SD = 5.90), and comprising 100 musicians and 100 non-musicians, evaluated the stimuli by reporting their perceived emotions through a forced-choice identification task, and by rating valence and arousal using the Self-Assessment Manikin. Among the original 116 stimuli designed, 38 musical stimuli were selected based on item response theory and confirmatory factor analysis for developing the Musical Emotion Evaluation Test. Additionally, musicians and non-musicians performed similarly when identifying emotions and perceiving emotional valence and arousal in musical stimuli. The validation of musical stimuli to evaluate emotions contributes to increasing the reliability and validity of the assessment in the research and clinical contexts. This study also provides insights into how musical expertise influences the perception of emotions in music.
Music Information Retrieval (MIR) is a rapidly advancing field dedicated to the scientific analysis of music across different cultures. This research explores the application of an Error-Correcting Output Codes (ECOC) framework combined with a Multilayer Perceptron (MLP) for the classification of Iranian classical music, an area that has been less explored in computational musicology. The novel approach leverages the NAVA dataset, a new standard in MIR specifically structured for this task, to address the challenge of distinguishing complex musical modes known as Dastgahs, which are integral to Iranian music. The proposed model reduces computational complexity while achieving higher classification accuracy compared to traditional deep learning models. This study demonstrates the potential of MLP-ECOC models to handle high feature dimensionality and various musical complexities effectively. This research contributes to the broader application of MIR technologies in non-Western musical contexts, aiming to enhance personalised music discovery and provide a deeper understanding of the cultural heritage embedded in music. This work sets a new standard in the field, promoting simpler, yet effective, ensemble learning techniques in music classification with a highest recorded accuracy of 96.32%.
This paper proposes group delay-based schemes for automatic tonic identification in Indian and Turkish classical music. Modified group delay functions (MODGD) have been employed for the proposed task. The tonic in Indian art music is the base pitch chosen by a performer in a concert. Tonic in Turkish music is one of the stable pitches of the performance, which serves as the reference throughout the performance. In the proposed work, tonic pitch in Indian art music is estimated by analysing the drone instrument in the non-melodic audio segments (segments with no lead melody line). Two schemes have been introduced for Indian art music. While the first scheme employs melodic pitch as an attribute to identify non-melodic segments, the second relies on audio descriptors. Modified group delay functions computed from the flattened music spectrum of these frames are bin-wise summed to form a summary-MODGD-gram. The location of the peak in the tonic range of summary-MODGD-gram is mapped to the tonic pitch. The experiment is also extended to estimate the tonic in Turkish makam music. The performance is evaluated on the Indian and Turkish classical music datasets, and the results show the potential of MODGD in tonic pitch estimation.
While several perceptual procedures exist to assess pitch accuracy of singing, computer-assisted methods are scarce. We developed and tested a method that combines automated frequency analysis with manual note segmentation. PitchAnalyzer2.2 software was used to extract the fundamental frequency and the musical pitches of the audio samples. The onsets, offsets, and steady-state sections of the sung syllables were detected manually to identify the keys or root notes and modulations that were dominant in a sample. Finally, the difference between the singing-sample melody and the target melody was computed to determine the overall scores for each performance. We used samples obtained from Chinese primary school children to test the new method. The samples were collected before and after the children participated in an extracurricular training programme on music versus foreign-language learning. Replicating previous findings, girls outperformed boys in singing accuracy. There were no effects of training programme or age. Testing the new method for analysis of singing-pitch accuracy for a target melody revealed that it is appropriate to generate overall performance scores for samples of children's singing, particularly when their singing sample contains multiple modulations.
Pytakt is a new Python library for the text-based description, algorithmic generation, and real-time processing of symbolic (event-level) music information. The library provides embedded textual description that makes it possible to represent scores containing chords, polyphony, and performance information, such as velocity or control changes, in a compact form. Scores can be concatenated, merged, or repeated with operators, and various score transformations, such as diatonic transposition and pattern replacement, are available. Pytakt also has real-time MIDI input/output functions with a priority queue, which are useful in interactive music applications, together with music-theoretic classes such as scales and chords. In addition, it incorporates basic analytic features, including the retrieval of active notes and controllers at any given time with a novel algorithm and a multi-track visualiser with a playback function. This paper introduces the design and features of Pytakt and presents its usage examples, including procedural music description and data preparation for machine learning. Furthermore, the results of performance comparison with other libraries are shown to confirm the lightweight property of our library.
This paper explores the relationships between tempo, tempo variability and musical genre through the corpus analysis of 222 recorded performances of Ecuadorian m & uacute;sica nacional. These recordings were released between the years 1942 and 1969, and include some of the most well-known performances of m & uacute;sica nacional. Results indicate that: (1) there are significant differences in average tempo between six m & uacute;sica nacional genres, suggesting that these have a specific tempo range associated to them in performance; (2) tempo variability is less strongly related to genre, with two genres displaying the most variability; (3) average tempo and tempo variability are sufficient to automatically classify songs into at least five m & uacute;sica nacional genres. The main tempo categories in m & uacute;sica nacional performances and historical trends in average tempo and tempo variability are also investigated through clustering and statistical analysis, respectively.
We present the development and demonstration of a transferable method for studying user experience (UX) with digital musical instruments (DMIs) over time. We introduce DMIs, music interaction, stakeholders, and observable experiential aspects of the user-instrument relationship (UIR), grounding the development of our method in theoretical frameworks from human-computer interaction and music technology. We discuss structured evaluation strategies for studying evolving experiential components of the UIR over time, noting the limitations of current approaches. We describe the development and structure of our method before reporting on the initial execution of the method in a limited context. Using a small sample of individuals with diverse musical backgrounds and a compressed time period, we demonstrate how the method can be used to collect rich qualitative data on dynamic aspects of the UIR with an unfamiliar DMI. Results from this initial demonstration suggest that the method is able to capture comparable experiential data from different perspectives and that participants' backgrounds played a central role in their emotional and cognitive experience. We reflect on the limitations and successes of the demonstration, and offer specific suggestions for expanding the method in future, to more widely assess its transferability to different DMIs, participants, and real-world musical contexts.
The revised version of the Short Test Of Musical Preferences (STOMP-R) is thought to reflect the general structure of musical tastes. Although it has been extensively used in empirical music research, the STOMP-R has not been formally introduced in a dedicated publication. Consequently, its exact contents, aims, scope, and psychometric properties remain unclear. The present research evaluated the usability, reliability, and validity of the STOMP-R through two studies: a systematic review of its use in 28 peer-reviewed studies (Study 1) and a psychometric investigation in a French university student sample (n = 1,515; Study 2). Study 1 revealed substantial inconsistencies in how the STOMP-R has been applied, including varying item pools and heterogeneous factor solutions - even when identical item sets were used. These findings raise concerns about the measure's cross-cultural usability, reliability, and validity. Study 2 further underscored these issues: at least 10% of participants were unfamiliar with nearly half of the items, reliability coefficients were very low, and factor analyses indicated poor model fit. Taken together, these findings cast doubt on the STOMP-R's assumed capacity to reflect a general structure of musical tastes. Results derived from the STOMP-R should therefore be interpreted with caution.
Tonal Pitch Space (TPS) has been proposed as a harmonic distance model, which provides a way to compute numerical distances between two chords with keys. Although the theory itself seems sound, it still contains arbitrary presuppositions in its definition, and its adequacy has not well been investigated. In this study, we propose a framework in which we construct, train, and evaluate a variety of distance functions that share the same domain and range as TPS. We define three basic harmonic features (mode, tonic, and degree), and utilise these to define various distance functions by combining them, and then we evaluate all simple combinations exhaustively. We then train and evaluate each function through the task of key estimation. This entire process is to provide a perspective from which to examine how well various features and their combinations can represent harmony through our distance models. Experiments confirmed that some aspects of the TPS are indeed well designed, but also showed that its performance can be surpassed with a considerably smaller number of parameters.
The development of wearable technology for tactile-augmented musical experiences is growing, and with it, the need to consider how we may best design tactile effects for the body. Since associations between auditory pitch and spatial elevation are strong across hearing and sight, we evaluated whether the same type of crossmodal correspondence might benefit music experiences across hearing and touch. Twenty-one participants (musicians and non-musicians) wore a wearable tactile device and rated a series of musical audio tactile stimuli based on which they preferred. The musical stimuli were three simple melodies, composed to facilitate the detection of comparatively high and low tones. Note by note, we varied the correspondence of the tactile stimuli's intensity, timing, and placement along the vertical axis of the back with the auditory stimuli's intensity, timing, and pitch, respectively. Though we hypothesised that the pitch-height mapping would have a positive effect on participant preference, the results showed that only intensity and timing served as strong predictors for tactile musical enhancement.
Nine experienced clarinettists were asked to produce a range of 'expressive goals' (EGs, related to emotions) through different performances of a set of melodies. Several performance features were analysed, including features concerned with rarely investigated aspects of the sustain portion of a note (SPOAN). In comparison to other EGs, Angry was played loud, maximum amplitude of SPOAN arriving early, and with a steep attack slope; Deadpan: with small amplitude variations of SPOAN; Expressive: with slow tempo, amplitude peak early in SPOAN; Fearful: with low amplitude and with large variations in SPOAN; Happy: with fast tempo, large amplitude variation in SPOAN; and, Sad: with slow tempo, maximum amplitude toward end of SPOAN, and slow attack. Within the anatomy of a note, attack slope, and amplitude related features (time of the peak in the envelope of SPOAN, envelope variability and curvature) are all varied considerably for different EGs, demonstrating how players of an instrument with sustained notes manipulate these features of individual notes, in addition to other features such as tempo and loudness, to communicate EGs.
The 'Polish Music Heritage in Open Access' project of the Fryderyk Chopin Institute represents a landmark effort in digitising and providing open access to a vast array of Polish musical sources, significantly enhancing the availability and study of Poland's musical heritage. This three-year EU funded project addresses the dispersion and loss of Polish music sources, scanning over 400,000 images from 25,000 archival sources across Poland. The repository contains over 200 TB of archival images and digital scores, ranking among the largest only Polish music collections. In this paper we focus on the preparation and presentation of 8,000 digital scores containing over 12 million sounding notes that were transcribed by a team of over 20 editors during the project. An important contribution of the project was adding 42,000 source records to the R & eacute;pertoire International des Sources Musicales (RISM) database, which was also utilised as the core metadata resource for the project.
Much psychological and musicological literature has been dedicated to the affective experience of music listeners, but considerably less attention has been given to the role of composers and their intentions in the process of conveying emotion in music. The present qualitative study sought to address this gap in the literature by seeking the perspective of composers in semi-structured interviews that probed their engagement with emotion during composition, the practical considerations shaping their deployment of musical elements, and their manipulation of those elements at a technical level to achieve the desired effects. We used flexible coding within a grounded theory approach to analyse a dataset of 22 in-depth interviews. Results showed that while contemporary composers predominantly conceived of their work in terms consistent with a primarily Western classical idiom, they differed considerably in their use of distinct musical elements to evoke affective responses in listeners, their approaches to listening to music qua composer versus qua listener, and the practical ways in which they differentiated affective reactions and cognitive emotion perceptions.
This article outlines a case for the building of a digital corpus of lute music of the period covering the sources of the works of John Dowland (1563-1626), made feasible by the existence of a large number of works in informal encodings by enthusiasts for distribution via the world wide web. Editorial work needs to be done on the basic texts, but the extra effort is likely to be less than that which would be demanded by a similarly comprehensive encoding initiative for keyboard music of the same period. Dowland's works are very widely spread in the manuscript and printed lute tablatures of his time, but relatively few pieces come from sources that are truly 'close' to the composer; most are transmitted in versions that sometimes vary significantly in detail. Dowland travelled widely in Europe before receiving his long-awaited English court appointment in 1612; several pieces exist solely in continental sources, and some of these are stylistically distinct from his early repertory. This article advocates the building of a corpus of relevant lute music which would allow a digital-humanities, computer-assisted approach to problems of attribution and style analysis.