
This study examines the perceptual boundaries between human-produced and AI-generated music through a controlled listening experiment involving 120 participants from the Sichuan Conservatory of Music. Participants evaluated 720 excerpts across six genres (Western Classical, Jazz, Rock, Pop, R&B and Chinese Traditional), with a proportional number of AI and human tracks. The AI music was generated using Suno AI and Udio AI, while the human tracks were sourced from commercial recordings. Identification accuracy was strongly genre- and model-dependent: Chinese Traditional and Western Classical were most often classified correctly, Jazz and R&B showed the lowest AI detectability, and human-produced Rock was frequently misidentified as AI-generated. Prior experience with AI music creation was associated with improved AI detection, indicating a familiarity effect. These results show that authorship judgements are shaped by sonic features, listener background and evolving AI capabilities, and that the direction of misattribution is an informative outcome. The study recommends genre-specific benchmarks for model evaluation, routine reporting of confusion matrices and longitudinal tracking so that AI development progress can be assessed against real listening behaviour. The study also supports institutional policy that pairs hands-on generation training with feature-level listening. Findings contribute to music cognition, media psychology and debates on creativity in an era of algorithmic production.
In this article, Xenakis's game of musical strategy Duel is examined and re-implemented as a computer-aided system for networked music performances. Motivated by an analysis revealing inconsistencies between game-theoretical formalisms and their translation to a game piece under Xenakis's interpretation, this work draws from the conflict between rational (assumed by game theory) and natural (driven by aesthetic preferences) behaviours. The revisited Duel highlights said conflict by design, because it comprises two principal modules: one simulating the conductors' decision-making, the other modelling the corresponding sonic events, originally performed by orchestras. Although the latter can be automatically rendered using predefined algorithmic processes, these are meant to be edited, mediated, or even entirely replaced in real time by live coders. This article provides a detailed description of the main author's system, describes the specifics of its premiere, and offers a reflection around the aesthetics of musical systems based on game theory, caught in the struggle between rational policies (to maximise rewards) and pleasing (but mathematically suboptimal) musical results.
Optical flow analysis, while initially developed for computer vision, has expanded its applications into various domains. While traditional cognitive models of musical expression focus on understanding the impact of music on listeners, they often show less interest in what the music itself expresses or fail to fully elucidate the complexities associated with this sphere. Moreover, these models tend to be coarse and predominantly top-down, categorising music based on standard emotion theories using general musical dimensions, without considering the intricacies of perceptual abstraction and idiosyncratic processing. In contrast, contemporary computer software allows for a more nuanced exploration of music aesthetics by integrating music visualisation into optical flow analysis. This article explores the potential of this integration, extending the application of algorithmic music visualisation to serve as a tool for feature detection and audiovisual priming at a finer scale. Optical flow analysis facilitates the collection and evaluation of data that transcends intuitive audiovisual experience, offering a method which may lead to a more “objective” understanding of aesthetic dimensions. By analysing micro-motion and micro-expression in abstract animations of the acoustic spectrum, this approach opens avenues to subtleties previously unexplored and deemed “ineffable”. Through two case studies involving Aphex Twin’s “Bucephalus Bouncing Ball” and Frédéric Chopin’s Prelude op. 28 no. 3, this article demonstrates the potential of optical flow analysis in uncovering hidden layers of meaning in musical expression via idiosyncratic, animated representations of spectral data.
Virtualising traditional musical instruments is gaining popularity as a means to preserve musical cultures. Various technologies are leveraged to support learning traditional musical instruments, which are neither popular nor easy to access. Despite this, many virtual instruments are often simplified in their designs and interactions to make them more attractive and accessible, disregarding the authenticity and naturalness of instrumental techniques and playability. This study presents a Digital Musical Instrument (DMI) design process based on the Malay gamelan instrument the bonang. Named the Air Bonang, the DMI is designed using our NEX2MI framework, which underlines three dimensions of designing traditional musical interfaces: they should be natural, expressive and explorative. The Air Bonang design process involved user requirements, development and validation. Using the User-Centred Design (UCD) method, quantitative and qualitative data were gathered through interactional video and a questionnaire. Results revealed that the Air Bonang can be used to learn bonang instrumental techniques while providing flexibility by introducing new gestural interactions unrestrained by the physicality of the instrument. The design process discussed in this article laid the foundations for design criteria for the bonang DMI.
This book arose out of international collaborations arranged for the occasion of the centenary year of Iannis Xenakis in 2022. Its remarkable breadth mirrors and examines the interests and international influence of the composer, one of the leading figures of the European post-war avant-garde, and a significant initiator of what have since become defining fields in twenty-first century aesthetics and musical thought. The relationship of Xenakis’s style and technique to mathematics, semiotics, architecture, computer design and the fledgling field of AI are all addressed, together with the impact that these factors in his philosophy and output have had on a range of international movements in music, technology and the arts.
In this article, we discuss the analytical findings from a qualitative investigation of musicking (Small, 1998) with a digital score. The main focus of this article is our methodology and the outline of our methods. The composition at the centre of this study is Nautilus – a digital score created using the Unity game engine. We discuss in detail the construction of a novel quantitative dataset, which has been designed as a standard structure for the analysis of different digital score case studies. Following this we present our analytical findings from the qualitative study and outline the themes that were created as part of a thematic analysis. As a conclusion, we assess the relevance of our findings against the core aims of the project, critically reflect on the methodology, and finally present some design considerations that emerged through this case study for those wishing to explore game engines as a platform for creative music-making.
In Book Eight of Athanasius Kircher’s Musurgia universalis (Rome, 1650), this Jesuit polymath describes a computing device for automated music composition, called the Arca musarithmica. A new software implementation of Kircher’s device in Haskell, a pure-functional programming language, demonstrates that the Arca can be made into a completely automatic computational system. Moreover, the project also demonstrates that the Arca in its original form already constituted a computational system that almost completely automatic though designed to be operated by a human user: as Kircher advertised, a completely “amusical” user could generate music simply by using the device according to his rules. The device itself served as a microcosm of Kircher’s goal in the Musurgia to encapsulate and codify all musical knowledge, and demonstrate that music manifested the underlying mathematical order of the Creation and its Creator. This article analyses the concepts and methods of computation in Kircher’s original system, in dialogue with the interpretation of his system in software. The Arca musarithmica, now available on the web, makes it possible actually to hear how well Kircher was able to reduce seventeenth-century music to algorithmic rules. The successes of the system are inseparable from many paradoxical elements that raises broader questions about how Kircher and his contemporaries understood the links between composition and computation, mathematics and rhetoric, traditional harmonic theory and emerging tonal practice, and concepts of “invention” and authorship.
Personalised creative computational or manual/performative exploration and perceptual experimentation with the basic sonic and structural materials of music can initiate novel expression. We propose a generalised metacultural approach that can encourage this process, while secondarily readying its users for intercultural music-making. Amongst such basic mutable musical elements we distinguish six: rhythm, pitch, timbre, dynamics, hierarchical structure and creator-interaction, each with their attendant structuring processes that anyone in any culture might consider as tools of expression. Formed cultures differ in their ranges of expectations as to stylistic fixity or flux; metaculture can freely choose its type and degree of variability. We propose five general exploratory principles, converting: discrete categories into usable continua; linearities or sequences into non-linearities or re-orderings; separations into overlays or vice versa (in space, time and other respects); or using: partial randomisation; and novel hierarchies. All the approaches we propose are susceptible to computational application, while most are underutilised. We present brief surveys of cognition of each of the sonic materials, illustrating that much remains only partially researched (whereas some dogma is often meekly accepted). This in turn supports the view that practical exploration by musicians can provide perceptible and usable creative innovations. We discuss briefly the sociocultural implications of such a novel generalised exploratory metaculture, likely requiring corresponding cognitive learning and adaptation. We conclude that in any current environment one could strive both to respect traditions and cultural sensitivities, and to evolve new musics.
In this article, we discuss the analytical findings from a qualitative investigation of musicking (Small, 1998) with a digital score. The main focus of this article is our methodology and the outline of our methods. The composition at the centre of this study is Nautilus – a digital score created using the Unity game engine. We discuss in detail the construction of a novel quantitative dataset, which has been designed as a standard structure for the analysis of different digital score case studies. Following this we present our analytical findings from the qualitative study and outline the themes that were created as part of a thematic analysis. As a conclusion, we assess the relevance of our findings against the core aims of the project, critically reflect on the methodology, and finally present some design considerations that emerged through this case study for those wishing to explore game engines as a platform for creative music-making.
In Book Eight of Athanasius Kircher’s Musurgia universalis (Rome, 1650), this Jesuit polymath describes a computing device for automated music composition, called the Arca musarithmica. A new software implementation of Kircher’s device in Haskell, a pure-functional programming language, demonstrates that the Arca can be made into a completely automatic computational system. Moreover, the project also demonstrates that the Arca in its original form already constituted a computational system that almost completely automatic though designed to be operated by a human user: as Kircher advertised, a completely “amusical” user could generate music simply by using the device according to his rules. The device itself served as a microcosm of Kircher’s goal in the Musurgia to encapsulate and codify all musical knowledge, and demonstrate that music manifested the underlying mathematical order of the Creation and its Creator. This article analyses the concepts and methods of computation in Kircher’s original system, in dialogue with the interpretation of his system in software. The Arca musarithmica, now available on the web, makes it possible actually to hear how well Kircher was able to reduce seventeenth-century music to algorithmic rules. The successes of the system are inseparable from many paradoxical elements that raises broader questions about how Kircher and his contemporaries understood the links between composition and computation, mathematics and rhetoric, traditional harmonic theory and emerging tonal practice, and concepts of “invention” and authorship.
Fuzzy relational music perception concerns the representation of congruent connections between musical features as fuzzy relations used to individuate and assemble concepts and conceptual hierarchies. This article presents two universal fuzzy domains of discourse, harmony H and grouping G, which partition sets using triangular norms (t-norms) based on generalised harmonic root support and generalised time regularity, respectively. Fuzzy relations between the sets of the domains are formed in the innate fuzzy neural architecture of a dedicated music faculty. Fuzzy relations are shown to be necessary representations for interconnection between the domains to individuate and assemble concepts. Concepts are individuated and assembled by virtue of fuzzy set resemblance relations between domains, or fuzzy logical implication relations in one or both domains through time. Fuzzy resemblance relations comprise the properties of weak reflexivity, weak symmetry and antitransitivity in a H ⨉ G Cartesian product space. Fuzzy implication relations involve fuzzy overlap (or continuation) of elements, calculated using a t-norm operator (min operator), in one or both domains of the product space. Supplementary theory is incorporated to explain polyphonic structure, involving pluralistic superimposition of independent fuzzy relational hierarchies. Broadly, fuzzy relational music perception is a rationalistic model that builds on generative theories and associative–statistical and connectionist approaches by providing a compact and coherent process for determining interaction across musical parameters.
We introduce the notion of multi-pattern, a combinatorial abstraction of polyphonic musical phrases. The interest of this approach to encode musical phrases lies in the fact that it becomes possible to compose multi-patterns in order to produce new ones. This composition is parametrized by a monoid structure on the scale degrees. This dives the set of the musical phrases into an algebraic framework since the set of the multi-patterns is endowed with the structure of an operad. Operads are algebraic structures offering a formalization and an abstraction of the notion of operators and their compositions. Seeing musical phrases as operators allows us to perform computations on phrases and admits applications in generative music. Indeed, given a set of initial multi-patterns, we propose various algorithms to randomly generate a new and longer phrase emulating the style suggested by the inputted multi-patterns. The designed algorithms use sorts of grammars working with operads and colored operads, known as bud generating systems.
This paper presents an attempt to employ the mask language modeling approach of BERT to pre-train a 12-layer Transformer model over 4,166 pieces of polyphonic piano MIDI files for tackling a number of symbolic-domain discriminative music understanding tasks. These include two note-level classification tasks, i.e., melody extraction and velocity prediction, as well as two sequence-level classification tasks, i.e., composer classification and emotion classification. We find that, given a pre-trained Transformer, our models outperform recurrent neural network based baselines with less than 10 epochs of fine-tuning. Ablation studies show that the pre-training remains effective even if none of the MIDI data of the downstream tasks are seen at the pre-training stage, and that freezing the self-attention layers of the Transformer at the fine-tuning stage slightly degrades performance. All the five datasets employed in this work are publicly available, as well as checkpoints of our pre-trained and fine-tuned models. As such, our research can be taken as a benchmark for symbolic-domain music understanding.
An important feature of the music repertoire of the Syrian tradition is the system of classifying melodies into eight tunes, called ’oktoe\={c}hos’. In oktoe\={c}hos tradition, liturgical hymns are sung in eight modes or eight colours (known as eight ’niram’ in Indian tradition). In this paper, recurrent neural network (RNN) models are used for oktoe\={c}hos genre classification with the help of musical texture features (MTF) and i-vectors.The performance of the proposed approaches is evaluated using a newly created corpus of liturgical music in the South Indian language, Malayalam. Long short-term memory (LSTM)-based and gated recurrent unit(GRU)-based experiments report the average classification accuracy of 83.76\% and 77.77\%, respectively, with a significant margin over the i-vector-DNN framework. The experiments demonstrate the potential of RNN models in learning temporal information through MTF in recognizing eight modes of oktoe\={c}hos system. Furthermore, since the Greek liturgy and Gregorian chant also share similar musical traits with Syrian tradition, the musicological insights observed can potentially be applied to those traditions. Generation of oktoe\={c}hos genre music style has also been discussed using an encoder-decoder framework. The quality of the generated files is evaluated using a perception test.
Abstract animation in the form of “visual music” facilitates both discovery and priming of musical motion that synthesises diverse acoustic parameters. In this article, two scenes of AudioVisualizer, an open-source Chrome extension, are applied to the nine musical poems of Robert Schumann’s Forest Scenes, with the goal to establish a basic framework of expressive cross-modal qualities that in audiovisual synchrony become apparent through visual abstraction and the emergence of defined dynamic Gestalts. The animations that build this article’s core exemplify hands-on how particular ways of real-time analogue music tracking convert score structure and acoustic information into continuous dynamic images. The interplay between basic principles of information capture and concrete simulation in the processing of music provides one crucial entry point to fundamental questions as to how music generates meaning and non-acoustic signification. Additionally, the considerations in this article may motivate the creation of new stimuli in empirical music research as well as stimulate new approaches to the teaching of music.
Generative musical models often comprise of multiple levels of structure, presuming that the process of composition moves between background to foreground, or between generating musical surface and some deeper and reduced representation that governs hidden or latent dimensions of music. In this paper we are using a recently proposed framework called Deep Musical Information Dynamics (DMID) to explore information contents of deep neural models of music through rate reduction of latent representation streams, which is contrasted with hight rate information dynamics of the musical surface. This approach is partially motivated by rate-distortion theories of human cognition, providing a framework for exploring possible relations between imaginary anticipations existing in the listener's or composer's mind, and the information dynamics of the sensory (acoustic) or symbolic score data. In the paper the DMID framework is demonstrated using several experiments with symbolic (MIDI) and acoustic (spectral) music representations. We use variational encoding to learn a latent representation of the musical surface. This embedding is further reduced using a bit-allocation method into a second stream of low bit-rate encoding. The combined loss includes temporal information in terms of predictive properties for each encoding stream, and accuracy loss measured in terms of mutual information between the encoding at low rate and the high rate surface representations. For the case of counterpoint, we also study the mutual information between two voices in a musical piece at different levels of information reduction.The DMID framework allows to explore aspects of computational creativity in terms of juxtaposition of latent/imaginary surprisal aspects of deeper structure with music surprisal on the surface level, done in a manner that is quantifiable and computationally tractable. The relevant information theory modeling and analysis methods are discussed in the paper, suggesting that a trade off between compression and prediction play an important factor in the analysis and design of creative musical systems.
Tempo and genre are two inter-leaved aspects of music, genres are often associated to rhythm patterns which are played in specific tempo ranges.In this paper, we focus on the Deep Rhythm system based on a harmonic representation of rhythm used as an input to a convolutional neural network.To consider the relationships between frequency bands, we process complex-valued inputs through complex-convolutions.We also study the joint estimation of tempo/genre using a multitask learning approach. Finally, we study the addition of a second input convolutional branch to the system applied to a mel-spectrogram input dedicated to the timbre.This multi-input approach allows to improve the performances for tempo and genre estimation.