
Abstract The aim of this project was to investigate how people are affected by noise pollution and how they relate to their acoustic community in the city of Morelia, in the state of Michoacán, Mexico. We achieved this by taking measurements of sound levels to create a noise map of Morelia while also recording the soundscapes of every location on the map. In addition, we recorded interviews with members of the public who live or work in each location on the map to discover how they relate to their acoustic community. The final stage of the project was to use the audio materials collected (soundscapes and interviews) to create artistic outcomes: electroacoustic works and mixed media compositions. The interviews showed that many people rarely consider sounds as noise pollution, and the ways in which these factors can impact their health. It is therefore necessary to promote awareness of these issues. The resultant works were presented in outreach concerts and workshops in different locations in the state of Michoacán, such as universities, rural areas, and cultural centers, to promote awareness of issues regarding noise pollution and to ignite discussions about soundscape with the Mexican public.
Soundscape composition has been rooted in the tradition of musique concr & egrave;te, whose basic principle was to break free from notating musical ideas with symbols and instead to use recorded real-world sounds in a musical way. The author's live-coding piece titled Latitude aims to reverse this principle and to propose a new form of soundscape composition in which place-based data, such as geolocation information, are used as symbols to create musical patterns. In addition to this, a conceptual framework contextualizes the piece, providing the extramusical element that is omnipresent in soundscape composition, but from a different angle. The musical and the extramusical exchange roles between input and output, as explained in the article, providing an additional dimension to the reversing of the roots. Along with Latitude, the existence of some pieces by other composers suggests a possible new subgenre of soundscape composition that is referred to here as "symbolic soundscape" music.
Artificial intelligence-powered assistants are revolutionizing the music industry by transforming how music is produced and experienced, opening new frontiers for creativity and innovation. Their application in supporting learners to play a musical instrument, however, has yet to be fully explored. Most current applications primarily focus on sound data to distinguish between different notes. This approach excludes the correct hand form, which is essential for learning to play an instrument. Furthermore, sound can be subject to background noise or compression. The current article presents a pipeline for guitar chord recognition based on 3-D data acquisition and color images. The developed method utilizes MediaPipe to estimate hand landmarks from color images and leverages depth images to retrieve real-depth information. In this study, various machine learning algorithms were compared to perform chord recognition from hand landmark information. The number of chords was limited to four, plus one class for unknown gestures. The classifier was trained with an RGB-D (red, green, and blue plus depth) video data set that features the hands of 18 individuals who performed the selected chords. The random forest classifier demonstrated remarkable performance in the classification task, achieving a balanced accuracy that exceeded 85% on unseen data under the same acquisition conditions and improving the state-of-the-art performance. Future developments will focus on expanding the number of supported chords so that the proposed approach could be used in real-world applications. Such applications might include not only guitar education but also transcription, control of sound synthesis, identification in ensembles, and other contexts where audio input alone is insufficient.
This article describes the semantics of notation systems for music as a network of transformations that connects the syntactic elements of the notation on one side and the semantic meaning, such as the parameters of a possible execution, on the other side. In this context, a digital encoding format for music notation can be defined by selecting a subset of the data nodes of this network for storage, leaving others to evaluation. We discuss how to integrate uncertainty, freedom of choice, and graphical attributes, and we give semantic properties of datasets that impact upon the practical costs of maintenance, migration, and extension, among other things. For their demonstration and evaluation, these dataset properties are applied to selected substrata of four widely used encoding systems: MIDI Standard File, LilyPond, MusicXML, and Music Encoding Initiative.
Long-form music generation remains a challenging task for generative models, particularly in capturing coherent musical structure over extended durations. This article introduces a hierarchical approach that combines a large language model (LLM) with a Transformer-based audio generator. The LLM designs the overall musical form and generates natural language prompts for each section, and the audio model synthesizes the corresponding music. To evaluate structural coherence, the current study proposes a new quantitative metric, in which the structural patterns of generated music are compared with those of human-composed pieces. The method is evaluated using both objective metrics and subjective listening tests. On analyzing the results, it appears that music with a more coherent and varied structure across temporal scales is produced. Additionally, a metaoptimization method is introduced, by which the prompt design process for the LLM is automated and generation quality is improved through iterative refinement.
Cross-modal interactions constantly occur when a musical instrument is played: Its perceived sound quality can be affected by its vibrotactile behavior, and its visual appearance may induce expectations that drive the musician’s way of playing. This article presents an exploratory setup in which the auditory, visual, and vibrotactile feedback of an electric guitar can be independently controlled by the player. The aim is to explore how these feedback modalities interact with each other in perceptual judgments. Verbal data and actions on the control interface were gathered during in situ interviews with five players. This pilot study suggests that the addition of sensory feedback (other than that resulting from the instrument’s usual sound production) can cause an increased feeling of immersion, an impression of sensory compensation, and an influence on the performer’s gestures and choice of repertoire.
This article presents an alternative approach for creating compositions for performers plus electronics, in contrast to the standard AI methods developed for music over the past 30 years. Despite their robust construction, the latter seem less useful to contemporary composers who pursue aesthetic directions that have shaped experimental music since the 1950s and that focus on texture and timbre. This study begins with an overview of the categories of AI-based autonomous and interactive composition, examining selected projects considered notable by scholars, with an aim to assess their usefulness for contemporary music. An alternative method for experimental composition is then introduced, using a Max for Live patch in which a simple neural network model controls an audio plug-in. Consisting of a single neuron, the model is not designed to emulate human intelligence but generates complex modulation functions. The patch's use is examined through results from two performances of the composition “Outside-Time Sketch.” The article describes artistic and technical implications, the project's relation to those discussed in the overview, and other methods for shaping and expanding the neuron's input. Future directions include the elaboration of this approach to use recorded sound samples in electronic music composition and spatial music.
This article describes a series of experiments and techniques for audiovisual performance composition using electromagnetic waves. These experiments represent the consecutive steps that led to the development of REBUS, a novel musical machine and interactive system that can be used to explore and expand upon contactless interaction techniques. Drawing on previous experience in designing simple synthesizers to transform light into sound, compositional systems that use invisible frequencies of the electromagnetic field were explored as both sonic material and interactive interfaces. A review and comparison of other sensing techniques is followed by a description of the potential of electromagnetic field sensing. Implementing state-of-the-art technology and using previously unexplored frequencies, REBUS is a novel digital computational instrument that creates a space where any subtle interaction is detected independently of external light or sound and where invisible affordances can be touched and manipulated with the hands and the body-almost as invisible strings.
This article introduces bellplay similar to, an open-source symbolic framework and software environment for algorithmically generating audio offline-that is, not in real time-originally developed for the realization of ludus vocalis, a 25-minute multimedia work for 8.1-channel audio and 4K video. bellplay similar to was leveraged to craft the entire audio component of the work, as well as to generate control data for automating visual parameters in TouchDesigner. The article begins with a contextual and conceptual overview of the development of bellplay similar to, followed by a discussion of its software architecture, scripting language, and key functionalities. Its application in ludus vocalis is presented as a case study, and the article concludes by reflecting on its effectiveness as a pedagogical tool within a university-level course on computer-assisted algorithmic composition.
This article investigates questions and concerns about the embedding of cultural and individual programmer biases into music programming languages (MPLs) and music software. A key contention is that MPLs and music software are extensions of natural language and thus inherit and transmit the cultural biases about music that are entrenched there, constraining one’s potential to explore music beyond conventional norms. Recent research is cited that shows there are indeed such effects on users. The primary source of insight into these biases and effects, however, comes from interviews with the developers of such technologies as Supercollider, ChucK, Max, and Kyma, whose responses constitute the bulk of this article. Although there is general agreement about the embedding and transmission of biases, the responses reveal different, compelling insights about this issue that should provide revelatory knowledge to developers and users of music technology alike. The interviews provide fertile ground for further reflection and analysis—especially about the need for greater openness in design and mutual exchange between developers and the communities they serve.
This article provides a retrospective analysis of China's efforts in applying digital technologies to guqin music research, tracing its evolution from early computer-assisted initiatives to contemporary computational methodologies. The guqin, an ancient seven-stringed zither, occupies a central position in Chinese musical heritages, with its repertoire historically preserved through the jianzipu tablature system. Although this notation documents the rich variety of finger techniques clearly, it lacks rhythmic detail, posing significant challenges for interpreting and reconstructing guqin music. Early efforts in the 1980s, such as Professor Chen Changlin's development of a computer processing system for the characters of the jianzipu tablature, was one of the pioneering computational approaches in this field. The establishment of China's first computer music laboratory in 1986 at Shanghai Jiaotong University further advanced this interdisciplinary research by uniting computer scientists and musicologists. Key milestones included the creation of encoding systems and software to translate pitch information from jianzipu into modern notation, alongside the application of statistical methods to analyze musical intonation features. These contributions have significantly shaped the digital humanities in China, particularly in guqin music research. The article concludes by considering the potential of emerging technologies—such as machine learning, artificial intelligence, and virtual reality—to revolutionize guqin music research and creativity, ensuring the preservation and revitalization of this invaluable Chinese cultural heritage.