Evolutionary systems have demonstrated remarkable results in creative domains, with recent applications in generative typography, design, and music. However, an open problem remains in designing fitness functions that effectively capture the desired aesthetics of abstract outputs. In this work, we explore two methods for evaluating the aesthetics of a population using Vision-Language Models (VLMs). The first method uses CLIP-IQA to predict an aesthetic score for each design. The second method instead pits candidates against each other, with winners determined by a VLM using a custom prompt specified by the user. The outcomes of these pairwise comparisons are then used to estimate a population ranking via the Glicko rating system. We present these methods in the context of a case study using a custom generative system and compare the resulting rankings with an artist's aesthetic ranking and those produced by other aesthetic evaluation techniques. Additionally, we document the artist's experience using these approaches to evolve designs, critically analysing the strengths and weaknesses of both methods.
This paper presents a practice-based case study that explores personalisation as a design methodology for embodied gestural instruments. Through a multi-year collaboration with a professional performer with physical disability, we employed research-through-design and co-design methods to iteratively develop a personalised instrument responsive to the performer’s unique movement style and creative vision. The system enabled real-time control of sound, stage lighting, and visualisations, and was ecologically validated in rehearsals and in an award-winning disability-led theatre production. Extending beyond the stage, we refined the instrument through workshops with young people with motor, sensory, and communication disabilities, demonstrating its adaptability to diverse bodies and creative practices. Rather than focusing on generalised solutions, this work advances methods for designing technologies that embrace difference, tailoring interaction to individual capabilities. We contribute to HCI research by articulating personalisation as a methodological approach to inclusive interaction design, expanding opportunities for creative expression among people with physical disability.
Recent advances in multimodal large language models (MLLM) for audio music have demonstrated strong capabilities in music understanding, yet symbolic music, a fundamental representation of musical structure, remains unexplored. In this work, we introduce MIDI-LLaMA, the first instruction-following MLLM for symbolic music understanding. Our approach aligns the MIDI encoder MusicBERT and Llama-3-8B via a two-stage pipeline comprising feature alignment and instruction tuning. To support training, we design a scalable annotation pipeline that annotates GiantMIDI-Piano with fine-grained metadata, resulting in a MIDI-text dataset. Compared with the baseline trained on converting MIDI into ABC notation under the same instruction-tuning procedure, MIDI-LLaMA substantially outperforms in captioning and semantic alignment in question answering. Human evaluation further confirms the advantages of MIDI-LLaMA in music understanding, emotion recognition, creativity, and overall preference. These findings demonstrate that incorporating symbolic music into large language models enhances their capacity for musical understanding.
Although directive prompting is the predominant way to interact with Large Language Models (LLMs), many creative practices rely on language that is open-ended, associative, phonaesthetic, symbolic, and that unfolds across multiple temporalities. In this work, we explore how creative practitioners might work with AI systems when language is treated not merely as instruction but as material. We conducted an ecological two-week study with four creative practitioners using a design probe: the Memetic Mixer, a tangible interactive device that constrains interaction with an LLM. Analysis of post-study interviews and device logs identified distinct modes of material language use and temporalities that shaped each participant's engagement with AI and their creative practice. We reflect on these findings and contribute design considerations that support open-ended interaction with AI in creative practice.
We report on two practice-based case studies in which personalised systems for interactive stage lighting were co-designed with performers, revealing how rehearsal-driven collaboration shaped both technical architectures and expressive mappings. Both movement-based systems, designed for real-time stage lighting interaction, were developed through a six-month co-design process involving iterative rehearsal sessions and interstitial prototyping cycles. This collaborative method ensured that the gestural mappings, lighting dynamics, and interaction designs were tightly coupled to each performer's expressive goals and movement vocabulary. We argue that the co-design process itself-not the final technical systems-is the primary contribution. This approach enabled mutual shaping between performer and system, allowing expressive systems for interactive stage lighting to emerge through cycles of rehearsal-based feedback, constraint negotiation, and creative exploration. Performer authorship was emphasised at every stage, with performers influencing gesture selection, lighting behaviour, and expressive mappings. Through these cases, we propose four generalisable design principles for personalised systems for interactive stage lighting: (1) choose sensing by expressive aim, (2) invest in isomorphic spatial models, (3) design modular pipelines for rehearsal-time adaptation, and (4) bring lighting into the performer's loop. These principles highlight how resonance and emergent instrumentality can be achieved through situated Research-through-Design practices in interactive performance.
The NIME conference traditionally focuses on interfaces for music and musical expression. In this paper we reverse this tradition to ask, can interfaces developed for music be successfully appropriated to non-musical applications? To help answer this question we designed and developed a new device, which uses interface metaphors borrowed from analogue synthesisers and audio mixing to physically control the intangible aspects of a Large Language Model. We compared two versions of the device, with and without the audio-inspired augmentations, with a group of artists who used each version over a one week period. Our results show that the use of audio-like controls afforded more immediate, direct and embodied control over the LLM, allowing users to creatively experiment and play with the device over its non-mixer counterpart. Our project demonstrates how cross-sensory metaphors can support creative thinking and embodied practice when designing new technological interfaces.
We present a graphical, node-based system through which users can visually chain generative AI models for creative tasks. Research in the area of chaining LLMs has found that while chaining provides transparency, controllability and guardrails to approach certain tasks, chaining with pre-defined LLM steps prevents free exploration. Using cognitive processes from creativity research as a basis, we create a system that addresses the inherent constraints of chat-based AI interactions. Specifically, our system aims to overcome the limiting linear structure that inhibits creative exploration and ideation. Further, our node-based approach enables the creation of reusable, shareable templates that can address different creative tasks. In a small-scale user study, we find that our graph-based system supports ideation and allows some users to better visualise and think through their writing process when compared to a similar conversational interface. We further discuss the weaknesses and limitations of our system, noting the benefits to creativity that user interfaces with higher complexity can provide for users who can effectively use them.
This paper presents Empathic Growth, an inter-species interactive art installation that explores plant perception from a more-thanhuman perspective. Grounded in Jakob von Uexkull's concept of Umwelt, the installation recognizes plants as active participants with their own perceptual worlds. By visualizing the biosignals of both humans and plants in a shared environment, Empathic Growth creates a visualized space of shared perception, encouraging participants to reflect on plant agency and fostering empathy between species. The installation exhibition revealed that participants developed a greater awareness of plant perception and a deeper empathy for plant life, encouraging a rethinking of humanplant relationships. Through three iterative exhibitions, Empathic Growth evolved from an initial experiment exploring emotional symbiosis between humans and plants into a real-time interactive hybrid perception interface, ultimately into digital hybrid forms as an agent-based artificial life with collective behaviors. This process realized an exploration of the interaction of human-plant empathy and the digital artificial life of growth.
Current approaches to music emotion annotation remain heavily reliant on manual labelling, a process that imposes significant resource and labour burdens, severely limiting the scale of available annotated data. This study examines the feasibility and reliability of employing a large language model (GPT-4o) for music emotion annotation. In this study, we annotated GiantMIDI-Piano, a classical MIDI piano music dataset, in a four-quadrant valence-arousal framework using GPT-4o, and compared against annotations provided by three human experts. We conducted extensive evaluations to assess the performance and reliability of GPT-generated music emotion annotations, including standard accuracy, weighted accuracy that accounts for inter-expert agreement, inter-annotator agreement metrics, and distributional similarity of the generated labels. While GPT's annotation performance fell short of human experts in overall accuracy and exhibited less nuance in categorizing specific emotional states, inter-rater reliability metrics indicate that GPT's variability remains within the range of natural disagreement among experts. These findings underscore both the limitations and potential of GPT-based annotation: despite its current shortcomings relative to human performance, its cost-effectiveness and efficiency render it a promising scalable alternative for music emotion annotation.
This paper investigates the importance of personal ownership in musical AI design, examining how practising musicians can maintain creative control over the compositional process. Through a four-week ecological evaluation, we examined how a music variation tool, reliant on the skill of musicians, functioned within a composition setting. Our findings demonstrate that the dependence of the tool on the musician's ability, to provide a strong initial musical input and to turn moments into complete musical ideas, promoted ownership of both the process and artefact. Qualitative interviews further revealed the importance of this personal ownership, highlighting tensions between technological capability and artistic identity. These findings provide insight into how musical AI can support rather than replace human creativity, highlighting the importance of designing tools that preserve the humanness of musical expression.
AbstractIn this chapter, we explore the possibilities of generative artificial intelligence (AI) technologies for building realistic simulations of real-world scenarios, such as preparedness for extreme climate events. Our focus is on immersive simulation and narrative rather than scientific simulation for modelling and prediction. Such simulations allow us to experience the impact and effect of dangerous scenarios in relative safety, allowing for planning and preparedness in critical situations before they occur. We examine the current state of the art in generative AI models and look at what future advancements will be necessary to develop realistic simulations.
This paper explores the design of an expressive visual instrument that embraces the unique movement style of a dancer living with physical disability. Through a collaboration between the dancer and an interaction designer/visual artist, the creative qualities of wearable devices for motion tracking are investigated, with emphasis on integrating the dancer’s specific movement capabilities with their creative goals. The affordances of this technology for imagining new forms of creative expression play a critical role in the design process. These themes are drawn together through an experiential performance which augments an improvised dance with an ephemeral real-time visualisation of the performer’s movements. Through practice-based research, the design, development and presentation of this performance work is examined as a ‘testbed’ for new ideas, allowing for the exploration of HCI concepts within a creative context. This paper outlines the creative process behind the development of the work, the insights derived from the practice-based research enquiry, and the role of movement technology in encouraging new ways of moving through creative expression.
Museums are permeated with digital technologies that have transformed them into unique cultural spaces by bringing an increasing variety of mediated experiences together in one place. This chapter explores how a subset of immersive technologies – including Virtual Reality, Artificial Intelligence, and remote sensing technologies such as LiDAR and Photogrammetry – affect how presence features in the aesthetic experiences associated with the museum today. The exposition frames the challenge of reinventing presence under postdigital conditions by discussing the work of three practitioner-researchers based at SensiLab (Monash University, Melbourne) who investigate technosocial co-presence through reimagining our relationship with the world in more integrative, bodily, and situated ways. Together, these creative research projects induce greater understanding of embodied experience for museums in relation to digital mediation by presenting novel ways in which cultural content can be materialised and experienced, while also surfacing productive interconnections between emerging technologies, people, and spaces that enrich our affective responses and connect us authentically with cultural experiences.
This paper presents the design and initial assessment of a novel device that uses generative AI to facilitate creative ideation, inspiration, and reflective thought. Inspired by magnetic poetry, which was originally designed to help overcome writer's block, the device allows participants to compose short poetic texts from a limited vocabulary by physically placing words on the device's surface. Upon composing the text, the system employs a large language model (LLM) to generate a response, displayed on an e-ink screen. We explored various strategies for internally sequencing prompts to foster creative thinking, including analogy, allegorical interpretations, and ideation. We installed the device in our research laboratory for two weeks and held a focus group at the conclusion to evaluate the design. The design choice to limit interactions with the LLM to poetic text, coupled with the tactile experience of assembling the poem, fostered a deeper and more enjoyable engagement with the LLM compared to traditional chatbot or screen-based interactions. This approach gives users the opportunity to reflect on the AI-generated responses in a manner conducive to creative thought.
We requested an interview with Jon McCormack after we encountered his work when looking for artists doing compelling work at the intersection of art and artificial intelligence (AI).
In this paper we present the design and development of the Transhuman Ansambl, a novel interactive singing-voice interface which senses its environment and responds to vocal input with vocalisations using human voice. Designed for live performance with a human performer and as a standalone sound installation, the ansambl consists of sixteen bespoke virtual singers arranged in a circle. When performing live, the virtual singers listen to the human performer and respond to their singing by reading pitch, intonation and volume cues. In a standalone sound installation mode, singers use ultrasonic distance sensors to sense audience presence. Developed as part of the 1st author's practice-based PhD and artistic practice as a live performer, this work employs the singing-voice to explore voice interactions in HCI beyond language, and innovative ways of live performing. How is technology supporting the effect of intimacy produced through voice? Does the act of surrounding the audience with responsive virtual singers challenge the traditional roles of performer-listener? To answer these questions, we draw upon the 1st author's experience with the system, and the interdisciplinary field of voice studies that consider the voice as the sound medium independent of language, capable of enacting a reciprocal connection between bodies.
This paper presents a study on the use of a real-time music-to-image system as a mechanism to support and inspire musicians during their creative process. The system takes MIDI messages from a keyboard as input which are then interpreted and analysed using state-of-the-art generative AI models. Based on the perceived emotion and music structure, the system's interpretation is converted into visual imagery that is presented in real-time to musicians. We conducted a user study in which musicians improvised and composed using the system. Our findings show that most musicians found the generated images were a novel mechanism when playing, evidencing the potential of music-to-image systems to inspire and enhance their creative process.
Image generation using generative AI is rapidly becoming a major new source of visual media, with billions of AI generated images created using diffusion models such as Stable Diffusion and Midjourney over the last few years. In this paper we collect and analyse over 3 million prompts and the images they generate. Using natural language processing, topic analysis and visualisation methods we aim to understand collectively how people are using text prompts, the impact of these systems on artists, and more broadly on the visual cultures they promote. Our study shows that prompting focuses largely on surface aesthetics, reinforcing cultural norms, popular conventional representations and imagery. We also find that many users focus on popular topics (such as making colouring books, fantasy art, or Christmas cards), suggesting that the dominant use for the systems analysed is recreational rather than artistic.
Anikó Ekárt合作论文数Computer and Automation Research Institute
Hungarian Academy of Sciences2
Stefano Cagnoni合作论文数Department of Engineering and Architecture, University of Parma2