We present a modular, sensor-based violin bow interface designed to augment musical performance and audiovisual composition. The system integrates real-time audio processing, visual feedback, and machine learning models for gesture recognition. Through an expert user study with seven violinists, we identified themes of ergonomic adaptation, co-agency, and creative potential. While results highlight the toolkit's promise as both instrument and collaborator, limitations include the absence of non-expert perspectives, training protocols, and audience evaluation. This work contributes to human-agent interaction by positioning augmented instruments as sites of improvisational dialogue and outlines future directions in unifying system components, quantitative evaluation, and broader user studies.
Existing research on music recommendation systems primarily focuses on recommending similar music, thereby often neglecting diverse and distinctive musical recordings. Musical outliers can provide valuable insights due to the inherent diversity of music itself. In this paper, we explore music outliers, investigating their potential usefulness for music discovery and recommendation systems. We argue that not all outliers should be treated as noise, as they can offer interesting perspectives and contribute to a richer understanding of an artist's work. We introduce the concept of 'Genuine' music outliers and provide a definition for them. These genuine outliers can reveal unique aspects of an artist's repertoire and hold the potential to enhance music discovery by exposing listeners to novel and diverse musical experiences.
This chapter examines how artist Wade Marynowsky’s recent robotic performance art projects are framed within musical genres: Opera, in Robot Opera (2015), Ambient/Glitch, in Synthesiser-Robot (2017); and Disco, in The Ghosts of Roller Disco (2020). By positioning the projects within known music genres, the research expands the canon of Cultural RoboticsCultural robotics by providing platforms that allow wider communities to understand the presentation of robotic performance as cultural events within a historical context. Notions of robotic agencyAgency, dramaturgy, choreographyChoreography, robotic musical gesture, and robotic musicianship are explored across three case studies, which are presented in the contexts of live performance festivals and durational exhibitions: (1) Robot Opera, a dramaturgically designed, interactive opera for eight, larger than life-sized robots; (2) Synthesiser-Robot, a solo autonomousAutonomous robot performance for a repurposed industrial robotIndustrial robot arm, theUR3, and a hardware-software interface, the Ableton Push; and (3) The Ghosts of Roller Disco, a choreographed performance for eight robotic roller skates. The research highlights the importance of robotic agencyAgency by applying autonomousAutonomous and interactive movementMovement, localised sound, and surround sound design in creating immersive and engaging robotic performance art experiences.
We propose a novel, cost-effective soft sensor capable of detecting contact force, multiple touch points, and reflecting sensor interaction in real-time with a 3D virtual surface representation. Our fabrication process has been optimized for cost efficiency through careful material selection, utilization of automated machinery, and low-cost hardware. The sensor can be easily replicated without the need for complex laboratory equipment. The sensor employs trained neural network models for real-time signal translation into localization, force measurement, and deformation mapping. We have also developed an efficient data collection system that captures accurate 2D localization, force measurement, and 3D surface data to generate a high-quality pre-validated data set. This data set is filtered using prior knowledge before being fed to two neural network models. Our interactive prototype demonstrates the stability and accuracy of the low-cost soft sensor, delivering reliable results in both single-point and multi-point contact scenarios.
In this paper we describe a qualitative study investigating how artists work with a scalable and distributed audio-visual installation system that utilises IoT technology. With no prior experience of the system invited artists incorporated the new technology according to a creative brief for a public performance. We examined how they (i) built an understanding of the technology’s affordances, (ii) refined their creative goals, and (iii) deployed collaborative strategies to achieve creative outcomes. We examine how the artists worked from the examples we provided to integrate our audio-visual system and develop their creative work. We identify three distinct creative strategies and use these to suggest ways that the design of examples, presets and readymade configurations can be successfully integrated into interfaces for new creative technologies.
An Internet of Sounds (IoS) ecosystem for rendering object-based audio across wireless distributed devices powering lights and speakers is interfaced with wireless gestural controllers to support embodied interaction in audio spatialisation. The devices generate audio locally, supporting the design of abstract objects for affecting other media in a distributed system. Through contextualising ours with other gestural controllers interacting in spatial audio settings, we articulate how affordances of the IoS ecosystem and its components complement the embodied interaction that gestural control facilitates with object-based audio. The interactions we describe for controlling are imagined in two creative scenarios that we discuss in an envisioned application and future developments.
In most Music Emotion Recognition (MER) tasks, researchers tend to use supervised learning models based on music features and corresponding annotation. However, few researchers have considered applying unsupervised learning approaches to labeled data except for feature representation. In this paper, we propose a segment-based two-stage model combining unsupervised learning and supervised learning. In the first stage, we split each music excerpt into contiguous segments and then utilize an autoencoder to generate segment-level feature representation. In the second stage, we feed these time-series music segments to a bidirectional long short-term memory deep learning model to achieve the final music emotion classification. Compared with the whole music excerpts, segments as model inputs could be the proper granularity for model training and augment the scale of training samples to reduce the risk of overfitting during deep learning. Apart from that, we also apply frequency and time masking to segment-level inputs in the unsupervised learning part to enhance training performance. We evaluate our model on two datasets. The results show that our model outperforms state-of-the-art models, some of which even use multimodal architectures. And the performance comparison also evidences the effectiveness of audio segmentation and the autoencoder with masking in an unsupervised way.
Augmented reality (AR) technology has been investigated previously to enhance the experience within various tourism contexts. This paper reports on a study that aims to investigate user expectations when AR mediates exploration of historical artifacts in a museum context. The users in this study used an iOS software application to explore several historical artifacts using an Augmented Reality experience. Twenty-four participants were interviewed to evaluate their expectations and satisfaction with the experience. Analysis of the responses showed that users prioritised storytelling techniques within the AR experience to understand the history and significance of the artifacts.
Users are capable of noticing, listening, and comprehending concurrent information simultaneously, but in conventional speech-based interaction methods, systems communicate information sequentially to the users. This mismatch implies that the sequential approach may be under-utilising human perception capabilities and restricting users to seek information to sub-optimal levels. This paper reports on an experiment that investigates the cognitive workload experienced by the users when listening to a variety of combinations of information types in concurrent formats. Fifteen different combinations of concurrent information streams were investigated, and the subjective listening workload for each of the combination was measured using NASA-TLX. The results showed that the perceived workload index score varies in all concurrent combinations. The workload index score depends on the types and the amount of information presented to users. The perceived workload index score in concurrent listening remained the highest in Monolog with Interview (three concurrent talkers) combination, medium in Monolog with News Headlines (two talkers where one is intermittent) combination, and the lowest in Monolog with Music (one talker and a concurrent music stream) combination. Users descriptive feedback remained aligned with the NASA-TLX-based results. It is expected that the results of this experiment will contribute to helping digital content creators and interaction designers to communicate information more efficiently to users.
The high feature dimensionality is a challenge in music emotion recognition. There is no common consensus on a relation between audio features and emotion. The MER system uses all available features to recognize emotion; however, this is not an optimal solution since it contains irrelevant data acting as noise. In this paper, we introduce a feature selection approach to eliminate redundant features for MER. We created a Selected Feature Set (SFS) based on the feature selection algorithm (FSA) and benchmarked it by training with two models, Support Vector Regression (SVR) and Random Forest (RF) and comparing them against with using the Complete Feature Set (CFS). The result indicates that the performance of MER has improved for both Random Forest (RF) and Support Vector Regression (SVR) models by using SFS. We found using FSA can improve performance in all scenarios, and it has potential benefits for model efficiency and stability for MER task.
In human-computer interaction, particularly in multimedia delivery, information is communicated to users sequentially, whereas users are capable of receiving information from multiple sources concurrently. This mismatch indicates that a sequential mode of communication does not utilise human perception capabilities as efficiently as possible. This article reports an experiment that investigated various speech-based (audio) concurrent designs and evaluated the comprehension depth of information by comparing comprehension performance across several different formats of questions (main/detailed, implied/stated). The results showed that users, besides answering the main questions, were also successful in answering the implied questions, as well as the questions that required detailed information, and that the pattern of comprehension depth remained similar to that seen to a baseline condition, where only one speech source was presented. However, the participants answered more questions correctly that were drawn from the main information, and performance remained low where the questions were drawn from detailed information. The results are encouraging to explore the concurrent methods further for communicating multiple information streams efficiently in human-computer interaction, including multimedia.
In this paper we present creative practice-led research into building large, scalable "multiplicitous media" artworks in which many networked devices control lights and speakers and are coordinated over Wi-Fi to create holistic artistic and environmental experiences. We discuss competing constraints, in particular the creative constraints associated with the challenge of coding complex multi-device behaviors, maximizing creative freedom and simplifying complex engineering and design decisions. Based on recent experience building multi-device digital installation works, we propose an approach, the "broadcast-first recipe," that aims to simplify the space of creative possibilities, with a trade-off between expressive power and creative efficiency that we argue is worth adopting. We examine this approach in light of hard technical constraints such as central processing unit (CPU) and Wi-Fi bandwidth budgets. which we discuss in a concrete example. We consider how the effectiveness of the proposed approach could be further leveraged in the provision of support tools.
In this article we discuss our practice-based research into effective architectures and creative workflows for creatively coding massive multidevice light and sound installation artworks. We discuss the challenges of working with networked multidevice systems and illustrate these challenges with examples of the type of content that one may wish to display on these systems. We then consider how the structuring of a creative framework can strongly influence how an artist approaches the creation of such work, eases the process of creative search and discovery and reduces the time cost and risk of solving technical problems of architecture design. We take a design perspective on how to make effective creativity support tools and also consider a holistic perspective on creative practice that attempts to satisfy creative ideals grounded in the reality of practice.
Much research has sought to recognize and retrieve music on the basis of emotion labels. These labels are usually obtained from either subjective test or social tags. Researchers use social tags usually either by grouping tags to emotion categories or clusters directly, or by mapping tags to dimensional quadrants simply. Few research work have undertaken semantic analysis on social tags for projecting them into a dimensional emotion space, especially based on recent neural word embedding techniques using large-scale datasets. In this paper, we propose an effective solution to analyse music tag information and represent them in a 2-dimension emotion plane without limiting the corpus to contain only emotion terms. In our solution, we apply neural word embedding methods for tag representation, including Skip-gram, Continuous Bag-Of-Words (CBOW) and Global Vectors (GloVe). In our experiment, we compare these methods with traditional Latent Semantic Analysis (LSA) model based on Procrustes Analysis evaluation metrics. The results shows that neural tag embedding methods outperform LSA and represent tags with high approximations with classic circumplex emotion definitions.
In Music Emotion Recognition (MER) research, most existing research uses human engineered audio features as learning model inputs, which require domain knowledge and much effort for feature extraction. We propose a novel end-to-end deep learning approach to address music emotion recognition as a regression problem, using the raw audio signal as input. We adopt multi-view convolutional neural networks as feature extractors to learn feature representations automatically. Then the extracted feature vectors are merged and fed into two layers of Bidirectional Long Short-Term Memory to capture temporal context sufficiently. In this way, our model is capable of recognizing dynamic music emotion without requiring too much workload on domain knowledge learning and audio feature processing. Combined with data augmentation strategies, the experimental results show that our model outperforms the state-of-the-art baseline with a significant margin in terms of R2 score (approximately 16%) on the Emotion in Music Database.
This research aims to assist users to seek information efficiently while interacting with speech-based information, particularly in multimedia delivery, and reports on an experiment that tested two speech-based designs for communicating multiple speech-based information streams efficiently. In this experiment, a high-rate playback design and a concurrent playback design are investigated. In the high-rate playback design, two speech-based information streams were communicated by doubling the normal playback-rate, and in the concurrent playback design, two speech-based information streams were played concurrently. Comprehension of content in both the designs was also compared with the benchmark set from regular baseline condition. The results showed that the users’ comprehension regarding the main information dropped significantly in the high-rate playback and the concurrent playback designs compared to the baseline condition. However, in answering the questions set from the detailed information, the comprehension was not significantly different in all three designs. It is expected that such equeryfficient communication methods may increase productivity by providing information efficiently while interacting with an interactive multimedia system.
In this paper we discuss the results of a workshop study for the HappyBrackets system – a development framework for creatively coding multi-device musical performances, sound installations and interactive media artworks – in which new users using the system are invited to create new multi-device music compositions in a rapid creative and collaborative hacking session. We consider the types of works made, the problems encountered and the methods used, including how some of the new features we have added to the system support exploratory creative search. We develop our observations into design principles that we speculate will better support more rapid creative exploration of multi-device creative musical compositions.
As an important computer vision task, 3d human pose estimation in a multi-camera, multi-person setting has received widespread attention and many interesting applications have been derived from it. Traditional approaches use a 3d pictorial structure model to handle this task. However, these models suffer from high computation costs and result in low accuracy in joint detection. Recently, especially since the introduction of Deep Neural Networks, one popular approach is to build a pipeline that involves three separate steps: (1) 2d skeleton detection in each camera view, (2) identification of matched 2d skeletons and (3) estimation of the 3d poses. Many existing works operate by feeding the 2d images and camera parameters through the three modules in a cascade fashion. However, all three operations can be highly correlated. For example, the 3d generation results may affect the results of detection in step 1, as does the matching algorithm in step 2. To address this phenomenon, we propose a novel end-to-end training scheme that brings the three separate modules into a single model. However, one outstanding problem of doing so is that the matching algorithm in step 2 appears to disjoint the pipeline. Therefore, we take our inspiration from the recent success in Capsule Networks, in which its Dynamic Routing step is also disjointed, but plays a crucial role in deciding how gradients are flowed from the upper to the lower layers. Similarly, a dynamic matching module in our work also decides the paths in which gradients flow from step 3 to step 1. Furthermore, as a large number of cameras are present, the existing matching algorithm either fails to deliver a robust performance or can be very inefficient. Thus, we additionally propose a novel matching algorithm that can match 2d poses from multiple views efficiently. The algorithm is robust and able to deal with situations of incomplete and false 2d detection as well.
The Internet of Musical Things is an emerging field of research that intersects the Internet of Things, humancomputer interaction, ubiquitous music, artificial intelligence, gaming, virtual reality and participatory art through device multiplicity. This paper introduces a paradigm whereby data points and variable parameters can be strategically mapped or bound using aliases, data types and scoping as an alternative to flat address-structured mapping. The ability to send and/or access complex data types as complete entities rather than lists of parameters promotes data abstraction and encapsulation, allowing greater flexibility through modular architecture as underlying data structures can change during the lifestyle or evolution of a computer based composition. Additionally, the facility to define data accessibility, and the ability to reuse human readable names based on a variable’s scope is a common feature of most programming languages. This paradigm has been extended in that scoping a variable can be dynamically bound or addressed to specific objects, class types, devices or globally on an entire network. We describe the evolution of this paradigm through its development via various project requirements.