The random perturbations of pitch and amplitude are always an integral part of quasi-periodic vocal musical waves. Perturbations in pitch and amplitude are one of the important measures of timbre in vocal music. These perturbations in pitch and amplitude are known as jitter and shimmer, respectively. The results, inter alia, indicate that these are basically aerodynamic in nature, caused by the physical properties of the voice source, particularly the non-rigidity of vocal folds. Measures of such perturbations reveal the physical property, hence the timbre characteristics of the voice source. In this work, an attempt has been made to study the effect of vocal F0 on jitter and shimmer in the signals of four trained singers of both sexes and three non-singers, all are male subjects. Intracorrelations between the two parameters were studied. Highly trained singers are known to possess a fine skill of pitch control. It has been found that the values of jitter and shimmer are higher in the case of trained singers than in that of non-singers, and the interdependence between the parameters is significantly higher in the case of trained singers. This is due to larger control of singers on the vocal apparatuses by rigorous training. The values of jitter and shimmer and the corresponding standard deviation are very similar for different individuals in the same class. For non-singers, jitter and shimmer are lower than those of the singers. It seems that the significantly large values of jitter and shimmer for singers indirectly point towards the more relaxed vocal folds, presumably because of long periods of vocal training. We have observed no significant difference in these parameters between male and female subjects.
Timbre is one of the most important features in qualifying the sound of a vocalist. This study is focused on identifying the timbre characteristics of vocal sound signals of eminent vocalists in Hindustani music and classifying their voices accordingly. In our study we have considered the following timbral parameters, viz. Irregularity of the partials, Brightness, Tristimulus1 (T1), Tristimulus2 (T2), Spectral Centroid, Jitter or Irregularity of the frequencies of the partials, Shimmer or Irregularity on the amplitude evolution of the partials, Odd parameter and Even parameter. In this study we have taken the sung voices of three male and three female singers. Each of them has sung four ragas. Due to bad rendering one signal was rejected. Therefore, twenty-three (23) signals of equal durations were considered for analysis. Since the objective of the work is to classify the voices with similar timbre, raga wise classification was not attempted. The study was concentrated on the singer wise data for each male and female singer. For each singer these parameters were measured, studied and compared separately among the three male singers named as A, B and C and three female singers named as D, E and F and classification of these singers were attempted. Findings of these parameters were compared statistically based on Minkowski distance to determine nearest neighbor or nearest points. A and B among male singers and E and F among female singers were found to be nearest neighbor of each other.
A nasal voice is a type of speaking voice characterized by speech with a nasal quality. For those who sound nasal, air comes through the nose caused by limited space in the back of the throat. In that case the naso-pharyngeal cavity is connected to the oral cavity by the opening of the velum. Role of nasality in Hindustani vocal music is controversial, since most of the schools refer to it as an error. In this study we have tried to see the acoustic changes which occur during rendition of two ragas by two different singers in both oral and nasal voice. A database consisting ten signals, sung by two singers, out of which five signals were of the raga Bhairav and the other five signals were of the raga Darbari was taken. Each singer sang the same phrase of the given raga firstly in their normal voice and added nasality gradually to the rest four signals in an increasing order. Three formants (F1, F2 and F3) along with the pitch and the notes were extracted from each of the ten signals and the spectral cues were studied for both normal and nasal singing. Notes were extracted from the steady state pitch values. The variation of amplitude with frequency was studied.
Music is a performing art. From the ancient time musical evaluation visa-vis perfection is subjective. Usual methods of assessment in vogue are subjective and depend largely on the personality and the background as well as the mood and temperament of the assessor. Therefore, a quantitative method for the assessment of the closeness of the performance of a student with respect to his/her teacher/idol in Hindustani classical music is a necessity. In this paper, we study the methods to evaluate the features which might estimate the closeness of two music signals. Here we compare some features extracted from analyzing the audio signals of both the teacher and the student and discuss how these features can be used as a measure of closeness. These are beat, note duration, the time between two successive notes, frequency of each note, frequency shift between successive notes, attack time of each note and mel frequency.
In Hindustani music (HM) melody is the most important factor to any performance. Usually, in Khayal singing the entire performance is based on a single raga (melodic mode/scale). Only a small portion of the total rendition is pre-composed; that is Bandish or compositions. To embellish the composition musicians use different types of musical patterns within the raga frame which are known as Alankar or ornaments (musical). There are a wide variety of ornaments. Its application is to some extent depending on gharana and of course on the singer’s cognitive and imaginative abilities. There are two fundamental elements used in improvisation: ‘alankars’ (musical ornamentation) and tanas (as per the old concept it was also a part of alankar). However, in this paper, we have intended to search for these musical ornamentations (alankaran/alankar) from some song signals. We have primarily considered meend (uninterrupted gliding between two or more notes) (Ghosh Laxmi Narayan (1975): Geet – Baddyam (part-1): Pub: Pratap Narayan Ghosh, 1st edition, Calcutta.), andolan (oscillations in the note region), gamaka (two notes alternation; i.e. S R-R-R S) and Murki. The key phrases and transition between notes provide strong cues to identify the underlying raga in Hindustani styles of Indian music. In this work, we consider the segmentation of selected aalap parts from audio signals of three renowned vocalists, and by computing on time series of automatically detected pitch values we have investigated the style of their vocal performance. The methods are investigated in the context of detecting the style of Hindustani vocal musicians from their performances.
Indian twin drums mainly bayan and dayan (tabla) are the most important percussion instruments in India popularly used for keeping rhythm. It is a twin percussion/drum instrument of which the right-hand drum is called dayan and the left-hand drum is called bayan. Tabla strokes are commonly called as 'bol' and constitute a series of syllables. In this study we have studied the timbre characteristics of nine strokes from each of five different tablas. Timbre parameters were calculated from the LTAS of each stroke signals. The study of timbre characteristics is one of the most important deterministic approaches for analyzing tabla and its stroke characteristics. Statistical correlations among timbre parameters were measured and also through factor analysis we get to know about the parameters of timbre analysis which are closely related. Tabla strokes have unique harmonic and timbral characteristics at mid-frequency ranges and have no uniqueness at low-frequency ranges.
Indian Classical Music (ICM) has been known to convey a variety of emotional responses among the listeners since time immemorial, but neural attributes of these emotional experiences are largely unexplored till date. One of our earlier studies based on acoustical time series analysis reported that the timbre of an instrument can play significant role in conveying different emotions through various raga clips. This study aims to explore the neuro-cognitive attributes of the same. Taking two-minute aalaap (opening) sections of six popular ragas (three of them are conventionally known to portray happiness and the rest three sadness) of ICM, played in three different musical instruments (sitar, sarod and flute) by three maestros of ICM, this study attempted to look for the neural cues for these two contrasting emotions and if (and how) the timbre of the instrument plays any significance in their perception. EEG data was recorded from 10 (8 male, 2 female) participants while they listened to these music clips and their brain responses for each clip were analyzed with nonlinear chaos-based MFDFA technique in a comparative manner in order to look for specific arousal activities in different lobes of the brain corresponding to the contrasting music-evoked emotions. Our findings revealed that the multifractal spectral width obtained from alpha/theta frequency ranges of EEG data can be developed as a parameter for the development of an automated emotion recognition system. This study may prove to have far reaching implications in the development of an automated emotion classifier algorithm in future.
Music of any form is a time series variation of different note combinations, where each note has a particular frequency. Neural responses from different brain parts are recorded with the help of an EEG experiment in response to this time-varying combination of notes. These responses are interpreted in the forms of non-linear, non-stationary time series. Since music and brain signals, both are complex time series, it would be interesting to study whether any kind of correlation exists between these two series or not. Using Multifractal Detrended Cross-Correlation Analysis (MFDXA), we have tried to compute the cross-correlation coefficients (γx) between these two complex time signals. Two Indian string instruments, Sitar and Sarod were chosen and certain Alaap sections from live performances of maestros were selected to prepare the clips to be used for the experiment. From the audience response survey, the emotional contents of the prepared clips were marked. Finally, using happy and sad clips of these two string instruments as the input signals, EEG was performed on 2 musicians (M) and 2 non-musicians (NM). The γx values between audio inputs and extracted EEG responses, as well as between EEG responses of different lobe pairs were computed. γx essentially computes the degree of correlation between the source audio signals and the output EEG signals, while the correlation between the lobes essentially provides a cue to the varying neural connections happening while listening to a music. Finally, the comparative natures of cross-correlation trends for different emotions and different audience categories were studied in details. The goal of this pilot study is to develop a novel approach to classify and characterize emotions (happy-sad) and audience categories (M-NM) depending upon the nature of both audio-EEG cross-correlation as well as inter-lobe EEG cross-correlation.
In Hindustani music, a small segment is the precomposed but the remains involve with a gradual systematic exploration of the raga by various forms of improvisation keeping the grammar of the raga intact. Performers are free to explore their improvisational patterns based on their experience and expertise. According to the artist's cognitive and imaginative abilities, the quality and nature of improvisation differ among artists. One of the most important tools used in improvisation is musical ornamentation (alankar/alankaran). Extempore variations that a musical performer creates while performing within the limit of melodic pattern and the rhythmic cycle could be termed as alankar or alankaran (ornamentation). These variations in performance embellished and enhanced the beauty of the raga. Here, in this study, we are interested in finding the hidden musical patterns and structures in Indian classical music which are fundamental in unfolding the musical ornamentation of a raga. Here, we have discussed only two ornamentations, viz. meend and andolan. The study leads to identifying the style of vocal performers.
Can the sound of a vocalist be qualified by timbre? Can we identify the style of a vocalist from timbre? In this work, our objective is to identify timbre characteristics of vocal sound signals of eminent vocalists in Hindustani music, and we tried to identify the style of the artist from those timbre structures. The descriptive timbre parameters that help with identifying the style of vocalists are brightness, odd and even harmonics, irregularity among partials and spectral centroids. Beside these, shimmer, jitter and frequency per partial number were also descriptive. The study concludes that timbre is an important feature to distinguish or identify the style of a vocalist.
The works of Rabindranath Tagore have been sung by various artistes over generations spanning over almost 100 years. there are few songs which were popular in the early years and have been able to retain their popularity over the years while some others have faded away. In this study we look to find cues for the singing style of these songs which have kept them alive for all these years. For this we took 3 min clip of four Tagore songs which have been sung by five generation of artistes over 100 years and analyze them with the help of latest nonlinear techniques Multifractal Detrended Fluctuation Analysis (MFDFA). The multifractal spectral width is a manifestation of the inherent complexity of the signal and may prove to be an important parameter to identify the singing style of particular generation of singers and how this style varies over different generations. The results are discussed in detail.
It is already known that both auditory and visual stimulus is able to convey emotions in human mind to different extent. The strength or intensity of the emotional arousal vary depending on the type of stimulus chosen. In this study, we try to investigate the emotional arousal in a cross-modal scenario involving both auditory and visual stimulus while studying their source characteristics. A robust fractal analytic technique called Detrended Fluctuation Analysis (DFA) and its 2D analogue has been used to characterize three (3) standardized audio and video signals quantifying their scaling exponent corresponding to positive and negative valence. It was found that there is significant difference in scaling exponents corresponding to the two different modalities. Detrended Cross Correlation Analysis (DCCA) has also been applied to decipher degree of cross-correlation among the individual audio and visual stimulus. This is the first of its kind study which proposes a novel algorithm with which emotional arousal can be classified in cross-modal scenario using only the source audio and visual signals while also attempting a correlation between them.
Color perception is a major guiding factor in the evolutionary process of human civilization, but most of the neurological background of the same are yet unknown. This work attempts to address this area with an EEG based neuro-cognitive study on response of brain to different color stimuli. With respect to a Grey baseline seven colors of the VIBGYOR were shown to 16 participants with normal color vision and corresponding EEG signals from different lobes (Frontal, Occipital & Parietal) were recorded. In an attempt to quantify the brain response while watching these colors, the corresponding EEG signals were analysed using two of the latest state of the art non-linear techniques (MFDFA and MFDXA) of dealing complex time series. MFDFA revealed that for all the participants the spectral width, and hence the complexity of the EEG signals, reaches a maximum while viewing color Blue, followed by colors Red and Green in all the brain lobes. MFDXA, on the other hand, suggests a lower degree of inter and intra lobe correlation while watching the VIBGYOR colors compared to baseline Grey, hinting towards a post processing of visual information. We hope that along with the novelty of methodologies, the unique outcomes of this study may leave a long term impact in the domain of color perception research.
At present emotion extraction from speech is a very important issue due to its diverse applications. Hence, it becomes absolutely necessary to obtain models that take into consideration the speaking styles of a person, vocal tract information, timbral qualities and other congenital information regarding his voice. Our speech production system is a nonlinear system like most other real world systems. Hence the need arises for modelling our speech information using nonlinear techniques. In this work we have modelled our articulation system using nonlinear multifractal analysis. The multifractal spectral width and scaling exponents reveals essentially the complexity associated with the speech signals taken. The multifractal spectrums are well distinguishable the in low fluctuation region in case of different emotions. The source characteristics have been quantified with the help of different non-linear models like Multi-Fractal Detrended Fluctuation Analysis, Wavelet Transform Modulus Maxima. The Results obtained from this study gives a very good result in emotion clustering.
In the last few decades, nonlinear science and chaos theory has provided several robust non-deterministic tools by means of which the complexity of a nonlinear audio waveform can be measured precisely. On the other hand, sound signal analysis in linear deterministic approach has reached a new dimension where a number of well equipped software have been developed which can minutely measure and control the basic parameters of sound like pitch, intensity, tempo etc. The main objective of the present work is to quantitatively study the changes in acoustic signal complexity (measured using chaos based fractal technique) with individual variation in pitch, loudness and timbre of a sound signal. EEG (Electroencephalography) was also performed on 10 participants to see how the neuro-cognitive attributes of a sound change, i.e. when these basic components - pitch, loudness and timbre of the sound vary, one at a time. Single strokes of a piano were recorded where pitch and loudness of the sound signals were varied one at a time keeping the other parameters fixed. Then the sounds of 14 different musical instruments playing the same pitch at same loudness were recorded, which effectively served the purpose of timbre variation. EEG experiment was conducted with these audio signals as stimuli for the participants. The multifractal spectral widths were calculated for all the music signals as well as the corresponding EEG signals using Multifractal Detrended Fluctuation Analysis (MFDFA) and compared with each other. The results point towards the direction of a correlation between the conventional linear parameters and the latest nonlinear features in the acoustic domain, while the changes in the multifractal values of the different EEG waves reveal new information about the cognition of the basic features of sound in human brain. This study is a novel attempt to provide new data in engulfing apparent objective (acoustics) - subjective (EEG) connection, which is highly needed for building any model for perception-cognition connectivity. (C) 2020 Elsevier B.V. All rights reserved.
Music is often considered as the language of emotions. It has long been known to elicit emotions in human being and thus categorizing music based on the type of emotions they induce in human being is a very intriguing topic of research. When the task comes to classify emotions elicited by Indian Classical Music (ICM), it becomes much more challenging because of the inherent ambiguity associated with ICM. The fact that a single musical performance can evoke a variety of emotional response in the audience is implicit to the nature of ICM renditions. With the rapid advancements in the field of Deep Learning, this Music Emotion Recognition (MER) task is becoming more and more relevant and robust, hence can be applied to one of the most challenging test case i.e. classifying emotions elicited from ICM. In this paper we present a new dataset called JUMusEmoDB which presently has 400 audio clips (30 seconds each) where 200 clips correspond to happy emotions and the remaining 200 clips correspond to sad emotion. For supervised classification purposes, we have used 4 existing deep Convolutional Neural Network (CNN) based architectures (resnet18, mobilenet v2.0, squeezenet v1.0 and vgg16) on corresponding music spectrograms of the 2000 sub-clips (where every clip was segmented into 5 sub-clips of about 5 seconds each) which contain both time as well as frequency domain information. The initial results are quite inspiring, and we look forward to setting the baseline values for the dataset using this architecture. This type of CNN based classification algorithm using a rich corpus of Indian Classical Music is unique even in the global perspective and can be replicated in other modalities of music also. This dataset is still under development and we plan to include more data containing other emotional features as well. We plan to make the dataset publicly available soon.
The work explores the variation of scaling exponents in a song, recitation and reading with the same lyrical content. Detrended Fluctuation Analysis (DFA) have been employed to find the long range temporal correlations (or the Hurst Exponent) present in each form of the auditory signal. Perceptually, it is known that the addition of rhythm and pitch, amplitude modulation distinguish between these different forms of audio signals, but the mathematical analogue of the same is still unknown. In this work, recordings were taken for 2 artists (1 male, 1 female) who were asked to read, recite and sing the entire lyrics of two (2) self chosen Tagore songs, each of which were later put to analysis. The rationale behind choosing Tagore’s works is because of its strong lexical content. It was seen that Hurst Exponent is found to be the lowest in case of reading while it maximizes in case of song implying that the amount of long range correlation increases consistently with addition of rhythmic content and pitch, amplitude modulation in the audio signals. With this work, we tried to establish a critical value of Hurst exponent, above/below which transformation occurs to other forms analogous to critical temperature in phase transition of matter. Keywords: Speech, song, recitation, Rabindranath Tagore, Nonlinear analysis, DFA
The relationship between color and music as part of the complex system consisting of visual and auditory domain has been investigated in this study. As both the stimulus forms are processed in the same part of the human body, i.e., the brain, it will be really interesting to examine whether they share a similarity in perception. Needless to say that color and music both have strong impact on emotion and feelings & also a few studies have been reported in literature to explore causal relationship between color and emotion. This work reports a neuro-cognitive study on response of brain to two different stimulus and their cross-modal associations. In this study the correlation between emotional arousal and the effect of audio and visual stimuli has been studied from a new perspective. 93 participants were asked to hear 6 different music pieces (each of 30 s duration). The type of emotion elicited by different music pieces were identified by the participants from a given collection of possible emotional responses. Then they are asked to assign a color associating the emotion from a given color wheel (structured according to Munsell color system/RGB color space). Each color, associated with a particular music piece, is a mixture of specific Red, Green and Blue values (RGB triplet) and has a specific HEX number (hexadecimal representation), which is recorded for each response. Then, the musical pieces used were further zoomed with the help of fractal technique to identify different emotions related to music in a quantitative approach. Here, to analyze the complexity of the sound signal (which are non-stationary and scale varying in nature), we have used Multifractal detrended fluctuation analysis (MFDFA), which is capable of determining multifractal scaling behavior of non-stationary time series. From the experimental data, it is seen that the visual and emotional response to the auditory stimulus follows a specific trend which is directly related to the stimulus complexity. (C) 2019 Elsevier B.V. All rights reserved.
Ajoy K. Datta合作论文数Computer Science2