In the present study we aim to capture rhythmic and melodic patterning in speech and singing directed to infants. We address this issue by exploring the acoustic features that best predict different classification problems. We built a database composed by infant-directed speech from two Portuguese variants (European vs Brazilian Portuguese) and infant-directed singing from the two cultures, comprising 977 tokens. Machine learning experiments were conducted in order to automatically discriminate between language variants for speech, vocal songs and between interaction contexts. Descriptors related with rhythm exhibited strong predictive ability for both speech and singing language variants’ discrimination tasks, presenting different rhythmic patterning for each variant. Common features could be used by a classifier to discriminate speech and singing, indicating that the processing of speech and singing may share the analysis of the same stimulus properties. With respect to discriminating interaction contexts, pitch-related descriptors showed better performance. We conclude that prosodic cues present in the surrounding sonic environment of an infant are rich sources of information not only to make distinctions between different communicative contexts through melodic cues, but also to provide specific cues about the rhythmic identity of their mother tongue. These prosodic differences may lead to further research on their influence in the development of the infant’s musical representations.
We aim to model infants’ perception and representation of temporal information that is present in infant directed speech and singing, using connectionist computational models (neural networks). In our approach, we consider the sound patterning, present in both speech and singing, in terms of timing and accent. The model receives audio on the input. Subsequently, different features are computed, according to different processes operating in parallel. Finally, we compute a representational transition, which learns categorical structured representations in terms of communication purposes from unstructured examples. In addition, we propose experiments to perform with the model. With these experiments we aim to study the development of representations from undifferentiated whole sounds to the relations between attributes that compose those sounds. Keywords-rhythm; music; speech; development; representation; computation. I. EXTENDED ABSTRACT Music and speech, from the perspective of a pre-verbal infant, may be perceived as sound sequences that unfold in time, following patterns of rhythm, stress and melodic contours. Therefore, it is likely that they share processing mechanisms and representation structures [1]. In this context, we aim to model infant’s perception and representation of temporal information that is present in infant directed speech and singing. In our approach, we consider the sound patterning, present in both speech and singing, in terms of timing and accent. We leave the grouping or phrasal patterning for a later stage. The model receives audio on the input. Subsequently, different features are computed, according to different processes operating in parallel. Finally, we compute a representational transition, which learns categorical structured representations in terms of communication purposes from unstructured examples. The model has been conceptualized around two main modules: the Perception Module and the Representation Module (see Figure 1 and Figure 2). In the Perception Module (see Figure 1), we hypothesize that for the perception of rhythmic information duration, intensity and pitch must be extracted from the audio input. Therefore, we have extracted this information from the vocalic intervals present in each input sound. Infants are able to segment vocalic intervals from the speech stream [2]. Vowels are perceptually relevant regarding rhythm since in languages with rhythmic patterns close to stressed-timing, which is the case, stress has a strong influence on vowel duration and the marking of certain syllables within a word as more prominent than others leads to vowels’ duration fluctuation. In addition, the usage of vocalic intervals allows building a parallel with music since musical notes can be compare to syllables and vowels form the core of syllables [3]. We have used Prosogram [4] for segmenting vocalic intervals and extract duration, intensity and pitch values for each interval. This information that is extracted for each sound input, either speech or singing, will feed four elements that we called Contrast, Mode, Beat and Stress. Each of these elements will be described next. • Contrast is based on the variability of the duration of the input units. For each input sound, nPVI is computed, which yields the index of contrast between neighboring durations. Contrast outputs “constant” for lower values of nPVI and “variant” for higher values of nPVI. • Mode detects the input sound velocity by computing speech rate. Speech rate represents the number of durational units (vowels, in this case) per second. Mode outputs “fast” for higher speech rate values and “slow” for lower speech rate values. • Stress identifies salient events in the input sounds. Salient events are marked whenever simultaneous peak values are spotted in Duration – Intensity or Pitch – Intensity pairs. • Beat is fed by Stress and Contrast and detects regularity in the prominent events. Beat outputs “regular” for input sounds with beat and “irregular” for input sounds with no beat. Figure 1– Perception Module. Therefore, in our account, we employ a very primitive set of contrastive features which are fast/slow, constant/variant and regular/irregular. These features are extracted for each object that is presented as input sound in the Perception Module. All objects fit in to a high-level representation that is composed by four categorical objects: Prohibitive Speech, Affectionate Speech, Play-song Singing and Lullaby Singing
Acknowledgments: Barcelona Media http://www.barcelonamedia.org EmCAP Project http://emcap.iua.upf.es Music Technology Group http://mtg.upf.edu/ Pitch representation develops depending on time or age, experience and learning. Infants show developmental shifts in the focus on absolute or relative pitch information ANN have demonstrated success as being suitable for solving cognitive developmental modelling problems
Sequence learning is an important process involved in many cognitive tasks, and is probably one of the most important processes governing music processing tasks. In this work we build and evaluate computational models addressed to solve a tonesequence learning task in a framework which simulates forcedchoice tasks experiments. The specific approach we have selected is that of Artificial Neural Networks in an on-line setting, which means the network weights are always updated when new events are presented. Here, we aim at simulating the findings obtained by Saffran, Johnson, Aslin, and Newport (1999). We propose a validation loop that follows the experimental setup that was used with human subjects, in order to characterize the networks’ accuracy to learn the statistical regularities of tone sequences. Tone-sequence encodings based on pitch class, pitch class intervals and melodic contour are considered and compared. We simulate the forced-choice task by selecting the attended tone-word which is best expected among a tone-word pair. The experimental setup is extended by introducing a pre-exposure forced-choice task, which makes it possible to detect an initial bias in the model population prior to exposure. Two distinct models (i.e. Simple Recurrent Network or a Feedforward Network with a time window of one event) lead to similar results. We obtain the best match with the ground truth using an encoding based on Pitch Classes, that is, based on an absolute pitch representation instead of intervals. Furthermore, we highlight the impact of tone sequence encoding in both initial model bias and post-exposure discrimination accuracy and suggest that melodic encoding should be further investigated in the modeling of psychological experiments involving musical sequences.