
Although yes/no questions are one of the most frequently occurring question types in English, research on the development and production of yes/no questions–in particular in young English learners–is still very limited. For example, we know very little about potential errors young L2 learners make when they produce yes/no questions–an area that is crucial in order to provide useful feedback in different learning environments, including computer-based applications. This paper reports on an exploratory study conducted with Can you guess who I am? , an interactive, spoken-dialogue-based speaking activity that allows young English learners to practice yes/no questions. After intro-ducing the SDS-based speaking activity, we present the findings from a systematic investigation of the output produced by 27 young English learners in Germany (ages 9-11) who engaged with the activity. A particular focus in the analysis was placed on the types of yes/no questions elicited and the types of errors made by the young learners. The findings provide further empirical support for a six-stage framework for the development of question formation in L2 learners [14]. Moreover, they offer insights into the types of errors young EFL learners make in forming polar interrogatives such as systematic confusion with regard to the auxiliaries “to be” and “to do”. The findings are discussed in terms of (a) how they contribute to a more comprehensive understanding of young learner’s speech and (b) how they will be used to inform further development of more targeted feedback options that can be implemented into the SDS-based speaking activity in order to harness its full potential for L2 learning. English—an underexplored area in young L2 learner research.
This paper presents Figurines, an offline framework for narrative creation with tangible objects, designed to record storytelling sessions with children, teenagers or adults. This framework uses tangible diegetic objects to record a free narrative from up to two storytellers and construct a fully annotated representation of the story. This representation is composed of the 3D position and orientation of the fig-urines, the position of decor elements and interpretation of the storytellers' actions (facial expression, gestures and voice). While maintaining the playful dimension of the storytelling session, the system must tackle the challenge of recovering the free-form motion of the figurines and the storytellers in uncontrolled environments. To do so, we record the storytelling session using a hybrid setup with two RGB-D sensors and figurines augmented with IMU sensors. The first RGB-D sensor completes IMU information in order to identify figurines and tracks them as well as decor elements. It also tracks the storytellers jointly with the second RGB-D sensor. The framework has been used to record preliminary experiments to validate interest of our approach. These experiments evaluate figurine following and combination of motion and storyteller's voice, gesture and facial expressions. In a make-believe game, this story representation was re-targeted on virtual characters to produce an animated version of the story. The final goal of the Figurines framework is to enhance our understanding of the creative processes at work during immersive storytelling.
Speaker recognition is a well established area for research but it mainly focuses on adult speech. Recent work on children’s speech shows that not all the findings from speaker recognition on adult speech are directly applicable on children’s speech. There are a variety of applications for speaker recognition from children’s speech, for example it could be used as a safeguard for a child during her/his interactions on social media network-ing websites. It could also be used as one of the main blocks in automatic tutor systems for educational purposes at schools. In this research we have evaluated two scoring method for speaker recognition within the i-vector framework using two simulated environments; in a classroom (contains 30 students) and in a school (contains 288 students). The first method is based on the PLDA scoring approach and the second method is based on the cosine similarity measure. Results show that the first method outperforms the second approach in a simulated school, but it is the other way around for the recognition of a child in a classroom in which the second scoring method performs better. focused on both speaker identification and verification for text-independent mode of operation.
We present a preliminary report on developing technology for an application that supports shared book reading. We dis-cuss how speech processing technology can be used to auto-mate different components of the system for oral reading fluency evaluation during shared book reading and the challenges posed by this new context in comparison to other automated reading tutor systems. We also present performance evaluation of the baseline system on a corpus of read speech. of our is to technology to support the
Motivated by theories of early language development in children we investigate the contribution of affective features to early acquisition of lexical semantics. For the task of semantic similarity between words, semantic and affective spaces are modeled using network-based distributed semantic models. We propose a method for constructing semantic activations from a combination of lexical and affective relations and show that affective information plays a prominent role in our lexical development model.
We analyze here readings of the same reference text by 116 children. We show that several factors strongly impact subjective rating of fluency, notably number of correct words, repetitions, errors, syllables spelled per minute. We succeeded in predicting four subjective scores – rated between 1 and 4 by human raters – from such objective measurements with a rather high precision (R > .8 for 3 out of 4 scores). This open the way for automatic multidimensional assessment of reading fluency using calibrated texts.
By combining visual-feedback and motivational elements, a speech therapy computer-based system can offer new approaches with various advantages when compared to traditional speech therapy techniques. Through visual-feedback and adaptation of traditional speech sound exercises, it is possible to create an engaging environment with motivation focused elements. These elements can be used in an interactive environment that motivates the therapy attendee towards better performances. Hereby we present an interactive gamified environment for speech therapy that combines visual-feedback and motivational components. The results from a survey and a usability study suggest that children can show more interest in the speech therapy sessions when the proposed environment is used.
In research on children’s language development, joint book reading is appreciated as being a situation beneficial for language learning. Motivated by this string of research, our aim was to explore whether the situation of joint book reading can be applied to a child–robot interaction. Before investigating whether and how this situation – when applied in child–robot interaction – can be used as a language learning scenario, the interactional requirements for a successful dialogue have to be studied. Our main aim in this paper is to present a study design for a child–robot interaction, in which a robot is introduced as a learner that acquires new color words, and the child is asked to teach the robot those words within a familiar interaction format of joint book reading. We then report the observations that we made in a single-case pilot study conducted with a 4;8-year-old child, and discuss these observations in terms of how the robot’s interactional behavior needs to be shaped to successfully participate in the situation of joint book reading.
This paper presents Figurines, an offline framework for narrative creation with tangible objects, designed to record storytelling sessions with children, teenagers or adults. This framework uses tangible diegetic objects to record a free narrative from up to two storytellers and construct a fully annotated representation of the story. This representation is composed of the 3D position and orientation of the figurines, the position of decor elements and interpretation of the storytellers' actions (facial expression, gestures and voice). While maintaining the playful dimension of the storytelling session, the system must tackle the challenge of recovering the free-form motion of the figurines and the storytellers in uncontrolled environments. To do so, we record the storytelling session using a hybrid setup with two RGB-D sensors and figurines augmented with IMU sensors. The first RGB-D sensor completes IMU information in order to identify figurines and tracks them as well as decor elements. It also tracks the storytellers jointly with the second RGB-D sensor. The framework has been used to record preliminary experiments to validate interest of our approach. These experiments evaluate figurine following and combination of motion and storyteller's voice, gesture and facial expressions. In a make-believe game, this story representation was re-targeted on virtual characters to produce an animated version of the story. The final goal of the Figurines framework is to enhance our understanding of the creative processes at work during immersive storytelling.
This paper presents Kin-LDD (stands for Kinaesthetic Learning Difficulties Diagnosis), which is a tool that supports the special educators during the assessment process of children’s learning difficulties. Children using Kin-LDD, instead of participating in a tedious and extensive process, they are playing a game. The tool is using a natural user interface for the children-computer interaction, combining gestures and typical mouse usage. Kin-LDD provides a set of activities, by presenting the material in text, images, and sounds. Kin-LDD is also available to school teachers and parents, for the early identification of learning disabilities before engaging a special educator, but mostly is a tool for special educators to include the ‘fun’ factor into the diagnostic process. The tool offers activities for spatial orientation, time orientation and storyboard sequencing and reports a set of key performance indicators, related to each child’s performance in these activities, to the special educators helping them towards the diagnosis.
Stuttering is a common speech disfluency that may persist into adulthood if not treated in its early stages. Techniques from spoken language understanding may be applied to provide auto-mated diagnoses of stuttering from voice recordings; however,there are several difficulties, including the lack of training data involving young children and the high dimensionality of these data. This study investigates how automatic speech recognition(ASR) could help clinicians by providing a tool that automatically recognises stuttering events and provides a useful written transcription of what was said. In addition, to enhance the performance of ASR and to alleviate the lack of stuttering data, this study examines the effect of augmenting the language model with artificially generated data. The performance of the ASR tool with and without language model augmentation is com-pared. Following language model augmentation, the ASR tool’s performance improved recall from 38% to 62.2% and precision from 56.58% to 71%. When mis-recognised events are more coarsely classified as stuttering/ non-stuttering events, the performance improves up to 73% in recall and 84% in precision.Although the obtained results are not perfect, they map to fairly robust stutter/ non-stutter decision boundaries.
We report results for an online multi-keyword spotter in a game that contains overlapping speech, off-task side talk, and keyword forms that vary in completeness and duration. The spotter trained on a data set of 62 children, and expectations for online performance were established by 10-fold cross-validation on that corpus. We compare the post hoc data to the recognizer’s performance online in a study in which 24 new children played with the real-time system. The online system showed a non-significant decline in accuracy which could be traced to trouble understanding the jump keyword and the pre-dominance of younger children in the new cohort. However, children adjusted their behavior to compensate, and the overall performance and responsiveness of the online system resulted in engaging and enjoyable gameplay.
For individuals affected by Autism Spectrum Disorder (ASD), the inability to make eye contact is a significant barrier to their engagement in social environments. This lack of eye contact limits their ability to read social and emotional cues exhibited through facial expressions resulting in a corresponding decrease in social engagement. The use of interactive virtual environments (VEs) as a therapeutic protocol is a growing field of study. In recent studies, individuals with ASD were placed in VEs and engaged with avatars controlled by a human in the background, resulting in improvements in eye contact and engagement for some subjects. This paper is the first in a series of experiments exploring the potential of virtual avatars controlled through software agency, rather than human control, as a therapeutic tool for ASD. This paper examines if a subject could learn to make eye contact with an avatar and consequently recognize and respond to emotional cues expressed by the avatar. Results indicate that children with ASD can learn to recognize the emotional cues of the virtual avatar, and that their reactions to the avatar’s needs as well as their eye contact with the avatar improved over the course of the VE experiment. This study sets the stage for future exploration into therapeutic use of agent-based virtual avatars, including transference of emotional cues from avatars to humans in the real world.
Curiosity plays a crucial role in learning and education of children. Given its complex nature, it is extremely challenging to automatically understand and recognize it. In this paper, we discuss the contexts under which curiosity can be elicited and provide an associated taxonomy. We present an initial empirical study of curiosity that includes the analysis of co-occurring emotions and the valence associated with it, together with gender-specific analysis. We also discuss the visual, acoustic and verbal behavior indicators of curiosity. Our discussions and analysis uncover some of the underlying complexities of curiosity and its temporal evolution, which is a step towards its automatic understanding and recognition. Finally, considering the central role of curiosity in education, we present two education-centered application areas that could greatly benefit from its automatic recognition.
Acoustic models for state-of-the-art DNN-based speech recognition systems are typically trained using at least several hundred hours of task-specific training data. However, this amount of training data is not always available for some applications. In this paper, we investigate how to use an adult speech corpus to improve DNN-based automatic speech recognition for non-native children's speech. Although there are many acoustic and linguistic mismatches between the speech of adults and children, adult speech can still be used to boost the performance of a speech recognizer for children using acoustic modeling techniques based on the DNN framework. The experimental results show that the best recognition performance can be achieved by combining children's training data with adult training data of approximately the same size and initializing the DNN with the weights obtained by pre-training using the full training set of the adult corpus. This system can outperform the baseline system trained on only children's speech with an overall relative WER reduction of 11.9%. Among the three speaking tasks studied, the picture narration task shows the largest gain with a WER reduction from 24.6 % to 20.1%.
This paper examines the extent to which computer speech recognition errors for children's speech can be attributed to common phonological effects associated with language acquisition.Recognition results are presented for three corpora of children's speech, two comprising recordings of American English spoken by five-to nine-year-olds and one comprising recordings of British English speech from children aged five and six.The results are compared with adult reference confusion matrices based on TIMIT for the first two experiments and with confusion matrices for British adults and children with good speech for the third.They appear to be influenced by three factors: (i) confusions that are predictable from phonological factors associated with language acquisition also arise from acoustic confusability (e.g./k/ → /t/) , (ii) the frequency of the phonological errors is expected to decrease with increasing age, and (iii) an accurate recogniser is more likely to detect a phonological error when it occurs than a less accurate one.Overall the percentage of errors attributable to phonological processes remains approximately constant in each experiment.However, the proportion of these errors that differ significantly from reference patterns increases with recognition accuracy and is greater for children who are judged to have poor speech.
Understanding the language environment of early learners is a challenging task for both human and machine, and it is critical in facilitating effective language development among young children. This papers presents a new application for the existing diarization systems and investigates the language environment of young children using a turn taking strategy employing an i-vector based baseline that captures adult-to-child or child-to-child conversational turns across different classrooms in a child care center. Detecting speaker turns is necessary before more in depth subsequent analysis of audio such as word count, speech recognition, and keyword spotting which can contribute to the design of future learning spaces specifically designed for typically developing children, or those at-risk with communication limitations. Experimental results using naturalistic child-teacher classroom settings indicate the proposed rapid child-adult speech turn taking scheme is highly effective under noisy classroom conditions and results in 27.3% relative error rate reduction compared to the baseline results produced by the LIUM diarization toolkit.
We explored the automatic analysis of vocal non-verbal cues of a group of children in the context of engagement and collaborative play. For the current study, we defined two types of engagement on groups of children: harmonised and unharmonised. A spontaneous audiovisual corpus with groups of children who collaboratively build a 3D puzzle was collected. With this corpus, we modelled the interactions among children using network-based features representing the centrality and similarity of interactions. The centrality measures how interactions among group members are concentrated on a specific speaker while the similarity measures how similar the interactions are. We examined their discriminative characteristics in harmonised and unharmonised engagement situations. High centrality and low similarity values were found in unharmonised engagement situations. In harmonised engagement situations, we found low centrality and high similarity values. These results suggest that interactional network features are promising for the development of automatic detection of engagement at the group level.