Bidirectional temporal contingency in infant-caregiver interaction is believed to be instrumental in shaping the development of turn-taking and speech. Likewise, caregivers are known to modulate their voices in a distinctive register when speaking to infants. Yet, less is known about whether these modulations are features of caregivers' timely responses to infant vocalisations, and whether such features support infant participation in turn-taking bouts. We analysed 231 daylong home recordings of 111 mother-infant dyads (4-34 months). Mother and infant vocalisations were automatically identified and infant vocalisations classified as cries or non-cries. Contingent responses were vocalisations which followed the vocal offset of the preceding speaker within 2 s, allowing overlapping responses. Features of infant-directed speech (pitch, pitch variability, intensity, voice quality, duration and speech rate) were extracted from maternal vocalisations. Furthermore, alignment in pitch, intensity and voice quality of mother responses to infant vocalisations was calculated. Maternal responses were also automatically classified as single words, questions, declaratives, or imperatives. Mothers' contingent responses showed a simplified and acoustically salient profile compared to non-contingent speech: they were shorter and slower, with higher pitch variability, intensity and voice quality, and they were more likely to be single words and questions. Contingent responses that were shorter, slower and clearer (higher intensity and voice quality) were more often followed by infant vocalisation. These findings suggest that beyond temporal contingency, caregivers' brevity, acoustic clarity and question-asking may also support the turn-taking exchanges through which early communicative competence develops.
The formation of a subjective sense of confidence often requires the integration of signals from multiple sources of evidence. This is of particular relevance when one needs to determine with a high degree of certainty whether a multisensory stimulus is present or absent. Understanding the mechanisms underlying this ability to map evidence strength from multiple modalities into a single confidence value is therefore central to the study of metacognition. To this end, we asked healthy adults to detect the presence or absence of near-threshold stimuli that could be visual, auditory, or both, and then rate their confidence. In two pre-registered experiments (N = 48 and N = 54), audiovisual stimuli were better detected than unimodal ones, but were not associated with better metacognitive performance. Surprisingly, participants were more confident in their absence than presence judgments. To explain these results, we fitted a Bayesian evidence accumulation model in which sensory evidence is available for presence only, rendering decisions about absence dependent on counterfactual inference. The model reproduced decision patterns by assuming that a stimulus was perceived if sensory evidence from either modality exceeded a threshold (a disjunctive integration rule). In contrast, it reproduced confidence judgments by assuming that high confidence requires that the two modalities align (conjunctive for presence, disjunctive for absence). Together, these findings reveal that distinct computational mechanisms drive perception and confidence when detecting near-threshold multisensory signals.
Vocal contingency between caregivers and infants is a fundamental building block forsocio-communicative development. This paper investigates vocal contingency in homesettings and how it varies as a function of the physical distance between infants andtheir caregivers. We recorded 61 mother–infant dyads when infants were 5 to 14months old during daylong recording sessions, using wearable audio and proximitysensors. In total, 737 hours of indoor audio data were processed with machinelearning–based pipelines to extract and classify vocalisations automatically. When weconsidered all ages and proximities, we observed that mothers vocalised contingentlyto their infants, but that infants did not. Unexpectedly, above-chance contingency didnot increase with age for either partner. However, contingency varied with proximity ina quadratic fashion: it strengthened when dyads were closer, peaking at approximately0.4 m, and decreasing thereafter. This profile remained stable across ages.Importantly, above-chance contingency was also observed in infants (in addition tomothers) at proximities that fell between 0.3–0.5 m. The selective emergence of infantvocal contingency at these proximities suggests that face-to-face exchanges maypotentiate vocal contingency in infancy.
Bidirectional temporal contingency in infant–caregiver interaction is believed to be instrumental in shaping the development of turn-taking. Caregivers are known to modulate their voices in a distinctive register when speaking to infants. Yet, less is known about whether features of caregivers’ responses to infant vocalisations, beyond their temporal contingency, may support infant’s participation in turn-taking bouts. We analysed 230 daylong home recordings of 111 mother-infant dyads (4-34 months). We used machine learning tools to identify infant and maternal vocalisations and classify non-cries. Pitch, pitch variation, intensity, voice quality, and duration were extracted. Pitch, intensity and voice quality alignment of maternal responses was estimated by comparing absolute differences in acoustic values to a permuted chance-level baseline. Maternal responses were automatically transcribed and classified as questions, declaratives, imperatives and single-word utterances.Contingent responses showed a simplified and acoustically salient profile: shorter, slower, and of higher pitch, intensity, and voice quality. Pragmatically, they were more likely to be questions and less likely to be imperatives. Among contingent responses, shorter, but not slower, responses were more often followed by infant vocalisation. Responses that were clear and more pitch-aligned to infant vocalisations, rather than merely high-pitched, were associated with infant response. Questions became more likely to be followed by infant responses over age, consistent with caregivers calibrating their response complexity to their infant’s developing communicative capacity. Findings suggest that beyond temporal contingency, brevity and alignment in pitch of caregiver responses to infant vocalisations may support the turn-taking exchanges through which early communicative competence develops.
Despite strong evidence that children learn more effectively from face-to-face interactions than from screens, we still understand relatively little about the dynamic, adaptive processes through which inter-personal contingency enhances attention and learning during live interactions. In this study, we investigate how social signals during early interactions operate across multiple hierarchical levels, ranging from low-level salience cues to higher-order features. Specifically, we examine how mothers dynamically and reciprocally adjust their behaviours across these levels in response to their infants’ attention during play. To achieve this, we developed a suite of novel, information-theory-based methods to quantify naturalistic audio-visual-semantic behaviours. Using time-series analyses, we assessed moment-by-moment associations between infant attention and both lower-order features (e.g., spectral flux of ambient noise and maternal vocalizations, maternal face and hand movement) and higher-order features (e.g., speech information rate, facial expression novelty, semantic surprisal, and toy naming) in tabletop interactions involving 67 mother-infant dyads (5- and 15-month-olds). Our findings suggest that, from early infancy, the information infants perceive is continuously and dynamically modulated across multiple hierarchical levels, contingent on their behaviour and attention. When infants focus on objects, mothers reduce low-level sensory input, minimising distractions. Conversely, increases in object naming and high-level information content associate with increases in sustained attention. These results indicate that maternal behaviours are both driven by and predictive of infant attention, and that, even from early development, attention involves interactive processes which unfold across multiple levels, from salience to semantics.
Almost all early cognitive development takes place in social contexts. At the moment, however, we know little about the neural and micro-interactive mechanisms that support infants’ attention during social interactions. Recording EEG during naturalistic caregiver-infant interactions (N=66), we compare two different accounts. Traditional, didactic perspectives emphasise the role of the caregiver in structuring the interaction, whilst active learning models focus on motivational factors, endogenous to the infant, that guide their attention. Our results show that, already by 12-months, intrinsic cognitive processes control infants’ attention: fluctuations in endogenous oscillatory neural activity associated with changes in infant attentiveness. In comparison, infant attention was not forwards-predicted by caregiver gaze or vocal behaviours. Instead, caregivers rapidly modulated their behaviours in response to changes in infant attention and cognitive engagement, and greater reactive changes associated with longer infant attention. Our findings suggest that shared attention develops through interactive but asymmetric, infant-led processes that operate across the caregiver-child dyad.
Naturalistic day-long recordings offer researchers the opportunity to study the physical and social environments surrounding infants and children. These recordings allow for analyses across modalities, over time, and at multiple hierarchical levels. In this paper, we focus on audio data to present automated methods to measure patterns ranging from low-level acoustic features to high-level semantic complexity and surprisal. Our goal is to illustrate potential untapped applications of wearable devices, particularly audio recordings, in developmental research. Using environmental predictability as an example, we demonstrate how complex constructs that unfold over different timescales can be meaningfully analysed. We also discuss the key considerations and limitations associated with this approach, highlighting both its utility and the practical challenges it entails.
Predicting how well they will perform or have performed in a task – i.e., metacognitive monitoring—can allow students to engage in self-regulation during learning. Research has demonstrated that students are often prone to overconfidence and underconfidence biases however, and it remains unclear whether teachers can use simple pedagogical tools to help students (re)calibrate their metacognitive monitoring directly in the classroom. Here, we test the efficacy of a 3-step intervention that consists in routinely asking students to predict their ability to conduct a cognitive task, to retrospectively evaluate their performance, and finally, to compare their metacognitive judgments with their actual performance. We tested the efficiency of this intervention in a large-scale training study. In a test group (N = 297, 2nd to 12th grade), students used a metacognitive tool that allowed them to report their confidence before performing a learning task, after performing it, and after receiving feedback on their performance. In a control group (N = 278), students performed an equally engaging, non-metacognitive task. Predictive metacognitive monitoring improved in the test group, especially for the most vulnerable students. By contrast, retrospective metacognition did not improve. Using the tool did not improve academic outcomes in the test group compared to the control group; potential explanations for this are discussed. Our findings suggest that using simple metacognitive tools as part of regular pedagogical practices has the potential to enhance vulnerable students’ metacognitive monitoring, but that this might not be sufficient to help them improve their academic performances.
Content produced for young audiences is structured to present opportunities for learning and social interactions. This research examines multi-scale temporal changes in predictability in Child-directed songs. We developed a technique based on Kolmogorov complexity to quantify the rate of change of textual information content over time. This method was applied to a corpus of 922 English, Spanish, and French publicly available child and adult-directed texts. Child-directed song lyrics (CDSongs) showed overall lower complexity compared to Adult-directed songs (ADsongs), and lower complexity was associated with a higher number of YouTube views. CDSongs showed a relatively higher information rate at the beginning and end compared to ADSongs. CDSongs and ADSongs showed a non-uniform information rate, but these periodic oscillatory patterns were more predictable in CDSongs compared to ADSongs. These findings suggest that the optimal balance between predictability and expressivity in information content differs between child- and adult-directed content, but also changes over timescales to potentially support multiple children’s needs. In a multilingual corpus of child-directed songs, an analysis of the cumulative-compressibility of the lyrics reveals multiscale complexity patterns that could support different relevant fuctions, e.g. attention, learning and bonding.
Infant vocal production is closely linked to autonomic arousal yet the developmental transition from reflexive vocal production to vocalisations produced independent of internal state remains unclear. Using naturalistic daylong recordings from wearable devices, we examined the relationship between vocal output and heart rate in 92 infants aged 4 to 21 months. Early vocalisations occurred predominantly at heightened levels of autonomic arousal, but this coupling weakened progressively with age. Daylong heart rate time series associated with vocalisation timing above chance (ROC AUC = 0.64) and hierarchical linear mixed effects modelling confirmed a progressive developmental decoupling. Arousal also strongly associated with the acoustic features of both cry and non-cry vocalisations, with higher arousal associated with higher and more variable pitch, more harmonicity and longer duration. Notably, the influence of arousal diminished with age for non-cry vocalisations but remained stable for cries. This differential developmental trajectory suggests the emergence of parallel vocal pathways: one remaining tied to physiological state for fixed affective signals such as cries and laughs, and another increasingly decoupled from internal states for the functionally flexible signals critical to speech communication. Importantly, infants whose mothers responded contingently to a greater proportion of non-cry vocalisations showed significantly greater decoupling, suggesting that caregiver responsivity plays a role in this developmental trajectory.
Agents engaged in creative joint actions might need to find a balance between the demands of doing something collectively, by adopting congruent and interacting behaviors, and the goal of delivering a creative output, which can eventually benefit from disagreements and autonomous behaviors. Here, we investigate this idea in the context of collective free improvisation-a paradigmatic example of group creativity in which musicians aim at creating music that is as complex and unprecedented as possible without relying on predefined plans or individual roles. Controlling for both the familiarity between the musicians and their physical copresence, duos of improvisers were asked to freely improvise together and to individually annotate their performances with a digital interface, indicating at each time whether they were playing with, against, or without their partner. At an individual level, we found that musicians largely intended to converge with their coimproviser, making only occasional use of noncooperative or noninteractive modes such as playing against or playing without. By contrast, at the group level, musicians tended to combine their relational intents in such a way as to create interactional dissensus. We also demonstrate that copresence and familiarity act as interactional smoothers: They increase the agents' overall level of relational plasticity and allow for the exploration of less cooperative behaviors. Overall, our findings suggest that relational intents might function as a primary resource for creative joint actions.
The question of how young infants learn to imitate others' facial expressions has been central in developmental psychology for decades. Facial imitation has been argued to constitute a particularly challenging learning task for infants because facial expressions are perceptually opaque: infants cannot see changes in their own facial configuration when they execute a motor program, so how do they learn to match these gestures with those of their interacting partners? Here we argue that this apparent paradox mainly appears if one focuses only on the visual modality, as most existing work in this field has done so far. When considering other modalities, in particular the auditory modality, many facial expressions are not actually perceptually opaque. In fact, every orolabial expression that is accompanied by vocalisations has specific acoustic consequences, which means that it is relatively transparent in the auditory modality. Here, we describe how this relative perceptual transparency can allow infants to accrue experience relevant for orolabial, facial imitation every time they vocalise. We then detail two specific mechanisms that could support facial imitation learning through the auditory modality. First, we review evidence showing that experiencing correlated proprioceptive and auditory feedback when they vocalise - even when they are alone - enables infants to build audio-motor maps that could later support facial imitation of orolabial actions. Second, we show how these maps could also be used by infants to support imitation even for silent, orolabial facial expressions at a later stage. By considering non-visual perceptual domains, this paper expands our understanding of the ontogeny of facial imitation and offers new directions for future investigations.
We know little about the mechanisms through which leader-follower dynamics during dyadic play shape infants' language acquisition. We hypothesized that infants' decisions to visually explore a specific object signal focal increases in endogenous attention, and that when caregivers respond to these proactive behaviors by naming the object it boosts infants' word learning. To examine this, we invited caregivers and their 14-mo-old infants to play with novel objects, before testing infants' retention of the novel object-label mappings. Meanwhile, their electroencephalograms were recorded. Results showed that infants' proactive looks toward an object during play associated with greater neural signatures of endogenous attention. Furthermore, when caregivers named objects during these episodes, infants showed greater word learning, but only when caregivers also joined their focus of attention. Our findings support the idea that infants' proactive visual explorations guide their acquisition of a lexicon.
Children raised in chaotic households show affect dysregulation during later childhood. To understand why, we took day-long home recordings using microphones and autonomic monitors from 74 12-month-old infant-caregiver dyads (40% male, 60% white, data collected between 2018 and 2021). Caregivers in low-Confusion Hubbub And Order Scale (chaos) households responded to negative affect infant vocalizations by changing their own arousal and vocalizing in response; but high-chaos caregivers did not, whereas infants in low-chaos households consistently produced clusters of negative vocalizations around peaks in their own arousal, high-chaos infants did not. Their negative vocalizations were less tied to their own underlying arousal. Our data indicate that, in chaotic households, both communicating and responding are atypical: infants are not expressing their levels of arousal, and caregivers are under-responsive to their infants' behavioral signals.
In this article we examine how contingency and synchrony during infant-caregiver interaction helps children to learn to pay attention to objects; and how this, in turn, affects their ability to direct caregivers’ attention, and to track communicative intentions in others. First, we present evidence that, early in life, child-caregiver interactions are asymmetric. Caregivers dynamically and contingently adapt to their child more than the other way around, providing higher-order semantic and contextual cues during attention episodes which facilitate the development of specialised and integrated attentional brain networks in the infant brain. Then, we describe how social contingency also facilitates the child’s development of predictive models; and, through that, goal-directed behaviour. Finally, we discuss how contingency and synchrony of brain and behaviour can drive children's ability to direct their caregivers’ attention voluntarily; and how this, in turn, paves the way for intentional communication.
Neural entrainment to slow modulations in the amplitude envelope of infant-directed speech is thought to drive early language learning. Most previous research with infants examining speech-brain tracking has been conducted in controlled, experimental settings, which are far from the complex environments of everyday interactions. Whilst recent work has begun to investigate speech-brain tracking to naturalistic speech, this work has been conducted in semi-structured paradigms, where infants listen to live adult speakers, without engaging in free-flowing social interactions. Here, we test the applicability of mTRF modelling to measure speech-brain tracking in naturalistic and bidirectional free-play interactions of 9-12-month-olds with their caregivers. Using a backwards modelling approach, we test individual and generic training procedures, and examine the effects of data quantity and quality on model fitting. We show model fitting is most optimal using an individual approach, trained on continuous segments of interaction data. Corresponding to previous findings, individual models showed significant speech-brain tracking at delta modulation frequencies, but not in alpha and theta bands. These findings open new methods for studying the interpersonal micro-processes that support early language learning. In future work, it will be important to develop a mechanistic framework for understanding how our brains track naturalistic speech during infancy.### Competing Interest StatementThe authors have declared no competing interest.
Curious information-seeking is known to be a key driver for learning, but characterizing this important psychological phenomenon remains a challenge. In this article, we argue that solving this challenge requires qualifying the relationships between metacognition and curiosity. The idea that curiosity is a metacognitive competence has been resisted: researchers have assumed both that young children and non-human animals can be genuinely curious, and that metacognition requires conceptual and culturally situated resources that are unavailable to young children and non-human animals. Here, we argue that this resistance is unwarranted given accumulating evidence that metacognition can be deployed procedurally, and we defend the view that curiosity is a metacognitive feeling. Our metacognitive view singles out two monitoring steps as a triggering condition for curiosity: evaluating one's own informational needs, and predicting the likelihood that explorations of the proximate environment afford significant information gains. We review empirical evidence and computational models of curiosity, and show that they fit well with this metacognitive account, while on the contrary, they remain difficult to explain by a competing account according to which curiosity is a basic attitude of questioning. Finally, we propose a new way to construe the relationships between curiosity and the human-specific communicative practice of questioning, discuss the issue of how children may learn to express their curiosity through interactions with others, and conclude by briefly exploring the implications of our proposal for educational practices.
Over the last few decades, developmental (psycho) linguists have demonstrated that perceiving talking faces audio-visually is important for early language acquisition. Using mostly well-controlled and screen-based laboratory approaches, this line of research has shown that paying attention to talking faces is likely to be one of the powerful strategies infants use to learn their native(s) language(s). In this review, we combine evidence from these screen-based studies with another line of research that has studied how infants learn novel words and deploy their visual attention during naturalistic play. In our view, this is an important step toward developing an integrated account of how infants effectively extract audiovisual information from talkers' faces during early language learning. We identify three factors that have been understudied so far, despite the fact that they are likely to have an important impact on how infants deploy their attention (or not) toward talking faces during social interactions: social contingency, speaker characteristics, and task- dependencies. Last, we propose ideas to address these issues in future research, with the aim of reducing the existing knowledge gap between current experimental studies and the many ways infants can and do effectively rely upon the audiovisual information extracted from talking faces in their real-life language environment.
Temporal coordination during infant-caregiver social interaction is thought to be crucial for supporting early language acquisition and cognitive development. Despite a growing prevalence of theories suggesting that increased inter-brain synchrony associates with many key aspects of social interactions such as mutual gaze, little is known about how this arises during development. Here, we investigated the role of mutual gaze onsets as a potential driver of inter-brain synchrony. We extracted dual EEG activity around naturally occurring gaze onsets during infant-caregiver social interactions in N=55 dyads (mean age 12 months). We differentiated between two types of gaze onset, depending on each partners’ role. ‘Sender’ gaze onsets were defined at a time when either the adult or the infant made a gaze shift towards their partner at a time when their partner was either already looking at them (mutual) or not looking at them (non-mutual). ‘Receiver’ gaze onsets were defined at a time when their partner made a gaze shift towards them at a time when either the adult or the infant was already looking at their partner (mutual) or not (non-mutual). Contrary to our hypothesis we found that, during a naturalistic interaction, both mutual and non-mutual gaze onsets were associated with changes in the sender, but not the receiver’s brain activity and were not associated with increases in inter-brain synchrony above baseline. Further, we found that mutual, compared to non-mutual gaze onsets were not associated with increased inter brain synchrony. Overall, our results suggest that the effects of mutual gaze are strongest at the intra-brain level, in the ‘sender’ but not the ‘receiver’ of the mutual gaze.