
Human communication relies on multiple modal-ities-language, gestures, facial expressions, and more-each conveying different types of information over varying time scales. Over time, humans gradually learn to integrate these modalities to understand others and interpret their commands. While recent advances in machine learning have enabled robots to fuse information from images, videos, and language commands, effectively integrating human gestures and linguistic instructions in real-world scenarios remains a challenge. In this paper, we present a simulated dataset of varying complexity, including ambiguous language commands and ambiguous pointing gestures, which might be used for evaluating different learning approaches. On this dataset, we quantitatively compare 3 different approaches for combining ambiguous information from language and pointing gestures - both learning-based and deterministic ones. Furthermore, we demonstrate how the mapping between these two modalities can be gradually learned with an increasing number of observations and samples.
Transcribing naturalistic caregiver-child interactions is a labor-intensive task in developmental research. While automatic speech recognition (ASR) models offer potential solutions, current ASR systems, including OpenAI’s Whisper, were primarily trained on adult-directed speech (ADS) and have not been systematically evaluated for infant-directed speech (IDS). IDS differs from ADS in prosody, phonetics, utterance structure, and lexical patterns, posing unique challenges for ASR transcription. In this study, we benchmarked Whisper’s transcription accuracy on a dataset of naturalistic caregiverinfant recordings, comparing ASR-generated transcripts to human-annotated gold-standard transcriptions. We evaluated Whisper’s performance using word error rate (WER) and accuracy, and conducted a failure analysis to identify systematic errors made by Whisper, including utterance segmentation mismatches, high error rates in short and repetitive speech and inconsistent filler word handling. Results indicated that Whisper struggles with short, contextually ambiguous utterances and fails to reliably segment IDS utterances. Based on these results and early user experience, we propose a preliminary pipeline for integrating ASR technology with human transcription workflows to enhance the efficiency of processing large-scale naturalistic speech data in developmental research.
This study investigated how action histories – unfolding sequences of actions with objects – provide a context for both attentional allocation and linguistic repair strategies. Building on theories of enactive cognition [34] and sensorimotor contingency theory [40], we experimentally manipulated action sequences (action history) to create either simple or rich “situational models,” and investigated how these models interact with attention and reflect in linguistic processes during human–robot interaction. Participants (N = 30) engaged in a controlled object placement task with a humanoid robot, where the action (manner) information was either provided or omitted. The omission elicited repair behaviors in participants that were in focus of our investigation. For rich models (competing action possibilities) participants demonstrated: a) increased attentional reorientation, reflecting active engagement with the situational model b) preference for restricted repairs, targeting the specific source of trouble in action selection. Conversely, a simple situational model led to more generalized attention patterns and open repair strategies, suggesting weaker constraints on internal processing. These findings highlight how situational structures emerge externally to scaffold internal cognitive processes, with action histories serving as a crucial context for the interface between perception, action, and language. We discuss how to implement such a tight loop in the assistance of a system.
This paper presents a novel approach to address personalized motion retargeting (PMR) from humans to robots. In contrast to most existing retargeting methods that overlook personalized human motion features, we propose a continuous learning framework that enables robots to learn and refine retargeting through human-robot interaction. Within the framework, we design a bidirectional imitation loop to encode the user’s preferences: The robot imitates the human first, and then the human replicates the robot’s movements. By being imitated, the robot gains insight into the effect of its own action, allowing the robot to adjust its retargeting accordingly. Furthermore, to enhance the interaction efficiency, our method significantly reduces training time by optimizing the robot’s local behavior rather than learning from scratch, leading to rapid and responsive personalization in seconds.
This study investigates how school-aged children ($\mathbf{N} \boldsymbol{=} \mathbf{6 0}$, ages 8–12) from a non-WEIRD population in Egypt balance breadth and depth during self-directed exploration. Using a novel “TRIVIA” paradigm, children freely explored topics across three days, with their choices analyzed through linear and cumulative link mixed models. Results revealed that socioeconomic status (SES) significantly influenced exploration patterns: lower-SES children explored more questions on the second day compared to higher-SES peers ($\beta=-0.17, 95\% CI [-0.28, -0.05]$) and were more likely to visit previously explored questions ($\beta=-0.005, 95 \% CI [-0.009,-0.002]$). While age did not predict exploration quantity, children progressively shifted from in-breadth to in-depth exploration as the task progressed ($\beta=0.295,95 \% CI [0.092,0.498]$). This effect was moderated by SES, with higher-SES children’s exploration becoming more in-depth over time ($\beta=-0.008, 95\% CI [-0.012,-0.004])$. Motivation and active learning scores showed no significant effects. These findings highlight how SES and task structure shape children’s exploration patterns, offering insights into curiosity-driven learning in understudied cultural contexts.
The way in which humans perceive agency in robotic agents is dependent on a wide variety of factors. Whether that be the mannerisms, such as the way a robot can hold gaze, or the degree of relative relation to that of human function, such as a dog showing human mannerisms like ‘sit’. The subversion of these factors all feed in the Uncanny Valley effect; where if not properly accounted for, may lead to a sense of unease with these inanimate systems. This introduces the purpose of Hazel. Hazel is a novel zoomorphic robotic system able to effectively convey the emotional states of a domestic dog; including Happiness, Excitement, Anger, Fear and Possessiveness. Control through a neuromodulated system allows for these emotions to emerge appropriately; dictated by simulated concentrations of Cortisol, Adrenaline and Dopamine. These are key components of modulating emotional responses in dogs, providing the system biological grounding. To test the effectiveness of her ability to convey emotions and encounter the Uncanny, Hazel was shown to 100 children; aged 9 to 11, within their school. An evaluation was then conducted of Hazel’s effectiveness to express emotions, and the ability to perceive agency from these emotions, by gauging the manner in which the children interacted with her. Here, we found that all children were able to successfully identify emotional states in Hazel, with the younger children even believing that the emotions reflected her internal state. Questions by this age group suggested that with the emotional responses, they believed Hazel to have her own wants and desires, and showed affinity towards the system, overcoming the Uncanny Valley.
Human development relies on a fundamental mechanism that enables the creation of novel tasks and their solutions, shaping cognitive and motor learning. This mechanism allows for the progressive refinement of problemsolving abilities. Unfortunately, such a mechanism is missing in recent robotic systems except for several works in the area of intrinsic motivation. Thus, it is desirable to develop a robotic ability to automatically generate diverse and solvable tasks and validate their feasibility in a continuous space. In this work, we introduce a novel system called PRAG, which serves as a generative mechanism for constructing multi-step manipulation tasks. This tool mirrors cognitive developmental processes by autonomously producing novel, structured challenges that can be progressively solved through interaction. PRAG requires just a set of known atomic actions, objects, and spatial predicates (semantic knowledge) as a starting point to output solvable task sequences of specified complexity. Validation occurs in two stages: high-level symbolic validation ensures logical and operational consistency, akin to how cognitive development refines action representations, while physical validation confirms task feasibility in a robotic environment, resembling embodied learning in human development. The generated tasks provide structured training data, facilitating progressive learning through curriculum-based approaches, much like the way children build on prior knowledge to master increasingly complex motor and cognitive skills. We tested PRAG on sequences with increasing complexity and demonstrated its capacity to produce millions of unique, solvable tasks. By drawing parallels between developmental mechanisms and task generation, we propose that our framework can contribute to understanding how structured learning environments shape problem-solving abilities in both artificial and biological systems.
Infants develop visual abilities through different kinds of visual experience: for instance, they learn from playing with objects when at home and also when they see objects outdoors. Inspired by this dichotomy of experiences, we use machine learning (ML) to investigate how representations that bridge these two kinds of experiences can be learned. In ML parlance, the differences in the characteristics of these two experiences constitutes a distribution shift. ML research suggests that aligning learned representations across the two distributions can be beneficial for bridging the distribution shift. We use the supervised learning paradigm in this work; we assume we have labeled images from both distributions. In this paradigm, the most popular approach is to train a model to map different examples of a category to the same label, i.e., using a category-based learning signal. In contrast, we propose that comparison of image features across the two distributions can promote alignment, i.e., what we call a comparative learning signal. We propose 2 metrics to measure the cross-distribution alignment of image features. Using these metrics, we show that standard supervised learning using only categorical learning signals does not lead to aligned representations when training a popular neural network architecture (ResNet-18). Additionally, we show that using our comparative learning signal on only a small number of cross-distribution image pairs promotes alignment without affecting accuracy. Finally, we show that the diversity of images that form the image pairs contributes to the effectiveness of comparative learning signals.
This paper introduces Canalizing Babbling, a development-inspired approach for data collection in sensorimotor learning. The method draws inspiration from reflexes in newborns, which are here hypothesized to scaffold the acquisition of coordinated sensorimotor actions by facilitating the early experience of contingent sensory and motor events. In the presented approach, a visual saliency system selects targets in 3D space, guiding an inverse dynamics controller to generate coordinated movements across multiple body parts of the MIMo simulated agent. Statistical analysis shows that the visual and motor observations collected using Canalizing Babbling exhibit a higher degree of coordination compared to those obtained through non-curated strategies. These findings suggest that biologically inspired exploration techniques like Canalizing Babbling can lower sample complexity, potentially accelerate downstream learning in embodied agents, and provide a framework to further investigate the role of reflexes in developmental learning.
In humans, changes in the visual field or obstructions influence the perception and interpretation of facial emotions. A partial or altered visual area can lead to a misunderstanding of the facial expressions of the interlocutor. This problem also arises in the field of social robotics, where the visual perception of robots is often constrained by technical and environmental factors. With the rise of robots intended for the general public, this is a question that must now be taken into account in many common situations. In this study, we explore the impact of restricting the visual field of a robot equipped with developmental AI on its ability to recognize human facial emotions. We analyze how different limitations of the field of vision affect its reactions and modify its ability to recognize facial emotions. For this, we developed A architecture to explore this question, based on: (1) the recognition of primary emotions through an imitation game that enhances the robot’s ability to recognize emotions, and (2) an attention-focusing mechanism. The integration of these two components allows the robot to modulate its areas of interest before performing facial emotion recognition. This approach enables the exploration of various recognition scenarios, ranging from full-face analysis to a more targeted focus on the mouth—an area particularly emphasized by elderly individuals and bilingual children. Our results show that the restriction of the visual field leads to a decrease in the accuracy of emotion recognition, with variations depending on the type of expression and the degree of limitation of the visual field. These results are consistent with observations made in humans, suggesting the possibility of reproducing this mechanism in social robotics.
This paper explores gendered toy preference in parent-child interactions. We focus on free-play, which allows for unique natural and dynamic interactions in which toy preferences might be less constrained by the experimental setting. We operationalize toy preference through the child’s visual focus of attention (VFOA). Our analyses of 25 interactions of 12-13 minutes each reveal statistically significant differences between boys and girls in terms of time spent looking at a doll and a jump box. We then investigate whether these effects can also be obtained through automated analyses of the video data. To this end, we leverage an automated VFOA algorithm to predict which toys are attended to. Our automatic algorithm reveals similar patterns as when using manual annotations, albeit with less statistical power. This advancement holds promise for developmental research by providing efficient and objective assessments of children’s interactions, potentially guiding early developmental interventions and informing strategies to mitigate gender bias in play environments1.1https://github.com/Chelseapt/Auto-Parent-Child-Assessment
While recent advances in emergent language research have yielded significant insights into how artificial agents develop communication protocols, most studies have focused on reference games involving static objects, leaving the communication of dynamic concepts largely unexplored. This paper introduces a novel multi-agent framework for investigating how agents develop protocols to coordinate complex actions, such as navigation and object manipulation. Through systematic experimentation, we discover a fundamental property we term “prototype consistency” - the emergence of shared message prefixes for categorically similar tasks, resembling verb-like structures in natural language. Our ablation studies reveal that this linguistic organization emerges primarily from listeners’ learned behavioral patterns rather than from input representations or environmental cues. This finding provides empirical support for action-based theories of language, suggesting that fundamental aspects of linguistic structure may arise naturally from the requirements of coordinated behavior. Our results offer new insights into the relationship between action, cognition, and language in artificial systems, while providing a novel framework for investigating the origins of linguistic structure.
Infant motor development is characterized by high variability; during locomotion, spatial and temporal aspects of locomotor coordination seem to change on every step. Such continuous variability could serve an exploratory function analogous to exploration-driven reinforcement learning. However, we need further insight into the structure of this variability to enable careful comparison with computational models of learning and their underlying assumptions. In this study, we identify structured signatures of spatial and temporal variability in infant locomotor patterns. Natural walking data were collected from 16 children ($9-45$ months) and compared to 9 adults. Analyses revealed a surprising dissociation between spatial and temporal variability; while children, like adults, exhibit a strong speed-dependent stride length relationship, they exhibit a much weaker speed-dependent stride frequency relationship. This observed dissociation between spatial and temporal variability poses a question for models of locomotor learning: under what assumptions might models exhibit a similar dissociation? Future modeling studies could test multiple hypotheses to explain this observation such as distinct underlying mechanisms for spatial and temporal exploration or faster learning enabled by sequencing spatial and temporal exploration inspired by infants.
As social robots become household staples providing language input to infants, it is crucial to determine the conditions under which infants can learn from them. Here, we investigate whether infants process speech provided by a robot, and whether robot-provided social cues enhance their attention and speech processing performance. Specifically, we tested 8- to 13-month-old German-learning infants ($\mathbf{3 2 ~ n}$). We used a classical speech segmentation paradigm, modified such that a Furhat robot spoke the text. There were two familiarization conditions: one in which the Furhat recited text passages containing two target words while maintaining “eye contact” with the infants, and one in which the Furhat did not express any social cues. Afterwards, infants were tested for their recognition of target words vs. novel words. Linear mixed-effects models revealed a marginal preference in the condition without social cues, with longer looking times to the novel than to familiar words, but no preference in the condition with social cues. There was a marginal interaction between social cue condition and word type. Pearson correlations revealed no relationship between attention to the robot during the familiarization and speech segmentation performance. Although exploratory, the finding that infants in the control condition outperformed infants in the social cue condition raises questions about the effectiveness and naturalness of the provided cues, highlighting the importance of improving these aspects in future robot designs.
Children’s interactions with peers are central to their social, emotional, linguistic, and cognitive development. While research has focused on children’s pairwise associations, less attention has been given to the characteristics of larger groups of children. This study highlights the importance of a group perspective by exploring whether homophily—the tendency to interact with similar others—extends beyond pairs to groups of children. To identify pairwise interactions between children, we use social contact criteria, and for interacting groups, we operationalize them as F-formations, a widely used concept in computational research for detecting adult groups. We conducted a case study in an inclusive preschool classroom over two consecutive years, involving two different cohorts of children with hearing loss (HL) and with typical hearing (TH). Our findings show that children tend to form groups, particularly groups of size 3 to 5. Homophily is evident among children with TH in both pairs and groups, while for children with HL, homophily is suggested only at the group level. Group-level analysis also reveals patterns not observed in pairwise interactions, such as TH children’s lower overall likelihood of being in a group and their tendency to associate with groups containing a higher proportion of TH peers. These findings suggest that incorporating group interactions provides a more comprehensive understanding of children’s sociality, capturing patterns that cannot be explained by pairwise analysis alone.
This systematic integrative review examines how scaffolding, grounded in Bruner’s foundational functions, is operationalized in early childhood interventions. We analyzed 20 peer-reviewed studies published between 2005 and 2024, focusing on three primary conceptualizations-fading, enrichment, and comprehensive models-and their developmental impacts. Early childhood represents a critical window for cognitive, language, and socio-emotional growth; however, effective interventions often face challenges in resource-limited settings and heterogeneous educational environments. Our synthesis identified three dominant approaches. The fading approach, characterized by gradual withdrawal of support, consistently yielded higher mean effect sizes and more robust developmental gains. Enrichment strategies emphasized rich learning experiences but produced more variable outcomes. Comprehensive models aimed to incorporate all of Bruner’s functions; however, measurable impacts were often more diffuse. We highlight the pivotal roles of outcome measure sensitivity and intervention fidelity in capturing true intervention effects. Emerging technologies such as artificial intelligence (AI) offer transformative potential. AI’s capacity for continuous learner monitoring and real-time parametric adjustments positions it to implement fading strategies with precision, dynamically adapting support as competence increases. By merging sociocultural learning principles with computational power, AIenhanced scaffolding can foster highly individualized learning trajectories. These insights underscore the need for standardized reporting guidelines and longitudinal studies to optimize early childhood interventions, particularly as AI integration accelerates. Our review provides a nuanced understanding of scaffolding’s operationalization across contexts and over time and suggests future research directions for harnessing AI-enhanced scaffolding in early learning.
This paper describes a novel closed-loop manipulation approach based on internal simulation and synthetic perception rendering. In particular, we propose the use of an internal simulator as a forward model that builds synthetic perception that acts as an error signal to refine the robots’ world representation through an inverse model, ultimately improving its ability to manipulate objects. We first describe our baseline implementation of a comprehensive pipeline for open-vocabulary object detection, segmentation, and modeling tailored for the Guayabo domestic robot. Then, combining concepts from Intuitive Physics and internal models, we introduce our new closed-loop control approach. We show that by using this correction mechanism, our baseline object grasping method goes from a $\mathbf{5 4 \%}$ success rate to $\mathbf{9 2 \%}$.
There is an established body of research showing that both the quality and quantity of caregiver input plays an integral role in infants’ language development. Equally, infants play an active role in eliciting that input, with their interests also guiding what they learn. In this study, we examined monolingual German children to see whether caregivers’ perception of their child’s interest predicts the quality of caregiver speech as well as infant vocalisations during a shared book-reading interaction at two timepoints, 18 months and 24 months. We also aimed to examine longitudinal changes in prosodic qualities of infant-directed speech over infant development. Caregivers read two books to their child, one of which they rated to be of high interest to their child, while the other was rated to be of low interest to the child. We measured the pitch range, mean pitch and duration of caregiver utterances as well as the number of infant vocalisations during the book reading sessions. We found that the duration of utterance varied significantly, with caregivers producing shorter utterances when reading high interest books to their child at $\mathbf{1 8}$ months and producing longer utterances reading high interest books to their older, $\mathbf{2 4}$-month-old infant. Infants also vocalised more during high interest books, which we interpret in terms of their active elicitation of information during reading. Longitudinally, we observed that pitch range increased and mean pitch decreased (approaching adultdirected speech), aligning with previous research on infantdirected speech. These findings help underscore the dynamic feedback loops of caregiver-child interactions which play a pivotal role in early development, with children seeking information and caregivers responding contingently.
This study investigates the use of Large Language Models (LLMs) in Reinforcement Learning (RL) to mimic human adaptive reasoning, and hence the ability to adjust cognitive effort to task complexity. Inspired by the dual-system theory of cognition, our approach employs a pre-trained LLM to guide an RL agent between “fast” and “slow” thinking modes. These modes are policies that correspond to either rapid, intuitive responses or more calculated, deliberate actions. Our main contribution is demonstrating the LLM’s ability to dynamically select the appropriate mode for varying situations in a custom grid-world environment, without requiring additional training or fine-tuning. Moreover, we demonstrate that the LLM’s performance is comparable to that of a Q-learning agent trained for the same purpose. This research presents a proof-of-concept integration of LLMs in RL, offering preliminary insights into their potential for more complex decision-making.
Change occurring across multiple timescales, from milliseconds to year-long periods, is inherent to any developmental process. Short-time behavioral fluctuations on the scale of minutes illustrate flexibility and calibration to the environmental context; however, relatively little is known about how such behavior unfolds in less controlled conditions. The current study investigates how infants’ manual object sampling movements change within a short, spontaneous, free-flowing play session in a laboratory setting. Nine-month-old infants participated in a free-flowing dyadic play session with their caregiver, where they were free to move and interact. The analysis focused on how the duration of a sampling episode with an interactive object (a button-press toy that elicited visual feedback) changed during the 5-minute-long task. A subset of infants (n = 51) engaged with the object in at least three separate 30-second windows. Their sampling episodes were categorized into these windows to assess changes in sampling duration over time. We predicted that infants would sample the object for shorter durations over time, reflecting within-task movement calibration to better match the object’s properties. However, contrary to our expectations, a linear mixed model indicated no significant differences in sampling duration, which was around 1 second across the three windows. These findings suggest that 9-month-old infants maintained a consistent sampling pattern throughout the task, with no major changes in duration. One possible interpretation is that brief, 1-second-long sampling episodes functioned as a behavioral attractor, guided by pre-existing movement strategies suited to the object’s properties. While more research is needed to fully understand how infants adapt their actions in real time, the current study is an initial step in examining within-task calibration in infants’ self-initiated manual sampling.