Theoretical accounts typically assume that key features of human socio-cognitive development are universal. This paper reports a large-scale cross-cultural study (17 communities, diverse ethnicities, N = 1,377, 709 female, mean = 5.50 years, -collected March 2022 to January 2024) on gaze following in early childhood. To test for universality, cognitive processing signatures were derived from a computational model treating gaze following as social vector estimation. Results showed substantial variation between communities and individuals. Yet, the processing signature was found in all communities. Individual differences in performance were related to children's familiarity with the data-collection device but not opportunities for social interaction. These results provide strong evidence for gaze following as a universal socio-cognitive process despite cultural and individual-level variation in absolute performance.
Large scale studies have documented socioeconomic (SES) and racial/ethnic disparities in children's standardized math achievement at kindergarten entry. These early math skills predict future mathematics achievement and career success. However, limited research has been conducted using large sample sizes to understand how SES and race/ethnicity are related to children's numerical skills at even younger ages. The current study aims to investigate sociodemographic variability in three fundamental areas of early numeracy: nonverbal numerosity discrimination, rote counting, and cardinal number word knowledge. In addition, we will examine if the relations between numerical skills might be explained by their shared correlations to sociodemographic factors and if differences in numerical skills between sociodemographic groups can be explained by variability in working memory. Finally, we also investigate whether childcare attendance moderates early sociodemographic differences in numerical abilities. To achieve these goals, data from children aged 2; 6-6; 0 will be gathered from ∼ 45 US sites, drawn from a larger multi-lab international project (ManyNumbers project). The findings of this research will enhance our understanding of early emerging variability in numerical skills and provide insights into developing responsive and inclusive educational practices that support diverse learning needs in the early years. SUMMARY: Early mathematical skills are crucial for long-term academic and career achievement. SES and race/ethnicity-related disparities in math achievement emerge as early as preschool. Most studies use standardized math assessments that combine different numerical skills to assess achievement gaps, leaving uncertain which specific skills vary with demographic variables. We explore disparities in developmentally significant numerical skills and their relation to demographic variables. We also report relations between WM, childcare attendance and numerical skills. Data from approximately N = 1080 children aged 2;6-6;0 will be collected from ∼ 45 US labs, including demographic information and numeracy measures.
Measuring children’s early vocabulary is crucial for developmental science and practical applications. The MacArthur–Bates Communicative Development Inventories (CDIs) are the most widely used caregiver-report checklist to assess children’s vocabulary across languages and dialects. Although full-length CDIs enable detailed analyses, completing checklists of several hundred words requires considerable time and effort from caregivers, often unnecessarily. Here, building on recent psychometric advances utilizing item response theory (IRT) in languages such as English and Spanish, we developed a computerized adaptive test of the Japanese CDI (Japanese CDI-CAT), which estimates total productive vocabulary size of the full-length CDIs from a reduced number of word items. Using a publicly available Japanese CDI dataset (N = 978) from children aged 12 to 38 months, we compared several IRT models and conducted simulations while varying the number of items administered (25–400). The results showed that the two-parameter logistic IRT model provided the best fit among the candidates. We also found that the CDI-CAT can successfully estimate children’s total vocabulary size while substantially reducing the number of items (e.g., r = .95 with the full-length CDI using only 25–50 items). By extending the CDI-CAT methodology to Japanese, we aim to improve the efficiency of vocabulary assessment and contribute to a more inclusive understanding of early language development across diverse linguistic and cultural contexts.
Children navigate a world full of auditory inputs that include both signal and noise, but what counts as signal vs. noise is dependent on their goals. We tested whether children recognize that particular acoustic contexts are consistent or inconsistent with goals, for example that music is a good environment for dancing but a bad environment for sleeping? In a series of experiments, we presented 3-5-year-old children (n = 168; 75 boys; 55.8% Caucasian/White (including Hispanic/Latino), 19.6% multiracial, 19.0% Asian/Pacific Islander, 2.2% African/Black, 0.05%, Native/Indigenous Peoples, 1.1% other, 1.0% not given) with auditory stimuli and asked them to select the best environment for each activity. By ages 3–5 years, children show some understanding of how noise in the auditory environment affects activities, and the robustness and flexibility of their reasoning about environmental noise improves during this period.
Given the inherently multimodal nature of human experience, vision-language models (VLMs) hold substantial promise for modeling human cognition as it grows and develops with experience. Realizing their potential requires tools for comparing VLMs with human cognitive development across tasks, ages, and populations. We present LEVANTE-bench, a benchmark based on tasks and data from the Learning Variability Network (LEVANTE), which distributes open-source tasks and data measuring children's cognition across languages and cultures. In LEVANTE-bench, we systematically assess VLMs on six tasks, comparing their alignment with children aged 5-12 (N = 1547) across three countries. We compare models at multiple scales, assessing their overall accuracy, their task- and item-level alignment with children, and how well they match children's trial-level error distributions. Alignment was heterogeneous across scales: at the level of tasks and items, more capable models aligned better with humans. However, match to human error distributions varied widely across tasks, and for several tasks, smaller models matched younger children's errors better. In addition, even the best-performing VLMs struggled on matrix reasoning and mental rotation tasks. Thus, current VLM architectures align only partially with the cognitive abilities of children.
Natural languages have been argued to evolve under pressure to efficiently compress meanings into words by optimizing the Information Bottleneck (IB) complexity-accuracy tradeoff. However, the underlying social dynamics that could drive the optimization of a language's vocabulary towards efficiency remain largely unknown. In parallel, evolutionary game theory has been invoked to explain the emergence of language from rudimentary agent-level dynamics, but it has not yet been tested whether such an approach can lead to efficient compression in the IB sense. Here, we provide a unified model integrating evolutionary game theory with the IB framework and show how near-optimal compression can arise in a population through an independently motivated dynamic of imprecise strategy imitation in signaling games. We find that key parameters of the model – namely, those that regulate precision in these games, as well as players' tendency to confuse similar states – lead to constrained variation of the tradeoffs achieved by emergent vocabularies. Our results suggest that evolutionary game dynamics could potentially provide a mechanistic basis for the evolution of vocabularies with information-theoretically optimal and empirically attested properties.
An abstract understanding of communication should support reasoning about both its success and failure: why it fails, what happens as a consequence, and how to fix it. Auditory noise frequently corrupts verbal communication, but little is known about how humans come to reason about it. The current work explored American 3- to 5-year-olds' third-party reasoning (Experiment 1, N = 168, 95 female) and communicative behaviors (Experiment 2, N = 48, 23 female) in noisy environments between 2021 and 2024. Children understood that auditory noise impedes others' hearing and prevents knowledge transmission, and they modified their own communication by gesturing more when their partner could not hear. Thus, even young children understand how noise disrupts communication and can communicate effectively in its presence.
Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning process look like? We analyzed first-person videos of young children's visual experience at home from the BabyView dataset (N = 31 participants, 868 hours, ages 5–36 months), using a supervised object detection model to extract common object categories from more than 3 million frames. We found that children's object category exposure was highly skewed: a few categories (e.g., cups, chairs) dominated children's visual experiences while most categories appeared rarely, replicating previous findings from a more restricted set of contexts. Category exemplars were highly variable: children encountered objects from unusual angles, in highly cluttered scenes, and partially occluded views; many categories (especially animals) were most frequently viewed as depictions. Surprisingly, despite this variability, detected categories (e.g., giraffes, apples) showed stronger groupings within superordinate categories (e.g., animals, food) relative to groupings derived from canonical photographs of these categories. We found this same pattern when using high-dimensional embeddings from both self-supervised visual and multimodal models; this effect was also recapitulated in densely sampled data from individual children. Understanding the robustness and efficiency of visual category learning will require the development of models that can exploit strong superordinate structure and learn from non-canonical, sparse, and variable exemplars.
Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research suggests current vision-language models (VLMs) trained on curated web data fail to generalize to the sparse, weakly-aligned egocentric streams produced by wearable devices, embodied agents, and infant head-cams – and no fixed evaluation pipeline exists for measuring progress on this regime. We train VLMs on datasets with varying degrees of semantic alignment between visual and linguistic inputs, including naturalistic infant and adult egocentric videos, and evaluate them with a comprehensive suite spanning multimodal language grounding and unimodal vision and language tasks. At the core of this suite is Machine-DevBench, a corpus-grounded benchmark of lexical and grammatical competence, automatically generated from the model's training vocabulary across logarithmic frequency bins to eliminate the train/eval mismatch and low statistical power of prior developmental benchmarks. Our results show that current VLM paradigms hinge on the tight semantic alignment of curated data and fail to exploit the weakly-aligned signal that dominates naturalistic egocentric input – the very regime in which humans thrive. To motivate progress, we introduce the EgoBabyVLM Challenge to drive the development of models capable of grounded language learning from the kind of naturalistic data that human infants experience.
Emotion displays are one of the most salient features of early social interactions - caregivers often smile, change their vocal register to convey positive emotions to their children. Despite the fact that these emotion displays often accompany the moments in which children learn language, prior work on the role of emotion in word learning has been sparse and has resulted in conflicting findings. The present investigation examines how emotion displays that vary in valence (positive vs. negative) and arousal (high vs. low) influence word learning in preschoolers and adults. Across five experiments, participants completed a word learning task in which novel object-label pairs were introduced by a character displaying one of four emotional states: high-arousal positive, low-arousal positive, high-arousal negative, or low-arousal negative. Results consistently showed that positive emotion displays facilitated word learning relative to negative displays in adults and preschoolers alike, with web-based eye tracking indicating that high-arousal negative displays modestly pull attention away from the labeled object. Following the study, preschoolers preferentially chose stickers depicting objects that had been paired with positive displays as a prize, suggesting that emotion signals shape not only which words children learn but also how they value the objects paired with emotion displays. Together, these findings establish that emotional signals are not merely a backdrop to early language input but an active ingredient in word learning. Specifically, positive emotion in caregivers’ communication to young children may serve a functional role in supporting language development.
Language users flexibly produce and interpret language in context, and over time modulate their language use to convey meanings with greater accuracy and efficiency. Adaptive language use is often studied using iterated reference games, which represent a controlled way of studying how people communicate in the absence of codified labels, and how conventionalised labels form over repeated reference to the same targets. While iterated reference games are a common paradigm, the research questions, pre-processing, and analytic approaches differ across papers so their results are challenging to synthesise using meta-analysis. Here, we present Refbank, a harmonised repository of iterated reference game data from 20 different papers, enabling quantitative synthesis in novel and larger-scale analyses. We conduct five mega-analyses that shed light on the robustness and variability of convention formation under different experimental settings. Refbank also has potential future uses for targeted experiment design, computational modelling, and language model evaluation. These directions demonstrate the utility of reusing data from iterated reference games, and we invite contributions of additional datasets to Refbank.
Early language skill is predictive of many later life outcomes and is thus of great interest to developmental psychologists and clinicians. The MacArthur-Bates Communicative Development Inventories (CDIs), parent report instruments typically containing inventories of hundreds of children's vocabulary words, have proven to be valid and reliable instruments for measuring children's early language skill. The CDIs have been adapted to many dozens of languages, and cross-linguistic comparisons show both consistency and variability in language acquisition trajectories. However, thousands of languages do not yet have CDIs, nor the early language corpora needed to create them, posing a significant barrier to increasing the diversity of languages that are studied. Here, we propose a method for selecting candidate words to include on new CDIs through analyzing psychometric properties of the translation-equivalent concepts that are frequently included on existing CDIs. Leveraging 32 datasets from existing CDIs, we propose a list of 100 concepts that have low variability in their cross-linguistic learning difficulty. This pool of common concepts-analogous to the Swadesh lists, which are basic vocabulary lists used in glottochronology for cross-language comparison-can be used as a starting point for future CDI adaptations. We show that the proposed Swadesh-CDI list generalizes well to data from 10 additional languages.
A decade of ManyBabies research, testing thousands of babies across hundreds of labs, has shown that some, but not all findings in infant research replicate well. Collectively, these projects have shown us that our methods carry limitations that larger samples alone cannot resolve. Here we present three lessons that point toward a more reliable, inclusive developmental science.
Being a fluent language user involves recognizing words as they unfold in time. How does this skill develop over the course of early childhood? And how does facility in word recognition relate to the growth of vocabulary knowledge? We address these questions using data from Peekbank, an open database of experiments measuring children’s eye movements during early word recognition. In an observational study of 26 datasets from over 2500 children ages 6 months to 6 years, we show that word recognition becomes faster, more accurate, and less variable across development, consistent with a process of skill learning. Factor analysis reveals covariation of word recognition speed and accuracy with children’s vocabulary size in cross-sectional analysis. Further, across a range of longitudinal models, speed, accuracy, and vocabulary were coupled. Children with overall faster word recognition tended to show faster vocabulary growth, though developmental growth in word recognition skill was not specifically associated with growth in vocabulary. Together, these findings support the view that word recognition is a skill that develops gradually across early childhood and that this skill is deeply intertwined with early language learning.
Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than they did from the one before. In contrast to children, language models – even those trained on child-directed speech – do not accelerate. Instead, they show constant proportional returns on new data, consistent with scaling laws. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation.
Inclusive global theories of child development require local researchers who spearhead research within their own communities. This means creating pathways for researchers from underrepresented contexts to lead projects, define priorities, and collaborate as equal partners, rather than merely act as data collectors. This capacity-building holds the promise of generating more accurate and equitable developmental science research.
Young children demonstrate early abilities to understand their physical world, estimating depth, motion, object coherence, interactions, and many other aspects of physical scene understanding. Children are both data-efficient and flexible cognitive systems, creating competence despite extremely limited training data, while generalizing to myriad untrained tasks – a major challenge even for today's best AI systems. Here we introduce a novel computational hypothesis for these abilities, the Zero-shot Visual World Model (ZWM). ZWM is based on three principles: a sparse temporally-factored predictor that decouples appearance from dynamics; zero-shot estimation through approximate causal inference; and composition of inferences to build more complex abilities. We show that ZWM can be learned from the first-person experience of a single child, rapidly generating competence across multiple physical understanding benchmarks. It also broadly recapitulates behavioral signatures of child development and builds brain-like internal representations. Our work presents a blueprint for efficient and flexible learning from human-scale data, advancing both a computational account for children's early physical understanding and a path toward data-efficient AI systems.
Learning language requires learning not only the content of language, but also how to use language to communicate. Iterated reference games provide a window into such skills, requiring rich communication as participants converge on mutually understandable names for initially novel referents. Some early experiments are interpreted as evidence that 4-5-year-old children cannot converge on the mutually understandable names needed to succeed in an iterated reference game. Here, we revisit young children's referential communicative abilities using a simpler, child-friendly version of the iterated reference game paradigm. Across 3 experiments and 82 pairs of children (N=158), we found that 4-5-year-olds could successfully establish reference with each other. Children were 81% accurate, and they often used descriptions similar to their partner's. These findings suggest that children’s capacity to construct effective referring expressions in novel contexts emerges earlier than previously thought, consistent with the view that children show early pragmatic competence in supportive contexts.
The effectiveness of social learning depends on whether learners receive help when they need it. In four preregistered studies, U.S. 4–6-year-olds (N = 244; 54% female, 27% White, 2% Black, 48% Asian, 9% Hispanic/Latino, 24% Multiracial/Other) interacted with an adult who either did or did not follow through on promised help. Experiment 1 tested the effect of reliable versus unreliable help on children’s future task choice; Experiment 2 examined its effect on children’s help-seeking and exploration of a novel toy. Children’s learning goals and strategies were modulated by the past reliability of help, suggesting that seemingly maladaptive decisions—such as avoiding a hard task—may be adaptive responses that balance the reliability of help against the utility of exploring alone.