Function words like `I' or `some' are highly frequent in children's linguistic input but are not learned early, so we ask: do they play any role in early word learning? Here we test a novel hypothesis that demonstratives, a class of function words like THIS and THAT in English, generate word learning opportunities for young children because they provide multimodal cues to joint attention on labeled objects. Combining head-mounted eye-tracking with machine-learning modeling of naturalistic interactions between parent and child (19 months), we show that demonstratives provide augmented inputs that scaffold word learning. Our results suggest that these function words orchestrate attentional dynamics in multimodal social interaction, expanding long-standing learning theories and showing that children can learn words from the full richness of everyday experience.
One of the primary aims of cognitive and behavioural sciences is to generate empirical findings that are both reproducible and generalizable to real-world settings. The present study investigates the extent to which results obtained from a structured laboratory task-parent-infant toy play-can be generalized to more naturalistic contexts and everyday activities. We focused on joint attention between parents and infants, a construct that has been extensively examined within developmental science. To characterize parent-infant joint attention during toy play, we recorded and analysed contingent gaze behaviour captured through dual head-mounted eye-tracking devices worn simultaneously by parents and their infants during spontaneous activities such as toy play and meal preparation. By continuously monitoring gaze locations and manual actions, we obtained fine-grained measures of how often dyads fixated on the same object concurrently and how their coordinated visual and manual behaviours contributed to the establishment and maintenance of joint attention. Our results suggest that laboratory findings can be both replicated and generalized when: (i) the study is designed to capture natural behaviours rather than to elicit specific, constrained responses, and (ii) the theoretical constructs are clearly defined and precisely measured through high-resolution behavioural data. This article is part of the theme issue 'Mechanisms of learning from social interaction'.
Is the sense of touch a mechanism for human babies' learning of visual concepts? If so, can we quantify its importance, and to what extent do babies rely on their sense of touch for visual learning? To approach these questions in a principled way, we propose a structured coding system for baby-centric touch events, yielding a dataset of 264k two-second clips of touch events coded according to this system. Using this dataset, we pretrain developmentally grounded models that reveal promising insights into the nature of baby learning from touch.
Infants' early learning is fast and efficient. Most research on infant learning from social contexts has focused on child-directed activities, such as toy play. However, learning in everyday contexts may not be limited to child-directed activities. Routines at home consist of other activities, such as meal preparation, laundry, and chores, that may still offer learning opportunities for infants. But infants can only learn from these moments if they are paying attention. Here, we provide the first examination of infant attention during non-child-directed everyday activity. In a home-like laboratory environment, parents prepared peanut butter and jelly sandwiches in front of infants who sat in a highchair across from them, as they might every day at home. Head-mounted eye trackers were used to record gaze data from infants, which was then temporally aligned with parent actions to measure infants' attention to the task. Our findings reveal that, despite the presence of a room full of distractors, infants most often focused on task-related objects. Although infants only tracked parents' actions for 40% of the session, they prioritized looking at parents' more complex actions and looked equally at parents' actions and the objects they themselves were engaged with. In addition, when parents were more engaged with their infants (e.g., looking at them and talking to them), infants paid more attention. These findings indicate that infants are attuned to the routines and activities of daily life and that parental engagement can play a critical role in enhancing this attention. This suggests that even routine tasks, often overlooked as learning opportunities, can be transformed into valuable moments for development. By examining how infants allocate their attention in naturalistic settings, we can better understand the role of everyday experiences in early learning and explore how parents can optimize these moments to support their infants' growth. SUMMARY: Head-mounted eye tracking captured real-time infant gaze during non-child-directed meal preparation in a home-like laboratory setting. Infants attended primarily to task-relevant objects and actions, even amid numerous visual distractions. Infants prioritized observing parents' complex, effect-producing actions over simple object handling. Everyday routines may serve as rich, natural input supporting infants' observational learning.
Recent latent visual reasoning methods achieve substantial gains by inserting continuous latent tokens into multimodal language models. These gains are commonly attributed to the tokens encoding visual evidence; recent analyses, however, reveal a paradox: the tokens are loosely tied to the image and contribute little to the answer. Critically, these analyses treat latent tokens as a single unit, obscuring the source of the gains. We therefore decompose latent tokens into three testable components: latent slots, boundary markers, and format, and develop a state-of-the-art method as a probe under favorable conditions. Across six method-stage settings and four perception-heavy benchmarks, latent slots fail every prediction of the visual-memory account. Strikingly, retaining only the boundary markers preserves 78 to 100
Joint attention (JA) is a critical developmental process, yet it has typically been studied through gaze-following experiments. However, infants rarely look at faces during play, suggesting that JA relies on multiple pathways. In 2 naturalistic studies-free play (N = 53, Mage = 16.70 months; 2018-2019; working-middle socioeconomic status [SES], White non-Hispanic families) and meal preparation (N = 42, Mage = 16.36 months; 2019-2023; mostly middle-high SES, White non-Hispanic families)-dyads used different pathways to achieve similar amounts of JA. Although pathway use varied across contexts, nearly all dyads used multiple strategies, and JA duration did not differ by pathway. These findings suggest it is not the specific pathway that matters, but the ability to flexibly coordinate attention using diverse behavioral cues.
Toddlers learn to recognize objects from different viewpoints with almost no supervision. During this learning, they execute frequent eye and head movements that shape their visual experience. It is presently unclear if and how these behaviors contribute to toddlers' emerging object recognition abilities. To answer this question, we here combine head-mounted eye tracking during dyadic play with unsupervised machine learning. We approximate toddlers' central visual field experience by cropping image regions from a head-mounted camera centered on the current gaze location estimated via eye tracking. This visual stream feeds a neural network model, which uses a biologically plausible unsupervised learning objective. Our experiments demonstrate that a few minutes of such first-person experience suffice to learn strong object representations permitting invariant object recognition. Importantly, by simulating alternative gaze behaviors we show that toddlers' eye movement patterns play a crucial role in this. Our analysis also reveals that the limited size of the central visual field where visual acuity is high plays an important role for successful learning. Together, this highlights the benefits of temporally structured visual experience arising from toddlers' natural interactions with objects. SUMMARY: We combine recordings of toddlers' first-person central visual field experience with biologically inspired self-supervised learning algorithms to model toddlers' development of invariant object recognition. Just a few minutes of toddlers' central visual field experience captured with head-mounted eye tracking suffice to learn strong object representations. Simulated alternative gaze behaviors produce weaker representations, demonstrating the importance of toddlers' active gaze strategies for learning. Our results emphasize the importance of toddlers' eye movements for learning object representations.
Though ample research and theory suggest that parents’ beliefs and cognitions are important predictors of their parenting behaviors, there is little understanding of Latino fathers’ perceived parenting motivations. We explored resident, first-time fathers’ motivations to be involved parents in a sample (N = 85) of socioeconomically diverse Latino fathers participating in a parenting intervention in the Washington D.C. area and southern California. Data were collected through structured interviews that were recorded during home visits when infants were 18-months old. Bilingual research assistants transcribed and translated into English fathers’ responses to the interview question, “What makes you want to be a good parent?” A thematic analysis revealed five main emergent themes: (1) personal rearing history, (2) desire to rear a well-adjusted child, (3) relationship with their child, (4) intrinsic motivations, and (5) sense of duty and responsibility. We further explored whether fathers’ perceived parenting motivations varied by their nativity status (i.e., U.S.-born or immigrant). We found variations in each of the themes, including that immigrant Latino fathers were more likely to prioritize their children’s morals and values, whereas U.S.-born Latino fathers emphasized their child’s future success. This study contributes to the limited research on Latino fathers’ parenting perceptions and beliefs. The findings can be used to inform programs geared at strengthening Latine family functioning in the face of adversity through leveraging the reasons behind why fathers want to be positively involved with their young children. This study used thematic analysis to explore the perceived parenting motivations of 85 first-time Latino fathers living in the U.S. We also explored variation in fathers’ parenting motivations by their nativity status (i.e., U.S.-born vs. immigrant). We identified five main themes, and various subthemes, in fathers’ perceived motivations, as well as differences by nativity status. The findings may inform parenting interventions, given the influence of fathers’ motivations on parenting behaviors.
Longitudinal measurements of brain structure and function are critical for understanding how humans change over time. Traditional longitudinal approaches sample sparsely across large windows of time to estimate coarse, long-term brain changes. This review showcases insights from dense longitudinal neuroimaging (DLN), an emerging approach that samples densely across relatively short windows of time to precisely estimate individual trajectories of brain change. DLN measures multiple samples from individuals throughout critical periods of rapid change. It allows precise estimates of nonlinear trajectories to advance a mechanistic understanding of brain change. Novel findings from this approach are improving our understanding of human cognition, such as the role of the motor system in visual development and learning.
Learning the meaning of a verb is challenging because learners need to resolve two types of ambiguity: (1) word-referent mapping-finding the correct referent event of a verb, and (2) word-meaning mapping-inferring the correct meaning of the verb from the referent event (e.g., whether the meaning of an action word is TURNING or TWISTING). The present work examines how adult learners solve this challenge by utilizing both in-the-moment linguistic information within individual learning situations and cross-situational statistical information across multiple learning situations. We investigate how different cues provided in the moment affect information selection and how cross-situational learning as a general computational mechanism allows for information integration over time. Two experiments were designed based on a Human Simulation Paradigm, in which adult learners were presented with a sequence of short videos from parent-toddler toy play and asked to guess a mystery verb the parent produced in each video. In Experiment 1, we compared individual learning situations containing linguistic information to the exact same learning scenes without linguistic information and found that linguistic information helped learners narrow down the meaning of a verb embedded in individual situations, which was consistent with prior research. In Experiment 2, the videos sharing the same target verb were presented in a blocked design to incorporate cross-situational statistics for the same verb. We measured the variability, convergence, and accuracy of participants' guesses. Within-trial linguistic information allowed learners to quickly narrow down their search space and focus on a few relevant aspects in a scene, while cross-situational learning allowed them to fine-tune their learning further across trials. Our findings support a unified account wherein within-trial linguistic information and cross-situational statistical information are integrated for more efficient verb learning.
According to the cross‐situational learning account, infants aggregate statistical information from multiple parent naming events to resolve ambiguous word‐referent mappings within individual naming events. While previous experimental studies have shown that infant and adult learners can build correct mappings based on statistical regularities encoded in multiple learning situations in an experiment, other studies that use more naturalistic stimuli (e.g., real‐world video) reveal poor performance in adults' ability to infer the correct referent. Based on those results derived from more naturalistic stimuli, the cross‐situational learning solution cannot be useful to solve the mapping problem in the real world because cross‐situational statistics from the real world are much more ambiguous than those created in experimental studies. To examine the feasibility of cross‐situational learning in everyday contexts, the present study aims to quantify visual‐audio statistics from one of everyday activities—parent–child toy play. We analyze parent naming events in a video corpus of infant‐perspective scenes during parent–child toy play in a naturalistic lab setting, where we found three distinct properties that characterize statistical regularities perceived by young learners: (1) there are a limited number of visual scene compositions perceived by young learners at the moments when they hear object names; (2) the frequencies of parent naming events are distributed in a skewed, Zipfian fashion; and (3) cross‐situational statistics in naturalistic toy play are comparable to those used in laboratory experiments. Our results underscore the importance of quantifying the statistical regularities in the input from the learner's perspective in order to shed light on the mechanisms supporting early word learning.
Recent self-supervised learning models simulate the development of semantic object representations by training on visual experience similar to that of toddlers. However, these models ignore the foveated nature of human vision with high/low resolution in the center/periphery of the visual field. Here, we investigate the role of this varying resolution in the development of object representations. We leverage two datasets of egocentric videos that capture the visual experience of humans during interactions with objects. We apply models of human foveation and cortical magnification to modify these inputs, such that the visual content becomes less distinct towards the periphery. The resulting sequences are used to train two bio-inspired self-supervised learning models that implement a time-based learning objective. Our results show that modeling aspects of foveated vision improves the quality of the learned object representations in this setting. Our analysis suggests that this improvement comes from making objects appear bigger and inducing a better trade-off between central and peripheral visual information. Overall, this work takes a step towards making models of humans' learning of visual representations more realistic and performant.
Inductive reasoning is an essential capability for large language models (LLMs) to achieve higher intelligence, which requires the model to generalize rules from observed facts and then apply them to unseen examples. We present Mirage, a synthetic dataset that addresses the limitations of previous work, specifically the lack of comprehensive evaluation and flexible test data. In it, we evaluate LLMs' capabilities in both the inductive and deductive stages, allowing for flexible variation in input distribution, task scenario, and task difficulty to analyze the factors influencing LLMs' inductive reasoning. Based on these multi-faceted evaluations, we demonstrate that the LLM is a poor rule-based reasoner. In many cases, when conducting inductive reasoning, they do not rely on a correct rule to answer the unseen case. From the perspectives of different prompting methods, observation numbers, and task forms, models tend to consistently conduct correct deduction without correct inductive rules. Besides, we find that LLMs are good neighbor-based reasoners. In the inductive reasoning process, the model tends to focus on observed facts that are close to the current test example in feature space. By leveraging these similar examples, the model maintains strong inductive capabilities within a localized region, significantly improving its deductive performance.
Picture book reading is an important word learning context for children. Using eye-tracking methodology, we studied infants’ attention to objects when parents named them. We took a novel perspective by examining the statistical regularities that are created when infants attend to the named target object and to non-target objects. We found that the objects infants attended to during parents’ naming were more semantically similar to the named object. Furthermore, we observed a reinforcing loop between parents’ naming and children’s attention: When infants attended to non-target objects, parents were more likely to name this object shortly after, and this increased infants’ gaze to the new target object. Moreover, the more adept parents were at naming the nontarget objects their infants attended to, the higher their infants’ overall attention levels during naming. Together, this study considers an often overlooked input for word learning, showing that attention to non-target objects may provide additional statistical regularities in space and time that infants can use to flexibly build word knowledge. The results also have implications for how parents can support word learning in book reading contexts and may contribute to the design of artificial intelligence that learns in more human-like ways.
To learn a word from an everyday context, infants need to be able to link the heard word with the correct object perceived. A prevailing view of the early learning environment is that infants' world is bombarded with objects and words. Therefore, it is difficult to find the named object from many possible candidates. However, building correct word-referent mappings relies on in-moment visual selection, it is not clear what infants attend to when learning words in a naturalistic context. Toward this goal, we conducted an eye-tracking experiment in which 12-month-old infants were presented with complex visual scenes extracted from infants' egocentric videos recorded during naturalistic parent-child toy play. These scenes were selected at naming moments when parents labeled a toy object during free-flowing play. We selected visual scenes from a mix of more or less ambiguous naming events that contained different visual properties of the named objects and measured infants' real-time object-looking behaviors. We found that, despite the different visual properties of infants' egocentric scenes, early visual attention is both selective and variable. Selective visual attention is highly constrained by the visual saliency of the learning scenes, but not influenced by labels or existing word knowledge. Infants are more likely to attend to the named object when it is salient in the egocentric view. Our results suggest that although infants' naturalistic learning environment appears to be messy in terms of the number of possible objects for a heard object name, their selective attention significantly reduces the in-moment uncertainty associated with object name learning.