We present Hierarchical Residual Networks (HiResNets), deep convolutional neural networks with long-range residual connections between layers at different hierarchical levels. HiResNets draw inspiration on the organization of the mammalian brain by replicating the direct connections from subcortical areas to the entire cortical hierarchy. We show that the inclusion of hierarchical residuals in several architectures, including ResNets, results in a boost in accuracy and faster learning. A detailed analysis of our models reveals that they perform hierarchical compositionality by learning feature maps relative to the compressed representations provided by the skip connections.
Adapting movements to rapidly changing conditions is fundamental for interacting with our dynamic environment. This adaptability relies on internal models that predict and evaluate sensory outcomes to adjust motor commands. Even infants anticipate object properties for efficient grasping, suggesting the use of internal models. However, how internal models are adapted in early childhood remains largely unexplored. This study investigated a naturalistic force adaptation task in 1.5-, 3-year-olds, and young adults. Participants opened a drawer with temporarily increased resistance, creating sensory prediction errors between predicted and actual drawer dynamics. After perturbation, all age groups showed lower peak speed, longer movement time, and more movement units with trial-wise changes analyzed as adaptation process. Results revealed no age differences in adapting peak speed and movement units, but 1.5- and 3-year-olds exhibited higher trial-to-trial variability and were slower in adapting their movement time, although they also adapted their movement time more strongly. Upon removal of perturbation, we found significant aftereffects across all age groups, indicating effective internal model adaptation. These results suggest that even 1.5-year-olds form internal models of force parameters and adapt them to reduce sensory prediction errors, possibly through more exploration and with more variable movement dynamics compared to adults.
Objectives: This study aimed to assess the effects of a virtual Mindful Self-compassion (MSC) intervention on mindfulness, self-compassion, empathy, stress, and well-being in Uruguayan primary school teachers. Methods: A quasi-experimental, longitudinal study was conducted with an active control intervention (Kundalini Yoga, KY). Uruguayan volunteer female teachers were randomly assigned to MSC or KY 9-weeks virtual training and completed self-reports and an empathy for pain task (EPT) at pre-, post-training, and follow-up (3 months). Results: After MSC training, mindfulness (ES: observing= -0.836; non-reactivity= -0.476; total mindfulness= -0.655), self-compassion (ES: self-kindness= 0.745; common humanity= -0.588; mindfulness= -0.487) and self-judgment (ES= -0.463) significantly (p<0.05) increased. Furthermore, perspective-taking increased (ES= -0.505) and personal distress decreased (ES= -0.587), while stress decreased (ES= -0.450) and well-being increased (ES= -0.612) after this training. At follow-up, observing (ES= -0.675) and total mindfulness (ES= -0.757) remained elevated and non-judging increased (ES= -0.667); self-compassion remained elevated (ES= -0.778) and personal distress remained decreased ( ES= -0.857). After MSC training, EPT intentionality comprehension accuracy significantly increased (SE= -0.588). After training, personal distress was higher in KY than MSC (ES= -0.344), while at follow-up observing (ES= -0.454) and total mindfulness (ES =-0.415) were higher in MSC. No differences between groups were found for the EPT. Conclusions: Virtual MSC training cultivated mindfulness and self-compassion associated with an increase in well-being and empathy, and a reduction of stress in Uruguayan primary school teachers.
Adapting movements to rapidly changing conditions is fundamental for interacting with our dynamic environment. This adaptability relies on internal models that predict and evaluate sensory outcomes to adjust motor commands. Even infants anticipate object properties for efficient grasping, suggesting the use of internal models. However, how internal models are adapted in early childhood remains largely unexplored. This study investigated a naturalistic force adaptation task in 1.5-, 3-year-olds, and young adults. Participants opened a drawer with temporarily increased resistance, creating sensory prediction errors between predicted and actual drawer dynamics. After perturbation, all age groups showed lower peak speed, longer movement time, and more movement units with trial-wise changes analyzed as adaptation process. Results revealed no age differences in adapting peak speed and movement units, but 1.5- and 3-year-olds exhibited higher trial-to-trial variability and were slower in adapting their movement time, although they also adapted their movement time more strongly. Upon removal of perturbation, we found significant aftereffects across all age groups, indicating effective internal model adaptation. These results suggest that even 1.5-year-olds form internal models of force parameters and adapt them to reduce sensory prediction errors, possibly through more exploration and with more variable movement dynamics compared to adults.
Binocular saccades in the three-dimensional world entail changes in the direction and depth of gaze. Newborns must autonomously learn to control these two types of eye movements concurrently. Here, we put forward a computational model of the joint development of conjugate saccades and vergence, resulting in the self-calibration of disjunctive saccades. This work builds on principles of information theory and efficient coding. We propose that conjugate saccades and vergence are learned together with the shared motivation of maximizing the mutual information between binocular visual inputs and their neural representations. We train and evaluate our model using an infant embodiment in a simulated playroom. Our results are compatible with experimental evidence and show that vergence predicts the binocular disparity at the target of a saccade.
Color constancy (CC) describes the ability of the visual system to perceive an object as having a relatively constant color despite changes in lighting conditions. While CC and its limitations have been carefully characterized in humans, it is still unclear how the visual system acquires this ability during development. Here, we present a first study showing that CC develops in a neural network trained in a self-supervised manner through an invariance learning objective. During learning, objects are presented under changing illuminations, while the network aims to map subsequent views of the same object onto close-by latent representations. This gives rise to representations that are largely invariant to the illumination conditions, offering a plausible example of how CC could emerge during human cognitive development via a form of self-supervised learning.
Human intelligence and human consciousness emerge gradually during the process of cognitive development. Understanding this development is an essential aspect of understanding the human mind and may facilitate the construction of artificial minds with similar properties. Importantly, human cognitive development relies on embodied interactions with the physical and social environment, which is perceived via complementary sensory modalities. These interactions allow the developing mind to probe the causal structure of the world. This is in stark contrast to common machine learning approaches, e.g., for large language models, which are merely passively “digesting” large amounts of training data, but are not in control of their sensory inputs. However, computational modeling of the kind of self-determined embodied interactions that lead to human intelligence and consciousness is a formidable challenge. Here we present MIMo, an open-source multi-modal infant model for studying early cognitive development through computer simulations. MIMo’s body is modeled after an 18-month-old child with detailed five-fingered hands. MIMo perceives its surroundings via binocular vision, a vestibular system, proprioception, and touch perception through a full-body virtual skin, while two different actuation models allow control of his body.We describe the design and interfaces of MIMo and provide examples illustrating its use. All code is available at https://github.com/trieschlab/MIMo .
During their first months of life, infants learn to coordinate their perceptions and actions across different modalities. For example, eye-hand coordination relies on combining visual and proprioceptive sensory inputs for controlling eye and hand movements. What drives the development and calibration of such coordination? Here, we put forward a multimodal hierarchical extension of the Active Efficient Coding framework to learn a simple form of eye-hand coordination. By learning to actively compress visual and proprioceptive inputs into a combined multimodal representation, our embodied infant model learns to make eye movements to track an object held in its hand. We find that the abstract multimodal representation improves the tracking accuracy, but only if it emerges after the establishment of the single-modality systems. This suggests the existence of a “less-is-more” effect for the development of coordinated multimodal sensorimotor behaviors.
Laboratory data from conflict tasks, e.g. Simon and Eriksen tasks, reveal differences in response time distributions under different experimental conditions. Only recently have evidence accumulation models successfully reproduced these results, in particular the challenging delta plots with negative slopes. They accomplish this with explicit temporal dependencies in their structure or activation functions. In this work, we introduce an alternative approach to the modeling of decision-making in conflict tasks exclusively based on inhibitory dynamics within a dual-route architecture. We consider simultaneous automatic and controlled drift diffusion processes, with the latter inhibiting the former. Our proposed Dual-Route Evidence Accumulation Model (DREAM) achieves equivalent performance to previous works in fitting experimental response time distributions despite having no time-dependent functions. The model can reproduce conditional accuracy functions and delta plots with positive and negative slopes. The implications of these results, including an interpretation of the parameters and potential links to perceptual representations, are discussed. We provide Python code to fit DREAM to experimental data.
How natural and artificial vision systems learn and develop depends on how they sample information from their environment. Humans actively do so through saccadic eye movements. The statistics of saccade amplitudes have been well-characterized in tightly controlled contexts such as viewing images on a computer screen. However, the degree to which such findings generalize to real-world contexts involving moving agents and objects is currently unknown. Here, we first analyze saccade amplitude statistics of both infants and adults during naturalistic free play. We find that these differ significantly from those previously reported for head-fixed picture viewing, with a relatively smaller/greater abundance of medium/large saccades. Next, we present a computational model that explains saccade amplitude statistics based on the foveated nature of vision and the associated space-variant magnification of different portions of the visual field in primary visual cortex. Finally, we demonstrate computationally efficient approximations to this space-variant sampling using a small number of discrete resolution levels.
Context-dependent computation is a relevant characteristic of neural systems, endowing them with the capacity of adaptively modifying behavioral responses and flexibly discriminating between relevant and irrelevant information in a stimulus. This ability is particularly highlighted in solving conflicting tasks. A long-standing problem in computational neuroscience, flexible routing of information, is also closely linked with the ability to perform context-dependent associations. Here we present an extension of a context-dependent associative memory model to achieve context-dependent decision-making in the presence of conflicting and noisy multi-attribute stimuli. In these models, the input vectors are multiplied by context vectors via the Kronecker tensor product. To outfit the model with a noisy dynamic, we embedded the context-dependent associative memory in a leaky competing accumulator model, and, finally, we proved the power of the model in the reproduction of a behavioral experiment with monkeys in a context-dependent conflicting decision-making task. At the end, we discuss the neural feasibility of the tensor product and made the suggestive observation that the capacities of tensor context models are surprisingly in alignment with the more recent experimental findings about functional flexibility at different levels of brain organization.