Autonomous agents capable of diverse object manipulations should be able to acquire a wide range of manipulation skills with high reusability. Although advances in deep learning have made it increasingly feasible to replicate the dexterity of human teleoperation in robots, generalizing these acquired skills to previously unseen scenarios remains a significant challenge. In this study, we propose a novel algorithm, Gaze-based Bottleneck-aware Robot Manipulation (GazeBot), which enables high reusability of learned motions without sacrificing dexterity or reactivity. By leveraging gaze information and motion bottlenecks—both crucial features for object manipulation—GazeBot achieves high success rates compared with state-of-the-art imitation learning methods, particularly when the object positions and end-effector poses differ from those in the provided demonstrations. Furthermore, the training process of GazeBot is entirely data-driven once a demonstration dataset with gaze data is provided.
How can the whole have a causal effect on its parts? This is considered impossible because, if the supervenient whole is completely determined by its parts, then the whole-to-parts causation would be redundant. However, the exclusion argument did not assume a hierarchy of multiple supervenient functions and the existence of an inter-level negative feedback control mechanism. Here, we propose that this mechanism enables a causal effect from the whole to its parts. Feedback control typically involves two mechanisms: observing and controlling of the feedback error. These mechanisms can be implemented at two different levels of the hierarchy. We assume that the macro-level consists of a set of mathematical functions that supervene on physical neural states. An algebraic structure of these functions describes a macro-level equation that determines the feedback error. This equation is independent of an external cause, introducing new causal power to the micro-level while avoiding overdetermination. Modifying the micro synaptic weights within the neural networks via inter-level negative feedback control is a whole-to-parts causation mechanism. It should be noted that this paper does not intend to take a position on the ontology surrounding downward causation.
Representation learning seeks meaningful sensory representations without supervision and can model aspects of human development. Although many neural networks empirically learn useful features, a principled account of what makes a representation "good" remains elusive. We study unsupervised categorization of transformations between pairs of inputs under algebraic constraints. Classical disentanglement favors mutually independent factors and fails when factors are coupled. Our prior Galois-theoretic approach decomposes a group via normal subgroups by learning a product of two transformations with one factor constrained to a normal subgroup, covering both commutative and non-commutative cases. That method, however, relied on auxiliary assumptions (e.g., motion and isometry restrictions) not required by decomposition theory, and ablations did not separate theory-based from auxiliary effects. We propose parameter division for a single transformation: we split its parameter into components, impose homomorphism constraints mapping the full transformation to one component, and identify the normal subgroup as the set of transformations when that component is fixed to the identity. This formulation drops the previous auxiliary assumptions and applies more broadly. We evaluate on image pairs involving rotation, translation, and scale; ablations show that group-decomposition constraints drive appropriate categorization.
Imitation learning has demonstrated impressive results in robotic manipulation but fails under out-of-distribution (OOD) states. This limitation is particularly critical in Deformable Object Manipulation (DOM), where the near-infinite possible configurations render comprehensive data collection infeasible. Although several methods address OOD states, they typically require exhaustive data or highly precise perception. Such requirements are often impractical for DOM owing to its inherent complexities, including self-occlusion. To address the OOD problem in DOM, we propose a novel framework, Exploration-assisted Bottleneck Transition for Deformable Object Manipulation (ExBot), which addresses the OOD challenge through two key advantages. First, we introduce bottleneck states, standardized configurations that serve as starting points for task execution. This enables the reconceptualization of OOD challenges as the problem of transitioning diverse initial states to these bottleneck states, significantly reducing demonstration requirements. Second, to account for imperfect perception, we partition the OOD state space based on recognizability and employ dual action primitives. This approach enables ExBot to manipulate even unrecognizable states without requiring accurate perception. By concentrating demonstrations around bottleneck states and leveraging exploration to alter perceptual conditions, ExBot achieves both data efficiency and robustness to severe OOD scenarios. Real-world experiments on rope and cloth manipulation demonstrate successful task completion from diverse OOD states, including severe self-occlusions.
A central challenge in consciousness research is the lack of agreement on what a theory of consciousness should explain, which makes it difficult to compare existing theories. We propose a framework for organizing explanatory targets of theories based on a minimal set of seven questions designed to be theoretically neutral, causally and functionally relevant, and applicable across different systems. We focus particularly on the role of causation based on the argument that causal relations cannot be fully specified within standard physical descriptions alone. Introducing an asymmetric causal structure allows internal mechanisms to be represented explicitly and helps distinguish between variable- and structure-level causation. As an example, we apply the proposed framework to analyzing the Dual-Laws Model. The aim of the framework is not to propose a definitive theory but to provide a common basis for analyzing and developing theories of consciousness.
The mirror self-recognition test evaluates whether a subject touches a mark on its own body that is visible only in a mirror, and is widely used as an indicator of self-awareness. In this study, we present a computational model in which this behavior emerges spontaneously through a single mechanism, the self-prior, without any external reward. The self-prior, implemented with a Transformer, learns the density of familiar multisensory experiences; when a novel mark appears, the discrepancy from this learned distribution drives mark-directed behavior through active inference. A simulated infant, relying solely on vision and proprioception without tactile input, discovered a sticker placed on its own face in the mirror and removed it in approximately 70
Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social dilemma situations. There, individual interests are misaligned with the common good and individual rationality leads to suboptimal group outcomes. In contrast, humans are able to achieve cooperation with one another in such situations. A common explanation for such cooperative behavior is that individuals have social preferences. In order to achieve cooperation in MARL, we design a new utility function integrating altruistic preferences (incentive for other's reward) and fairness preferences (incentive for equality) from social psychology and behavioral economics, namely, Altruistic and Fairness Preference (AFP), a reward-sharing mechanism which converts one's own and other's rewards to incentives for cooperative behavior. We performed comparative experiments with standard RL and inequity aversion agents in two challenging sequential social dilemma games, and showed that AFP agents successfully achieved mutual cooperation with more collective rewards and higher equity than the baselines. To further understand the progression of AFP during training, we subsequently explore the effects of altruistic preferences and fairness preferences on agents' behavior. The results suggest that altruistic preferences encourage agents to contribute to the public goods, and fairness preferences induce mutual behavior between agents.
The escalating energy demands for high-performance machine learning have sparked growing interest in unconventional computing paradigms rooted in physical systems. At the core of this emerging direction is the leveraging of physical properties to design computing frameworks that integrate strong learning capabilities, efficient training strategies, and practical implementability. Here, we present a simple, efficient, and scalable framework aimed at achieving this, and validate it on an optoelectronic platform. This framework features our proposed training mechanism—state-skipping direct feedback alignment—which eliminates access to intermediate states of the system and thereby significantly simplifies the error backpropagation process, enhancing both training efficiency and practical feasibility. Compared to conventional deep neural networks, our approach substantially reduces computational costs and training parameter count while achieving comparable performance. Furthermore, we integrate our scheme into modern architectures, attaining improved performance with an approximately 40% reduction in computational resources. Notably, our approach challenges the scaling laws of conventional counterparts, underscoring its strong scalability and practical promise.
This study proposes a foundational method for objectively classifying whether mothers can easily perceive fetal movements, a crucial indicator in perinatal care. Traditional fetal movement assessments rely on subjective maternal reports or wearable sensor-based methods susceptible to external noise.This paper applies the previously established Multi-Resolution Feature (MRF) method to independently quantify both fetal and maternal deformation from simulated fetal movement videos. We constructed a novel fetal-maternal deformation map using these deformation values as axes. The results demonstrate that this map can visually represent the relationship between fetal and maternal deformation. Furthermore, exploratory analysis of the data distribution on the map suggested the existence of distinct fetal movement patterns: movements accompanied by significant maternal abdominal wall deformation (likely perceived by the mother) and movements where the fetus is active but less transmission occurs to the mother (likely difficult for the mother to perceive). This method is expected to be a stepping stone towards establishing a new approach for more detailed assessment of fetal health and, consequently, improving the quality of perinatal medical care.
Mental events are considered to supervene on physical events. A supervenient event does not change without a corresponding change in the underlying subvenient physical events. Since wholes and their parts exhibit the same supervenience-subvenience relations, inter-level causation has been expected to serve as a model for mental causation. We proposed an inter-level causation mechanism to construct a model of consciousness and an agent's self-determination. However, a significant gap exists between this mechanism and cognitive functions. Here, we demonstrate how to integrate the inter-level causation mechanism with the widely known dual-process theories. We assume that the supervenience level is composed of multiple supervenient functions (i.e., neural networks), and we argue that inter-level causation can be achieved by controlling the feedback error defined through changing algebraic expressions combining these functions. Using inter-level causation allows for a dual laws model in which each level possesses its own distinct dynamics. In this framework, the feedback error is determined independently by two processes: (1) the selection of equations combining supervenient functions, and (2) the negative feedback error reduction to satisfy the equations through adjustments of neurons and synapses. We interpret these two independent feedback controls as Type 1 and Type 2 processes in the dual process theories. As a result, theories of consciousness, agency, and dual process theory are unified into a single framework, and the characteristic features of Type 1 and Type 2 processes are naturally derived.
Defining agency is an extremely important challenge for cognitive science and artificial intelligence. Physics generally describes mechanical happenings, but there remains an unbridgeable gap between them and the acts of agents. To discuss the morality and responsibility of agents, it is necessary to model acts; whether such responsible acts can be fully explained by physical determinism has been debated. Although we have already proposed a physical "agent determinism" model that appears to go beyond mere mechanical happenings, we have not yet established a strict mathematical formalism to eliminate ambiguity. Here, we explain why a physical system can follow coarse-graining agent-level determination without violating physical laws by formulating supervenient causation. Generally, supervenience including coarse graining does not change without a change in its lower base; therefore, a single supervenience alone cannot define supervenient causation. We define supervenient causation as the causal efficacy from the supervenience level to its lower base level. Although an algebraic expression composed of the multiple supervenient functions does supervenes on the base, a sequence of indices that determines the algebraic expression does not supervene on the base. Therefore, the sequence can possess unique dynamical laws that are independent of the lower base level. This independent dynamics creates the possibility for temporally preceding changes at the supervenience level to cause changes at the lower base level. Such a dual-laws system is considered useful for modeling self-determining agents such as humans.
While current deep learning models achieve high performance by learning statistical correlations from vast datasets,which stands in stark contrast to human learning. They lack the flexibility of humans-particularly preverbal infants-to autonomously acquire the underlying structure of the world from limited experience and adapt to novel situations. In this study, we propose an unsupervised representation learning method based on a hierarchical relationship in group operations, rather than statistical independence, aiming to build a computational model of the cognitive development of infants. The proposed model features an integrated architecture that simultaneously performs object segmentation and the extraction of motion laws from dynamic image sequences. By introducing the Homomorphism from algebra as a structural constraint within a neural network, the model structurally separates pixel-level changes into meaningful, decomposed transformation components, such as translation and deformation. Using interaction scenes (chasing and evading tasks) based on developmental science findings, we experimentally demonstrate that the model can segment multiple objects into individual slots without any ground-truth labels. Furthermore, we confirmed that relative movements between objects, such as approaching or receding, are accurately mapped and structured into a one-dimensional additive latent space. These results suggest that by introducing algebraic geometric constraints rather than relying solely on statistical correlation learning, physically interpretable "disentangled representations" can be acquired. This study contributes to the understanding of the process by which infants internalize environmental laws as structures and provides a new perspective for constructing artificial systems with developmental intelligence.
Motion retargeting from humans to human-like artificial agents is becoming increasingly important as humanoid robots grow more capable. However, most existing approaches focus only on reproducing kinematics and ignore the rich sensorimotor experience associated with human movement. In this work, we present a framework for simulating the multimodal sensorimotor experiences of infants using physical and virtual humanoids. From a single video, our method reconstructs the infant's body configuration by extracting its skeletal structure and estimating the full 3D pose from each frame. Then we map the reconstructed motion onto several developmental platforms: the physical iCub robot and the virtual simulators pyCub, EMFANT and MIMo. Replaying the retargeted motions on these embodiments produces simulated multisensory streams including proprioception (joints and muscles), touch, and vision. For the best-matching embodiment, the retargeting achieves sub-centimeter accuracy and enables a rich multimodal analysis of infant development as well as enhanced automated annotation of behaviors. This framework provides a unique window into the infant's sensorimotor experience, offering new tools for robotics, developmental science, and early detection of neurodevelopmental disorders. The code is available at https://github.com/ctu-vras/motion-retargeting/.
A self-determining system is defined as one in which causes originating within the system influence the system itself. This definition raises the question of how to specify system boundaries. Although the concept of "closure" is commonly used for this purpose, defining boundaries solely in terms of causal relations introduce challenges, such as how to handle external causes and circular causality. To address this issue, we introduce two types of asymmetric relations: causal and constitutive. We propose that system boundaries can be defined as closures of loops formed by these relations, referred to as causal-constitutive loops. By constraining constitutive relations, the resulting system necessarily includes internal causes and thereby satisfies self-determination. Furthermore, to prevent reduction to supervenience, constitutive relations must involve at least two independent variables. This minimal requirement leads to two interdependent loops, which implies a dual-process organization.
Understanding and predicting how mechanical systems respond to environmental variability is essential for advancing next-generation robotic systems with physical intelligence. In this study, we investigated the use of echo state networks (ESNs), a representative class of reservoir computing (RC) models, to predict the bifurcation structures of real-world mechanical systems from limited observations. We examined two representative cases: a simulated passive dynamic walking (PDW) robot with hybrid continuous-discrete dynamics and a real-world soft pneumatic artificial muscle (PAM) actuator whose electrical resistance undergoes complex changes under varying loads. To address the challenges posed by the PDW's hybrid dynamics, we proposed a hybrid ESN (HESN) model that integrates a knowledge-based touchdown detection mechanism with an ESN module. The HESN successfully reproduced the route-to-chaos bifurcation structure of the PDW, captured its multi-attractor dynamics, and demonstrated robustness against imperfect domain knowledge. For the PAM, where no reliable physical model is available, a purely data-driven ESN accurately predicted resistance bifurcations across changing environmental conditions. These results highlight the potential of RC models as flexible digital twins for mechanical systems, enabling parameter-aware modeling of bifurcations with limited training data and supporting the design of robust, adaptive robots capable of operating in complex environments.
Accurately predicting individual aesthetic evaluation for images is a fundamental challenge for AI. Various deep learning (DL)-based models have been proposed for this task, training on image evaluation data to extract objective low-level features. However, aesthetic preferences are inherently subjective and individual-dependent. Accurate prediction thus requires the extraction of high-level semantic features of images and the active collection of preference information from the target individual. To address this issue, we focus on the utility of Large Language Models (LLMs) pretrained on vast amounts of textual data, and develop an integrated DL-LLM system. The system actively elicits aesthetic preferences through LLM-based semi-structured interviews and predicts aesthetic evaluation by leveraging both low-level and high-level features. In our experiments, we compare the proposed system against conventional systems, human predictors, and the target individual's own re-evaluations after a certain time interval. Our results show that the proposed system outperforms all of them, with particularly strong performance on highly-rated images. Moreover, the prediction error of the proposed system is smaller than within-person variability, while human predictors show the largest error, likely due to the influence of their own aesthetic values. These results suggest that AI may be better positioned than others or one's future self to capture individual aesthetic preferences at a given point. This opens a new question of whether AI could serve as a deeper interpreter of human aesthetic sensibility than humans themselves.
Optimizing vision models purely for classification accuracy can impose an alignment tax, degrading human-like scanpaths and limiting interpretability. We introduce EVA, a neuroscience-inspired hard-attention mechanistic testbed that makes the performance-human-likeness trade-off explicit and adjustable. EVA samples a small number of sequential glimpses using a minimal fovea-periphery representation with CNN-based feature extractor and integrates variance control and adaptive gating to stabilize and regulate attention dynamics. EVA is trained with the standard classification objective without gaze supervision. On CIFAR-10 with dense human gaze annotations, EVA improves scanpath alignment under established metrics such as DTW, NSS, while maintaining competitive accuracy. Ablations show that CNN-based feature extraction drives accuracy but suppresses human-likeness, whereas variance control and gating restore human-aligned trajectories with minimal performance loss. We further validate EVA's scalability on ImageNet-100 and evaluate scanpath alignment on COCO-Search18 without COCO-Search18 gaze supervision or finetuning, where EVA yields human-like scanpaths on natural scenes without additional training. Overall, EVA provides a principled framework for trustworthy, human-interpretable active vision.
What exactly is the meaning of physical causal closure, a concept frequently discussed in the philosophy of mind? Jaegwon Kim explicitly adopts a conception of causation according to which physical causation is effectively identified with deterministic physical lawfulness, and on this basis equates physical determinism with physical causal closure. While this conception is internally coherent, it differs from the currently dominant theories of causation, which emphasize asymmetry between cause and effect grounded in manipulability and intervention widely employed in contemporary scientific practice. Physics and the theory of causation serve different descriptive purposes, and in this study we refer to them respectively as the Physical Stance and the Causal Stance. Within this framework, physical determinism is a notion that belongs to the Physical Stance, whereas physical causal closure is a notion defined only within the Causal Stance; consequently, the two should not be equated. Since causation is not explicitly defined within the language of physics, physical causal closure is not definable within the Physical Stance alone. By distinguishing between these two stances, this study reconstructs Davidson's Anomalous Monism as a materialist position that consistently acknowledges mental causation without contradicting physical determinism, and examines its relation to the Dual-Laws Model we propose. We further argue that, for the development of scientific theories of mind and consciousness, it is necessary to construct a linguistic framework within which physical causal closure does not hold in the Causal Stance, while physical determinism remains intact in the Physical Stance.
In general, objects can be distinguished on the basis of their features, such as color or shape. In particular, it is assumed that similarity judgments about such features can be processed independently in different metric spaces. However, the unsupervised categorization mechanism of metric spaces corresponding to object features remains unknown. Here, we show that the artificial neural network system can autonomously categorize metric spaces through representation learning to satisfy the algebraic independence between neural networks, and project sensory information onto multiple high-dimensional metric spaces to independently evaluate the differences and similarities between features. Conventional methods often constrain the axes of the latent space to be mutually independent or orthogonal. However, the independent axes are not suitable for categorizing metric spaces. High-dimensional metric spaces that are independent of each other are not uniquely determined by the mutually independent axes, because any combination of independent axes can form mutually independent spaces. In other words, the mutually independent axes cannot be used to naturally categorize different feature spaces, such as color space and shape space. Therefore, constraining the axes to be mutually independent makes it difficult to categorize high-dimensional metric spaces. To overcome this problem, we developed a method to constrain only the spaces to be mutually independent and not the composed axes to be independent. Our theory provides general conditions for the unsupervised categorization of independent metric spaces, thus advancing the mathematical theory of functional differentiation of neural networks.
Luc Berthouze合作论文数Department of Informatics, School of Engineering and Informatics, University of Sussex;Department of Developmental Neurosciences, Institute of Child Health, University College London7