Higher-order Thought (HOT) theories of consciousness typically claim that conscious experience is a matter of a mental state's being represented by a HOT which is itself normally an unconscious state. Self-representational (SR) theorists have objected to HOT theory on these grounds, citing our constant awareness of our own consciousness, which they claim cannot be explained by appeal to an unconscious HOT (though it can be explained by appeal to self-representing mental states). I consider two interpretations of this objection, one claiming that any content of which we are aware in experience must occur as part of a conscious mental state, and a stronger interpretation according to which the explanation of state consciousness cannot ultimately rest on appeal to unconscious representations. I then argue that in its stronger form, the objection applies equally to SR (entailing a regress of representational properties rather than of mental states), while in its weaker form, it poses no threat to HOT theory.
Recent advances in theoretical biology suggest that key definitions of basal cognition and sentient behavior may arise as emergent properties of in vitro cell cultures and neuronal networks. Such neuronal networks reorganize activity to demonstrate structured behaviors when embodied in structured information landscapes. In this article, we characterize this kind of self-organization through the lens of the free energy principle, that is, as self-evidencing. We do this by first discussing the definitions of reactive and sentient behavior in the setting of active inference, which describes the behavior of agents that model the consequences of their actions. We then introduce a formal account of intentional behavior that describes agents as driven by a preferred end point or goal in latent state-spaces. We then investigate these forms of (reactive, sentient, and intentional) behavior using simulations. First, we simulate the in vitro experiments, in which neuronal cultures modulated activity to improve gameplay in a simplified version of Pong by implementing nested, free energy minimizing processes. The simulations are then used to deconstruct the ensuing predictive behavior, leading to the distinction between merely reactive, sentient, and intentional behavior with the latter formalized in terms of inductive inference. This distinction is further studied using simple machine learning benchmarks (navigation in a grid world and the Tower of Hanoi problem) that show how quickly and efficiently adaptive behavior emerges under an inductive form of active inference.
This paper examines the constraints that the free-energy principle (FEP) places on possible model of consciousness, particularly models of attentional control and imaginative experiences, including episodic memory and planning. We first rehearse the classical and quantum formulations of the FEP, focusing on their application to multi-component systems, in which only some components interact directly with the external environment. In particular, we discuss the role of internal boundaries that have the structure of Markov blankets, and hence function as classical information channels between components. We then show how this formal structure supports models of attentional control and imaginative experience, with a focus on (i) how imaginative experience can employ the spatio-temporal and object-recognition reference frames employed in ordinary, non-imaginative experience and (ii) how imaginative experience can be internally generated but still surprising. We conclude by discussing the implementation, phenomenology, and phylogeny of imaginative experience, and the implications of the large state and trait variability of imaginative experience in humans.
As artificial agents enter open-ended physical environments – eldercare, disaster response, and space missions – they must persist under uncertainty while providing reliable care. Yet current systems struggle to generalize across distribution shifts and lack intrinsic motivation to preserve the well-being of others. Vulnerability and mortality are often seen as constraints to be avoided, yet organisms survive and provide care in an open-ended world with relative ease and efficiency. We argue that generalization and care arise from conditions of physical embodiment: being-in-the-world (the agent is a part of the environment) and being-towards-death (unless counteracted, the agent drifts toward terminal states). These conditions necessitate a homeostatic drive to maintain oneself and maximize the future capacity to continue doing so. Fulfilling this drive over long time horizons in multi-agent environments necessitates robust causal modeling of self and others' embodiment and jointly achievable future states. Because embodied agents are part of the environment, with the self delimited by reliable control, empowering others can expand self-boundaries, enabling other-regard. This provides a path from embodiment toward generalization and care based in shared constraints. We outline a reinforcement-learning framework for examining these questions. Homeostatic mortal agents continually learning in open-ended environments may offer efficient robustness and trustworthy alignment.
We provide a general account of propositional attitudes from the point of view of the predictiveprocessing (PP) framework. We first specify functional roles for conative and cognitive attitudetypes, in terms of the distinction between agent-driven and stimulus-driven control over internalstates. This accounts for the motivational role of conative attitudes as well as the rationalityconstraints characteristic of cognitive attitudes such as belief and perception. We then examinebelief and desire in more detail, and develop a broader taxonomy of attitude types, castinginstances of counterfactual exploration, such as supposition and imagination, as second-orderpropositional attitudes in which a cognitive attitude is the object of a conative one. Finally, weaddress worries about the suitability of PP’s representations—probabilistic graphical models—for implementing the attitudes.
This paper concerns structure learning or discovery of discrete generative models. It focuses on Bayesian model selection and the assimilation of training data or content, with a special emphasis on the order in which data are ingested. A key move - in the ensuing schemes - is to place priors on the selection of models, based upon expected free energy. In this setting, expected free energy reduces to a constrained mutual information, where the constraints inherit from priors over outcomes (i.e., preferred outcomes). The resulting scheme is first used to perform image classification on the MNIST dataset to illustrate the basic idea, and then tested on a more challenging problem of discovering models with dynamics, using a simple sprite-based visual disentanglement paradigm and the Tower of Hanoi (cf., blocks world) problem. In these examples, generative models are constructed autodidactically to recover (i.e., disentangle) the factorial structure of latent states - and their characteristic paths or dynamics.
Legal autonomy - the lawful activity of artificial intelligence agents - can be achieved in one of two ways. It can be achieved either by imposing constraints on AI actors such as developers, deployers and users, and on AI resources such as data, or by imposing constraints on the range and scope of the impact that AI agents can have on the environment. The latter approach involves encoding extant rules concerning AI driven devices into the software of AI agents controlling those devices (e.g., encoding rules about limitations on zones of operations into the agent software of an autonomous drone device). This is a challenge since the effectivity of such an approach requires a method of extracting, loading, transforming and computing legal information that would be both explainable and legally interoperable, and that would enable AI agents to reason about the law. In this paper, we sketch a proof of principle for such a method using large language models (LLMs), expert legal systems known as legal decision paths, and Bayesian networks. We then show how the proposed method could be applied to extant regulation in matters of autonomous cars, such as the California Vehicle Code.
Previous active inference accounts of emotion translate fluctuations in free energy to a sense of emotion, sometimes focusing exclusively on valence. However, in affective science, emotions are often represented as multidimensional. In this paper, we adapt a Circumplex Model of emotion to the Active Inference framework by demonstrating a mapping of free energy into valence and arousal, relating valence to utility less expected utility and arousal to the entropy of posterior beliefs. Under this formulation, we simulate artificial agents engaged in a search task and assign emotional states to them. We show that experimental manipulation of priors and object presence results in commonsense variability in these assignments.
This paper aims to assess whether the recently proposed "inner screen model" of consciousness that follows from the free-energy principle (FEP) can be regarded as a minimal unifying model (MUM) of consciousness, thereby providing a common foundational model for consciousness studies, and integrating approaches to consciousness based on the FEP. We first present the inner screen model, which follows from applying the quantum information theoretic version of the FEP to the known sparse (nested and hierarchical) neuroanatomy of the brain. We then review models of consciousness that are premised on the FEP. Specifically, we review Bayesian versions of the global workspace and attention schema theories, theories premised on world-models and self-models, and models formalizing the computational structure and properties of time-consciousness. We then discuss how extant FEP-theoretic models of consciousness can be situated with respect to the candidate MUM.
Although the latent spaces learned by distinct neural networks are not generally directly comparable, recent work in machine learning has shown that it is possible to use the similarities and differences among latent space vectors to derive "relative representations" with comparable representational power to their "absolute" counterparts, and which are nearly identical across models trained on similar data distributions. Apart from their intrinsic interest in revealing the underlying structure of learned latent spaces, relative representations are useful to compare representations across networks as a generic proxy for convergence, and for zero-shot model stitching. In this work we examine an extension of relative representations to discrete state-space models, using Clone-Structured Cognitive Graphs (CSCGs) for 2D spatial localization and navigation as a test case. Our work shows that the probability vectors computed during message passing can be used to define relative representations on CSCGs, enabling effective communication across agents trained using different random initializations and training sequences, and on only partially similar spaces. We introduce a technique for zero-shot model stitching that can be applied post hoc, without the need for using relative representations during training. This exploratory work is intended as a proof-of-concept for the application of relative representations to the study of cognitive maps in neuroscience and AI.
This paper presents a model of consciousness that follows directly from the free- energy principle (FEP). We first rehearse the classical and quantum formulations of the FEP. In particular, we consider the “inner screen hypothesis” that follows from the quantum information theoretic version of the FEP. We then review applications of the FEP to the known sparse (nested and hierarchical) neuro-anatomy of the brain. We focus on the holographic structure of the brain, and how this structure supports (overt and covert) action.
We develop an approach to policy selection in active inference that allows us to efficiently search large policy spaces by mapping each policy to its embedding in a vector space. We sample the expected free energy of representative points in the space, then perform a more thorough policy search around the most promising point in this initial sample. We consider various approaches to creating the policy embedding space, and propose using k-means clustering to select representative points. We apply our technique to a goal-oriented graph-traversal problem, for which naive policy selection is intractable for even moderately large graphs.
Capsule networks are a neural network architecture specialized for visual scene recognition. Features and pose information are extracted from a scene and then dynamically routed through a hierarchy of vector-valued nodes called 'capsules' to create an implicit scene graph, with the ultimate aim of learning vision directly as inverse graphics. Despite these intuitions, however, capsule networks are not formulated as explicit probabilistic generative models; moreover, the routing algorithms typically used are ad-hoc and primarily motivated by algorithmic intuition. In this paper, we derive an alternative capsule routing algorithm utilizing iterative inference under sparsity constraints. We then introduce an explicit probabilistic generative model for capsule networks based on the self-attention operation in transformer networks and show how it is related to a variant of predictive coding networks using Von-Mises-Fisher (VMF) circular Gaussian distributions.
In this article, we aim to conceptualize and formalize the construct of resilience using the tools of active inference, a new physics-based modeling approach apt for the description and analysis of complex adaptive systems. We intend this as a first step toward a computational model of resilient systems. We begin by offering a conceptual analysis of resilience, to clarify its meaning, as established in the literature. We examine an orthogonal, threefold distinction between meanings of the word "resilience": (i) inertia, or the ability to resist change (ii) elasticity, or the ability to bounce back from a perturbation, and (iii) plasticity, or the ability to flexibly expand the repertoire of adaptive states. We then situate all three senses of resilience within active inference. We map resilience as inertia onto high precision beliefs, resilience as elasticity onto relaxation back to characteristic (i.e., attracting) states, and resilience as plasticity onto functional redundancy and structural degeneracy.
There is a growing body of evidence suggesting that the neural processes underlying perception, learning, and decision-making approximate Bayesian inference. Yet, humans perform poorly when asked to solve explicit probabilistic reasoning problems. In response, some have argued that certain brain processes are Bayesian while others are not; others have argued that reasoning errors can be explained by either inaccurate generative models or limitations of approximation algorithms. In this paper, we offer a complementary perspective by considering how a Bayesian brain would implement conscious reasoning processes more generally. These considerations require making two distinctions, each of which highlights a fundamental reason why Bayesian brains should not be expected to perform well at explicit inference. The first distinction is between inferring probability distributions over hidden states and representing probabilities as hidden states. The former assumes that the brain’s dynamics instantiate a form of approximate Bayesian inference, premised on a model of how observations are generated by hidden states of the world. In contrast, the latter assumes the brain represents probabilities themselves as hidden states – namely, hypotheses about the correct answers to explicit reasoning problems. In this latter case, correctly inferring the most likely probability to report would implausibly require the brain to possess a generative model encoding Bayes’ theorem itself. The second distinction is between inference and mental action. In addition to state inference, consciously solving Bayes’ theorem requires the selection of a particular sequence of goal-directed cognitive actions (e.g., mental multiplication and addition, followed by division). While Bayesian brains infer probability distributions over action sequences, the possible sequences themselves often need to be learned. These considerations show that, regardless of the specific generative model in question or approximation algorithm employed, and even if all brain processes were Bayesian, an innate proficiency at solving explicit probabilistic reasoning problems should not be expected.
We challenge the authors' view that Markov blankets are illicitly reified when used to describe organismic boundaries. We do this both on general methodological grounds, where we appeal to a form of structural realism derived from Bayesian cognitive science to dissolve the problem, and by rebutting specific arguments in the target article.
This white paper lays out a vision of research and development in the field of artificial intelligence for the next decade (and beyond). Its denouement is a cyber-physical ecosystem of natural and synthetic sense-making, in which humans are integral participants—what we call “shared intelligence.” This vision is premised on active inference, a formulation of adaptive behavior that can be read as a physics of intelligence, and which inherits from the physics of self-organization. In this context, we understand intelligence as the capacity to accumulate evidence for a generative model of one’s sensed world—also known as self-evidencing. Formally, this corresponds to maximizing (Bayesian) model evidence, via belief updating over several scales, that is, inference, learning, and model selection. Operationally, this self-evidencing can be realized via (variational) message passing or belief propagation on a factor graph. Crucially, active inference foregrounds an existential imperative of intelligent systems; namely, curiosity or the resolution of uncertainty. This same imperative underwrites belief sharing in ensembles of agents, in which certain aspects (i.e., factors) of each agent’s generative world model provide a common ground or frame of reference. Active inference plays a foundational role in this ecology of belief sharing—leading to a formal account of collective intelligence that rests on shared narratives and goals. We also consider the kinds of communication protocols that must be developed to enable such an ecosystem of intelligences and motivate the development of a shared hyper-spatial modeling language and transaction protocol, as a first—and key—step towards such an ecology.
Active inference offers a unified theory of perception, learning, and decision-making at computational and neural levels of description. In this article, we address the worry that active inference may be in tension with folk psychology because it does not include terms for desires (or other conative constructs) at the mathematical level of description. To resolve this concern, we first provide a brief review of the historical progression from predictive coding to active inference, enabling us to distinguish between active inference formulations of motor control (which need not have desires under folk psychology) and active inference formulations of decision processes (which do have desires within folk psychology). We then show that, despite a superficial tension when viewed at the mathematical level, the active inference formalism contains terms that are readily identifiable as encoding both the objects of desire and the strength of desire at the psychological level. We demonstrate this with simple simulations of an active inference agent motivated to leave a dark room for different reasons. Despite their consistency, we further show how active inference may increase the granularity of folk-psychological descriptions by highlighting distinctions between drives to seek information vs. reward – and how it may also offer more precise, quantitative folk-psychological predictions. Finally, we consider how the implicitly conative components of active inference may have partial analogues (i.e., “as if” desires) in other systems describable by the broader free energy principle to which it conforms.
An approach to implementing variational Bayesian inference in biological systems is considered, under which the thermodynamic free energy of a system directly encodes its variational free energy. In the case of the brain, this assumption places constraints on the neuronal encoding of generative and recognition densities, in particular requiring a stochastic population code. The resulting relationship between thermodynamic and variational free energies is prefigured in mind–brain identity theses in philosophy and in the Gestalt hypothesis of psychophysical isomorphism.