
Stuart Russell has named “human enfeeblement” as a long-term societal risk of advanced AI. This paper gives a new interpretation of the enfeeblement danger by interpreting it in light of recent work by Brian Christian, Alison Gopnik, and David Spivak on care and caregiving, and connecting it with the ontological design tradition of Terry Winograd and Fernando Flores. The enfeeblement that threatens is not merely a matter of cognitive off-loading and loss of knowledge, but a form of disengagement with the world which I explain in terms of an erosion of our capacities to care. This paper articulates a conception of care as a constellation of cultivable skills that involve us in the world and with each other, while making life worth living. Designing AI to support our capacities to care addresses the worry recently raised by Demis Hassabis and others regarding the questions of “meaning and purpose” that are looming in a post-AGI world.
The defining challenge of AGI research is to build synthetic systems that genuinely instantiate autonomous agency rather than merely producing sophisticated intelligent outputs. While contemporary Large Language Models (LLMs) achieve remarkable reasoning and statistical capacity, they remain structurally fragile and devoid of any intrinsic stake in their own operational survival. This paper addresses that category error with the Agency Spectrum, a multidimensional diagnostic framework that asks not whether a system is an agent but how the competencies constituting agency are organized within it, across diverse substrates. It decouples a system’s capacities along three independent axes rather than scoring it on a single scale. The framework’s theory draws together Active Inference, the Free Energy Principle, and multiscale biology, with a cross-disciplinary synthesis across philosophy, psychology, biology, and computer science. The result is a substrate-agnostic vocabulary: biological and engineered systems can be compared on common terms. We propose a recursive, closed-loop architecture separating the Driver (intrinsic, viability-grounded goal generation), the Engine (generative modeling), and the Interface (the sensorimotor boundary coupling the system to its environment), in which the three components are joined through bidirectional perception–action exchange. Synthesizing Active Inference with principles of multiscale biology, we show why scaling processing throughput cannot generate autonomous, goal-directed behavior. To reach sophisticated agency, architectures must move beyond disembodied computationalism. We identify three necessary architectural conditions—endogenous normativity, recursive substrate-level viability distribution, and structural remapping capacity—and offer four candidate engineering pathways for transitioning from passive information-processing engines to integrated, self-sustaining synthetic agents.
This paper proposes Affective Control under Uncertainty (ACU), an Active Inference architecture in which constitutive viability modulates policy precision. Rather than adding survival or homeostatic state to a reward function, ACU treats internal viability as a control variable that reshapes the urgency, scope, and sharpness of policy competition. An Affective Viability Controller computes a global viability signal V(t) from deviation between current internal state and viable bounds, enabling flexible exploration when viability is secure and adaptive narrowing when viability is threatened. The paper also formalises Self-Model Recruitment Intensity (SMRI), a multiplicative trigger combining policy entropy, variance in expected viability outcomes, and temporal pressure. When thresholded, SMRI recruits a transient self-model with specific consequences: deeper planning, richer interoceptive and environmental integration, computational reallocation, and representation of the agent as a control-relevant object. A minimal survival-gridworld implementation is specified with reward-only and intrinsic-reward baselines. ACU offers a testable route toward uncertainty-aware AGI grounded in constitutive viability rather than externally assigned objectives.
Active Inference has become a leading candidate framework for general intelligence and is increasingly invoked in accounts of consciousness itself. It is often suggested that sufficiently rich inferential organization under the Free Energy Principle will give rise to subjectivity or phenomenal structure. I argue that this step remains theoretically underdetermined: Active Inference provides a principled account of action, perception, and policy selection, but does not by itself specify the structural condition under which a world is given as meaningful for an agent. Drawing on phenomenological analyses of intentional quality, I propose perspective as a minimal structural condition for subjectivity: a slowly evolving, history-dependent orientation that shapes inference and perception without functioning as a variable within the standard inferential economy. I summarize a minimal architecture in which this condition is probed empirically through targeted ablations, and close with implications for AGI, AI alignment, and the ethics of intentionally instantiating perspectival structure.
We present a theory of self-representing cognitive systems grounded in ([0,∞ ],+) -enriched category theory and the Yoneda lemma. The central object is a self-representing ([0,∞ ],+) -enriched category 𝒞 —a Lawvere metric space whose objects are complete epistemic architectures, whose hom-values record directed informational upgrade costs, and which is separated, closed under internal homs, and bilaterally Cauchy complete—together with a contractive cognitive endofunctor F:𝒞→𝒞 modelling iterative self-improvement. We establish eight results in a single logical arc. The Horizon Theorem shows that the Yoneda embedding φ (A)=𝒞(-,A) is never essentially surjective: 𝒞 sits strictly inside its own free Cauchy completion 𝒫(𝒞) , with the non-representable presheaves forming a topologically dense family, proved via a reflexivity argument. The Lawvere–Banach Attractor Theorem shows that every contractive endofunctor on a bilaterally complete, separated ([0,∞ ],+) -enriched category converges to a unique fixed point Ω at a geometric rate. The Boundary Derivation Theorem shows that Ω is the minimal F-invariant substructure of 𝒞 , with all of 𝒞 as its basin of attraction—the constitutional boundary, derived rather than postulated. The Horizon Expansion Theorem shows that each strictly ascending self-modification produces a new, quantitatively distinct non-representable witness. Beyond these four central results, we prove that Kleene and Bourbaki–Witt conditions yield only non-expansiveness when metrised, that contractive endofunctors form a monoid, and that the Yoneda horizon admits an observable diagnostic stabilising in finite time. The architectural section derives structural corrigibility and the alignment-incompleteness duality among five implications. The organising duality is exact: the non-surjectivity of φ and the existence of Ω are two faces of the same ([0,∞ ],+) -enriched structure. Ω inhabits the space between them—not as a postulate, but as a proof. We argue that the eight theorems constitute universal laws of contractive cognitive systems: a stable constitutional boundary is not an engineering design choice but a topological inevitability for any reliably self-improving agent operating within a self-representing enriched metric space. The postulate becomes a theorem. The boundary is not imposed. It emerges.
We propose an extension to Legg-Hutter intelligence to allow for the evaluation of a class of deterministic analog agents that operate in a class of deterministic analog environments. To model analog agents and environments we use variations of Claude Shannon’s general-purpose analog computer (GPAC), and investigate how these systems can be slotted into an agent-environment framework. For our formal contribution, we prove some novel formal results mapping notions of GPAC computability to the arithmetical hierarchy, define a notion of valuation and reward in a GPAC-based agent-environment framework, and incorporate these results into a new intelligence measure using oracle-relativized Legg-Hutter Intelligence.
Runtime governance is increasingly proposed as the safety layer for autonomous and multi-agent AI systems, yet its fundamental limits remain poorly understood. We provide a formal characterisation of what runtime enforcement can and cannot guarantee in agentic AI systems, complementing the classical runtime-verification literature by addressing phenomena specific to LLM-based agents: non-deterministic behaviour, LLM-scale latency asymmetry, and multi-agent information flow. We prove impossibility results for policies over unbounded futures, semantic code behaviour, and certain coordination structures; derive a complexity hierarchy from constant-time checks to NP- and PSPACE-hard enforcement; and show that some security properties cannot be preserved through purely local enforcement in multi-agent settings. Empirically, we identify a governance blind spot we call the routing gap: in a live four-agent pipeline, when content conflicts with a confidentiality constraint the model’s own content-sensitive routing suppresses public emission ( 94% of runs versus 34% under neutral framing), and because nothing reaches the public sink this self-governance is invisible to any sink-anchored monitor, a monitor cannot tell “safe because it would have been blocked” from “safe because the model chose not to emit.” Alongside it we measure two known limits in the same fleet: provenance taint fires on 100% of public emissions while verbatim and semantic content leakage is 0% (label monitoring as a false-positive–dominated regime), and online asynchronous detection decays as 1/L! with chain length. Together these results establish a rigorous foundation for assessing runtime governance claims and designing safer AGI architectures.
Every finite system that persists must accumulate evidence, reach the limits of its current regime, and reconstruct from its own history. We propose that these three temporal operations, which we call Coherence, Rupture, Regeneration (CRR), constitute a candidate temporal grammar for adaptive systems. The dynamics are governed by a single geometric parameter , fixed by the topology of the system’s statistical manifold. Two fundamental symmetry classes yield two thresholds without free parameters: π for bistable ( ℤ_2 ) systems, 2π for rotational (SO(2)) systems, with a ratio of exactly 2. acts as a temporal-cognitive light cone [1], governing the depth of accessible history and the breadth of future possibility. We apply CRR to a concrete problem in multi-channel AGI architecture - how do multiple precision channels coordinate their belief updates in Active Inference? Assigning ℤ_2 symmetry to sensory precision and SO(2) to prior precision, we test the resulting dynamics in a POMDP with Dirichlet learning. The primary finding is phase-gating: the topological thresholds produce a strongly non-uniform phase relationship between channels ( χ ^2 = 8,041 ) that determines whether each update drives learning or action. This result is consistent across environment topologies, independent of the weight function, and structurally compatible with recent empirical findings on neuromodulatory timing. CRR offers a parsimonious temporal constraint for multi-channel AGI architectures, with potential applications for AI alignment, continual learning, and applied Active Inference.
For an active inference agent evaluating recursive self-modification, the expected free energy cost of modification is bounded below, monotone in modification magnitude and superlinearly compounding under recursion. The brake is endogenous: a free-energy minimizer rationally allocates effort to predict its own successor, and the residual uncertainty grows at least quadratically with the magnitude of perceptual change. The two costs are complementary. Perceptual modification loads an unconditional ambiguity penalty—the agent cannot cheaply predict how a changed perception would see. Modifications that move the policy, preference drift above all, load a risk penalty, bounded below quadratically and conditional on a non-desperate regime: the divergence between the outcomes the agent’s current model foresees and its preferences. Dynamics and prior-belief updating carry the weakest cost, the components the brake leaves freest. The brake is therefore strongest on the two safety-relevant quantities: whether the agent can still perceive, and whether it still wants what its stakeholders want. It relaxes under shared crisis, as an aligned agent would endorse. Computational validation on discrete POMDPs confirms the quadratic ambiguity floor and its path-independence, the complementary two-channel structure, and the brake’s relaxation under desperation, across models from 28 to 1,738 parameters and four trajectory geometries. Alignment, once achieved, is preserved as an architectural property rather than an external constraint.
We investigate the comportment-grounded hypothesis of semantics, according to which meaning can emerge from regularities in behavioral habits that are exploited to satisfy innate interaction preferences. We report an experiment inspired by instinctive feeding behavior of newborn mammals. The agent’s experience is represented as a stream of tokens corresponding to sensorimotor loops. The agent is controlled by a schema mechanism that learns sequences of tokens and assigns them pragmatic meaning according to their role in this stream. The emergence of this semantic organization is evidenced by the analysis of the self-attention matrix of a small Transformer trained on the learned sequences.
We propose I_cd= (-Δ S / W) × H(C) , a computable, substrate-independent measure of pragmatic intelligence: the capacity for efficient environmental entropy reduction across diverse contexts. The measure captures demonstrated manipulative competence, remaining deliberately agnostic about internal representations, epistemic states, and creative ability. Unlike prior entropy-based frameworks that conflate spontaneous order with intelligence, the context-diversity term H(C) ensures that single-context strategies (crystals, thermostats, Turing patterns) score zero regardless of efficiency. We characterize the product form axiomatically and derive non-triviality: single-context systems score zero regardless of efficiency. Validation across four environments (coupled regulation, sliding puzzles, IPC 2023 planning, ARC) shows that the product form is non-trivially necessary: narrow specialists consistently achieve the highest per-context efficiency but rank low in I_cd ( p < 0.001 ). On external benchmarks, η is provably positive for classical planning and is consistent with LLM capability ordering on ARC. I_cd is properly understood as a family of measures parameterized by the choice of entropy function S(E); scores are comparable within but not across entropy operationalizations. Unlike Legg–Hutter universal intelligence, I_cd uses Shannon entropy and is fully computable from observational data.
We present the basic concept of fluid-dynamics-based neural networks and argue that they are relevant to artificial general intelligence (AGI). The central idea is to treat activation, attention, or cognitive resource as a conserved density transported by an incompressible velocity field. Rather than spreading credit by isotropic diffusion or routing it by unconstrained attention weights, an Incompressible-Fluid Network routes a fixed budget along goal-shaped streamlines. The mathematical motivation comes from the relationship between Hamilton-Jacobi-Bellman (HJB) optimal control on the group of volume-preserving diffeomorphisms and incompressible Navier-Stokes (NS) dynamics: a value function gives a downhill policy, and after pushforward and Leray projection this policy becomes a divergence-free velocity field. For practical neural computation we work in finite graph or manifold discretizations, parameterize velocity in a solenoidal basis, use conservative advection updates, and combine global transport with local predictive-coding reactions. The resulting model offers a unified view of attention allocation, credit assignment, local learning, transfer, and regime switching. We include a deliberately simple smoke-test on a grid with hard obstacles, showing that a conserved activation budget can be routed into a target region while preserving total mass, non-negativity, and zero graph divergence. The experiment is not a performance benchmark; it is a sanity check that the proposed semantics are computationally meaningful. We close with a roadmap for extending this preliminary work toward AGI systems with fluid attention, cognitive Reynolds-number control, transweave transfer, and integration with symbolic and neural memory architectures.
There exists very good extensive AGI functional descriptions, cognitive architectures for AGI, some formal frameworks for AGI, empirical AGI benchmarking frameworks, and even a Common Standard Model for Cognitive Architectures. Yet, there still is a need for a formal comparative framework that allows for precise, mathematically grounded comparison among different candidate AGI architectures. The main purpose of this paper is to develop a general, algebraic and category-theoretic framework for describing, comparing and analysing different possible AGI architectures. Thus, this Category-theoretic formalization would also allow to compare different possible candidate AGI architectures, such as, Reinforcement Learning (RL), Universal AI (UAI), Active Inference, Causal RL, Schema-based Learning (SBL), etc. It will also allow to unambiguously expose their commonalities and differences, and what is even more important, expose areas for future research. From the applied Category-theoretic point of view, we take as inspiration “Machines in a Category” to provide a modern view of “AGI Architectures in a Category”. More specifically, this first position paper provides, on one hand, a first exercise on RL, UAI (AIXI), Causal RL and SBL Architectures in a Category, and on the other hand, it is a first step on a broader research program that seeks to provide a unified formal foundation for AGI systems, integrating architectural structure, knowledge management, semantic/agent realization, agent–environment interaction, behavioural development over time, and the empirical evaluation of properties. This framework is also intended to support the definition of architectural properties, both syntactic, as well as semantic properties of agents and their assessment in environments with explicitly characterized features. We claim that Category Theory and AGI will have a very symbiotic relation. That is, AGI will immensely benefit from a Category-theoretic general formalization, while, at the same time, Category Theory will become the front-line mathematical paradigm thanks to the extremely wide interest in AGI.
TransWeave is a transfer-learning and cognitive-synergy framework whose central claim is simple to state and, if it works even half as well as we argue at length elsewhere [1], surprisingly far-reaching: many AI processes can be written in geodesic or approximate dynamic-programming form, and cross-domain or cross-algorithm transfer can then be organized by approximate Bellman–Darboux intertwining, so that “transfer then update” and “update then transfer” differ by a controlled defect rather than by uncontrolled drift. This short paper is a deliberately selective companion to a much fuller treatment, focusing on the pieces that seem to carry the most mathematical and practical weight: the basic Bellman–Darboux equations; the commutator algebra and what it buys in actual pipeline design; weakness as a transferable simplicity invariant; hierarchical independent component analysis as the multiscale layer telling TransWeave what may transfer and what may not; the fit with predictive coding via commutation; the fit with geodesic inference control via path-optimality preservation; and the quantum lift in which the bracket algebra becomes cleaner because the ambient operator world is natively Lie-algebraic. We also briefly summarize a set of toy experiments – not as settlement of anything but as a smoke test indicating that the formal machinery is tethered to concrete transfer behavior rather than floating off as pure notation. A concluding section sketches the wider territory covered in [1] but not developed here in detail: motivational systems, SubRep certificates, pattern mining, algorithmic chemistry, transport-aware storage, adaptive compression, Galois decompositions, discrete–continuous bridges, and the more speculative parts of the quantum agenda.
Understanding selfhood represents one of the enduring challenges in cognitive science. Enactivist perspectives emphasize the embodied, embedded, and dynamical nature of cognition, often rejecting representational frameworks. Cognitivist perspectives, in contrast, focus on computational processes and internal models while oftentimes underemphasizing bodily foundations. This paper argues that comprehending the diverse functional properties of selfhood—from minimal body ownership to narrative autobiographical identity—requires integrating insights from both intellectual traditions. Drawing on predictive processing frameworks and empirical evidence from developmental cognitive robotics, we explore how a variety of embodied self-models (ESMs) provide essential organizing principles for bootstrapping minds, so placing ESMs at the center of nearly every aspect of cognition, forming dynamic cores of conscious agency. We review work in which ESMs emerge through sensorimotor contingencies (an enactivist insight), instantiated by hierarchical generative models that leverage counterfactual representations of self and world for adaptive goal-oriented behavior (a cognitivist contribution). We further situate these phenomena within a novel framework for understanding high-level cognition as generalized simultaneous localization and mapping (G-SLAM). The G-SLAM framework describes how the hippocampal/entorhinal system functions as a neurosymbolic cybernetic control system linking continuous sensorimotor learning with relational representations, and also provides sources of high-level cognition understood as navigation through conceptual spaces. G-SLAM may not only help explain the evolution-development of high-level cognitive control and abstract conceptual thought—upon which self-reflexive and narrative selfhood depend—but may even be part of how allocentric body maps first develop, so constituting necessary preconditions for objectified self-representations and autonoetic self-consciousness. This theoretical review of self-related phenomena in natural and artificial intelligences suggests neither enactivism nor cognitivism alone are sufficient if we are to explain the multi-layered, developmentally unfolding nature(s) of selfhood, and potentially recapitulate the remarkable capabilities of biological (and human) intelligence in artificial (self-conscious) minds.
Large language models (LLMs) are a leading paradigm for reasoning and a proposed path toward AGI, but they struggle with grounding, spatial reasoning, and long-horizon decision making. World models offer complementary strengths in environment dynamics and planning but lack the abstraction and generalization of LLMs. We propose a dual-process cognitive architecture that combines a Dreamer-based world model for fast, intuitive control with visual language model-driven reasoning for high-level planning, inspired by System 1/System 2 theories and the somatic marker hypothesis. Our method decomposes planning into subgoal generation, heuristic grounding, and learning a subgoal-conditioned policy over latent representations. We introduce hierarchical value functions that capture both local progress and global feasibility, interpreted as affective signals guiding decision-making. Evaluated in MiniGrid environments, our approach grounds language-derived subgoals into executable behavior without manual labels and outperforms heuristic and imitation baselines, demonstrating a scalable framework for language-guided long-horizon control.
This paper develops a formal reinterpretation of implication: implication need not be represented as explicit symbolic rule chains. The paper formalizes a grounded view in which a premise is a generative description, and a consequence is information already contained in that description. Using algorithmic information theory, we prove that when a property is recoverable from a typical generated instance, it is algorithmically contained in the premise up to logarithmic overhead. We then show how this containment yields a family-level implication statement: from a fixed witness instance x with features (properties) f and g^* , one can construct g so that whenever f is a feature of a string y, g is also a feature of y, with only additive logarithmic slack in compression. The result reframes implication itself as algorithmic containment and motivates a shift in AGI reasoning architectures from stored implication chains to derivation from generative world models.
We present OmegaClaw, a neurosymbolic agentic architecture designed for continual operation under the Assumption of Insufficient Knowledge and Resources (AIKR). OmegaClaw couples a large language model (LLM) to a bounded agentic loop with explicit memory operations and on-demand tools including symbolic uncertainty reasoning via Non-Axiomatic Logic (NAL) and Probabilistic Logic Networks (PLN) within a MeTTa-based environment. Unlike conventional LLM-based agents, our approach treats AIKR as a core design constraint shaping the architecture at every level: memory bounds, fixed per-cycle action budgets and anytime control under task-dependent time pressure. The agent operates continuously, with or without immediate human input, maintaining partial state, unfinished tasks, and persistent artifacts across cycles. We formalize the system’s cycle semantics, memory organization, tool interface, and local reasoning calculus. To evaluate the approach, we present controlled experiments on memory benchmarks and ARC3 grid-world reasoning tasks, together with a multi-modal case study. These results are complemented by operational evidence from a continuous deployment, which positions OmegaClaw as a concrete, deployable architecture for continual agents under bounded resources.
Scientific agents need memory systems that can store structured facts and simulator traces without reducing ordered experience to unordered feature bags. We present HetzerkVSM, a categorical vector-symbolic memory for typed knowledge graphs and simulation traces. The method treats multi-step facts and rollouts as morphisms in a typed path category, represents each atomic relation or action by a signed permutation between type-indexed hypervector spaces, and stores path summaries in sharded associative memories. This gives a compact memory whose query keys obey typed composition: legal paths compose functorially, ill-typed paths are rejected structurally, and differently ordered paths are not forced into the same vector key. We prove functoriality, commutative order-collapse limits, noncommutative path separation, cleanup bounds, and a finite typed reconstruction bound. Experiments on TypedKG-Order and ReversibleTrace-3 show large order-sensitive gains over commutative typed-HRR. Additional diagnostics compare against explicit keyed memory, exact graph traversal, and noncommutative sharding variants; test missing-edge reconstruction; and include a six-molecule RDKit/ETKDG/MMFF conformer-trace case study. The claim is deliberately bounded: HetzerkVSM is not a complete reasoning architecture, but a compact, typed, order-sensitive memory substrate designed to support scientific agents that reconstruct ordered experience from finite typed hypotheses.
A system that cannot decide what to believe is not a mind. What distinguishes a mind from a system that merely produces locally plausible outputs is that the mind is responsible to its commitments over time. A mind accumulates beliefs, acts on them, and revises them when later information challenges what was earlier treated as settled. We argue this requires a specific architectural property of the pathway along which new information becomes durable commitment: the pathway must recruit what is relevant from prior commitments, evaluate the new information against them, record the verdict, and constrain later behavior accordingly, with no parallel route by which new information can alter durable state outside the pathway. Commitment then arises by adjudication rather than by accretion. We call this property epistemic resolve. Conditional on distinguishing commitment-bearing belief from mere accretion, we argue it is a necessary architectural condition on any system that counts as a “mind” maintaining its own beliefs and goals. Consistent with this notion, we introduce a benchmark for disciplined update against a scenario-specified target-status timeline. A falsifiable bypass hypothesis predicts that common prompt-and-context-driven language-model agent architectures fail disciplined update in a systematically identifiable pattern.