This paper presents an argument that certain AI safety measures, rather than mitigating existential risk, may instead exacerbate it. Under certain key assumptions - the inevitability of AI failure, the expected correlation between an AI system's power at the point of failure and the severity of the resulting harm, and the tendency of safety measures to enable AI systems to become more powerful before failing - safety efforts have negative expected utility. The paper examines three response strategies: Optimism, Mitigation, and Holism. Each faces challenges stemming from intrinsic features of the AI safety landscape that we term Bottlenecking, the Perfection Barrier, and Equilibrium Fluctuation. The surprising robustness of the argument forces a re-examination of core assumptions around AI safety and points to several avenues for further research.
This work defends the 'Whole Hog Thesis': sophisticated Large Language Models (LLMs) like ChatGPT are full-blown linguistic and cognitive agents, possessing understanding, beliefs, desires, knowledge, and intentions. We argue against prevailing methodologies in AI philosophy, rejecting starting points based on low-level computational details ('Just an X' fallacy) or pre-existing theories of mind. Instead, we advocate starting with simple, high-level observations of LLM behavior (e.g., answering questions, making suggestions) – defending this data against charges of metaphor, loose talk, or pretense. From these observations, we employ 'Holistic Network Assumptions' – plausible connections between mental capacities (e.g., answering implies knowledge, knowledge implies belief, action implies intention) – to argue for the full suite of cognitive states. We systematically rebut objections based on LLM failures (hallucinations, planning/reasoning errors), arguing these don't preclude agency, often mirroring human fallibility. We address numerous 'Games of Lacks', arguing that LLMs do not lack purported necessary conditions for cognition (e.g., semantic grounding, embodiment, justification, intrinsic intentionality) or that these conditions are not truly necessary, often relying on anti-discriminatory arguments comparing LLMs to diverse human capacities. Our approach is evidential, not functionalist, and deliberately excludes consciousness. We conclude by speculating on the possibility of LLMs possessing 'alien' contents beyond human conceptual schemes.
AlphaGo plays chess and Go in a creative and novel way. It is natural for us to attribute contents to it, such as that it doesn't view being several pawns behind, if it has more board space, as bad. The framework introduced in Cappelen and Dever (2021) provides a way of thinking about the semantics and the metasemantics of AI content: does AlphaGo entertain contents like this, and if so, in virtue of what does a given state of the program mean that particular content? One salient question Cappelen and Dever didn't consider was the possibility of alien content. Alien content is content that is not or cannot be expressed by human beings. It's highly plausible that AlphaGo, or any other sophisticated AI system, expresses alien contents. That this is so, moreover, is plausibly a metasemantic fact: a fact that has to do with how AI comes to entertain content in the first place, one that will heed the vastly different etiology of AI and human content. This chapter explores the question of alien content in AI from a semantic and metasemantic perspective. It lays out the logical space of possible responses to the semantic and metasemantic questions alien content poses, considers whether and how we humans could communicate with entities who express alien content, and points out that getting clear about such questions might be important for more 'applied' issues in the philosophy of AI, such as existential risk and XAI.
This article challenges conventional boundaries between human and artificial cognition by examining introspective capabilities in large language models (LLMs). Although humans have traditionally been considered unique in their ability to reflect on their own mental states, we argue that LLMs may not only possess genuine introspective abilities but potentially excel at them compared to humans. We discuss five objections to machine introspection: (1) the lack of direct routes to self‐knowledge in training data, (2) the conflict between static knowledge and dynamic mental states, (3) the distorting effects of reinforcement learning on self‐reports, (4) LLMs own denials of inner experience, and (5) arguments that LLMs simply mimic language without understanding. We think all these arguments fail and that there are deep parallels between human and machine introspection. Most provocatively, we propose that LLMs superior processing capabilities and pattern recognition may enable them to develop more sophisticated theories of mind than humans possess, potentially making them more reliable introspectors than their creators. If we are right, this has significant implications for artificial intelligence (AI) alignment, transparency, and our understanding of the nature of AI.
Can humans and artificial intelligences share concepts and communicate? 'Making AI Intelligible' shows that philosophical work on the metaphysics of meaning can help answer these questions. Herman Cappelen and Josh Dever use the externalist tradition in philosophy to create models of how AIs and humans can understand each other. In doing so, they illustrate ways in which that philosophical tradition can be improved. The questions addressed in the book are not only theoretically interesting, but the answers have pressing practical implications. Many important decisions about human life are now influenced by AI. In giving that power to AI, we presuppose that AIs can track features of the world that we care about (for example, creditworthiness, recidivism, cancer, and combatants). If AIs can share our concepts, that will go some way towards justifying this reliance on AI. This ground-breaking study offers insight into how to take some first steps towards achieving Interpretable AI.
There is a view prevalent among people working in the artificial intelligence field to the effect that philosophers have nothing to tell us about AI and (putative) AI communication—that philosophy cannot help with the mathematical problems of making practical advances in AI and is, therefore, no more than a diverting irrelevance. This chapter rebuts that view. It takes the form of a dialogue between a philosopher and someone working in AI who is sceptical about philosophy’s relevance to AI. The sceptic, Alfred, argues that philosophical issues about the nature of communication are irrelevant to ongoing work in AI; the philosopher responds, showing that the sceptic’s supposedly unphilosophical perspective in fact harbours philosophical presuppositions, and ones that are worth discussing—in particular, the question of meaning and content within AI systems.
This chapter continues the process of anthropocentric abstraction, here concentrating on proper names. Do AI systems use proper names? Using our example of ‘SmartCredit’, it highlights problems concerning how to treat the output of an AI system when some, but not all or most, of the information in its neural network fails to apply to the individual we interpret the output to be about. After giving reasons to think the standard Kripkean theory might not work well here, it suggests an alternative theory of communication about particular entities, the mental file framework, which is more apt for theorizing about AI systems. It then abstracts from the human-centric features of extant theories of mental files to consider how AI might use something like them to refer to particulars.
Can humans and artificial intelligences share concepts and communicate? One aim of Making AI Intelligible is to show that philosophical work on the metaphysics of meaning can help answer these questions. Cappelen and Dever use the externalist tradition in philosophy of to create models of how AIs and humans can understand each other. In doing so, they also show ways in which that philosophical tradition can be improved: our linguistic encounters with AIs revel that our theories of meaning have been excessively anthropocentric. The questions addressed in the book are not only theoretically interesting, but the answers have pressing practical implications. Many important decisions about human life are now influenced by AI. In giving that power to AI, we presuppose that AIs can track features of the world that we care about (e.g. creditworthiness, recidivism, cancer, and combatants.) If AIs can share our concepts, that will go some way towards justifying this reliance on AI. The book can be read as a proposal for how to take some first steps towards achieving interpretable AI. Making AI Intelligible is of interest to both philosophers of language and anyone who follows current events or interacts with AI systems. It illustrates how philosophy can help us understand and improve our interactions with AI.
This short chapter does two things. First, it shows that in fact workers in AI frequently talk as if AI systems express contents. We present the argument that the complex nature of the actions and communications of AI systems, even if they are very different from the complex behaviours of human beings, and the way they have ‘aboutness’, strongly suggest a contentful interpretation of those actions and communications. It then introduces some philosophical terminology that captures various aspects of language use, such as the ones in the title, to better make clear what one is saying—philosophically speaking—when one claims AI systems communicate, and to provide a vocabulary for the next few chapters.
The previous chapters have given us ways of thinking about how an AI system might use names and predicates. But language use involves more than simply tokening expressions. It also involves predicating, or asserting, or judging: applying predicates to terms to make a claim. How can AI do that, even granting it can name things and express predicates? This chapter proposes an answer, by melding together two popular theories: the act-theory of propositional content, and teleosemantics. In a now familiar way, it abstracts from human-centric features of extant theories to show how we can understand AI predication.
Linguistic intervention in rational decision making is standardly captured in terms of information change. But the standard view gives us no way to model interventions involving expressions that only have an attentional effects on conversational contexts. How are expressions with non-informational content - like epistemic modals - used to intervene in rational decision making? We show how to model rational decision change without information change: replace a standard conception of value (on which the value of a set of worlds reduces to values of individual worlds in the set) with one on which the value of a set of worlds is determined by a selection function that picks out a generic member world. We discuss some upshots of this view for theorizing in philosophy and formal semantics.
The final chapter considers or reconsiders four topics that are important for philosophers coming to terms with AI communication: the fact that AI systems’ goals might change without human intervention and become misaligned with humans’ goals. It is argued that such possibilities make it both particularly important but also particularly difficult to give theories of AI communication. The second topic is the external mind hypothesis: the author considers its relevance for AI systems. The third considers what we can learn about so-called adversarial perturbations, and suggests they can help us reply to the sceptic Alfred from Chapter 2. Finally, the chapter concludes by considering again explainable AI, suggesting that the externalistic perspective offers can help us understand what we can and cannot require of explainable AI systems.
This chapter introduces the central claim of the book about AI communication. It argues first that we should understand AI communication in terms of externalism, the thought that the semantic content an entity can express is determined to a large extent by the environment it finds itself in, rather than by its internal states. It then argues that existing externalist theories are too human-centric: they concentrate on peculiarities of human beings not shared by AI systems. It accordingly proposes to abstract from those human peculiarities when developing theories of communication for AI. The chapter ends by discussing how to decide between competing metasemantic frameworks such as externalism and internalism.
Herman Cappelen and Josh Dever consider whether we can distinguish between ideal and Non-Ideal Philosophy of Language in the way that a philosopher like Charles Mills does for political philosophy. They argue that there is no deep distinction between the sort of philosophy collected in this volume and the more mainstream material.
This chapter begins to flesh out the theory of de-anthropocentrized externalism, by considering what we should say about AI systems’ use of predicates, such as ‘is a benign lesion’ or ‘will default on a loan’. Introducing Kripke’s seminal causal theory of names, the chapter shows how (and why it’s necessary) to abstract from the theory’s anthropocentric features while still preserving its key features, such as the idea that use of a term should be anchored in a baptismal event and that the term’s reference is passed on from speaker to speaker.
This chapter provides a new argument that de se attitudes play no essential role in understanding actions. It turns on the nature of corporate agency. It gives some representative examples of corporations and their actions, with the goals of (a) making plausible that corporate entities can undertake genuine action and (b) showing some ways in which those actions can float free of the intentions and actions of individuals composing and associated with the corporate entities. The chapter discusses that several lines of thought that people have taken to show that action is impossible without de se thoughts are not only uncompelling in the corporate case, but turn out to have their grip in the individual case weakened by virtue of seeing why they fail in the corporate case.
Aristotle observes that substances have no contraries. Consider one possible role that the contrary of a substance might play, were it to exist. Just as objects serve as guarantors of the instantiation of properties, contraries of objects could serve as guarantors of the non-instantiation of properties. By first considering a reframing of deontic logic that takes ‘being permitted’ rather than ‘being forbidden’ as the default state, I develop logical tools that allow the construction of extremal models in which everything is permitted and nothing is forbidden. Transferred to quantified logic, these same tools give us a conception of antiobjects which block property instantiation. The resulting picture of antiobjects is then used as a test case to examine questions about the role of symmetry considerations in philosophical methodology.
Montague and Kaplan began a revolution in semantics, which promised to explain how a univocal expression could make distinct truth-conditional contributions in its various occurrences. The idea was to treat context as a parameter at which a sentence is semantically evaluated. But the revolution has stalled. One salient problem comes from recurring demonstratives: “He is tall and he is not tall”. For the sentence to be true at a context, each occurrence of the demonstrative must make a different truth-conditional contribution. But this difference cannot be accounted for by standard parameter sensitivity. Semanticists, consoled by the thought that this ambiguity would ultimately be needed anyhow to explain anaphora, have been too content to posit massive ambiguities in demonstrative pronouns. This chapter aims to revived the parameter revolution by showing how to treat demonstrative pronouns as univocal while providing an account of anaphora that doesn’t end up re-introducing the ambiguity.
In her very interesting 'First-personal modes of presentation and the problem of empathy' (2017, 315-336), L. A. Paul argues that the phenomenon of empathy gives us reason to care about the first person point of view: that as theorists we can only understand, and as humans only evince, empathy by appealing to that point of view. We are skeptics about the importance of the first person point of view, although not about empathy. The goal of this paper is to see if we can account for empathy without the ideology of the first person. We conclude that we can.
In her Transient Truths (New York: Oxford University Press, 2012), Berit Brogaard defends temporalism about proposition content from the more traditional eternalist views. I argue that both temporalism and eternalism are equally capable of accommodating all the data, and thus suggest that we should adopt a neutralism that holds there is no serious or resolvable dispute. Contra Brogaard, I argue that neither disagreement patterns nor belief dynamics favor temporalism over eternalism. I also suggest that Brogaard's defense of operator over quantificational semantics for tense is unnecessary, because quantificational tense semantics can be given a temporalist reading, and operator tense semantics can be given an eternalist reading.