In this paper I provide two simple and straightforward semantic analyses of Aristotle’s apodeictic syllogisms. It is shown that both an intensional and an extensional semantics can be given, such that for any two terms S and T, the intension of S is ‘part of’ the intension of T iff the extension of T is ‘part of’ the extension of S, just as Leibniz assumed for the semantics of assertoric syllogisms.
Large language models (LLMs) have achieved impressive progress in natural language processing tasks but still struggle with complex logical reasoning. We observe that in propositional logic question-answering (QA), LLMs' performance varies with the order of training samples during fine-tuning. Motivated by this, we propose a data-driven approach to automatically determine the fine-tuning sample order, enhancing the logical QA performance of LLMs. Specifically, we first quantify the logical reasoning complexity of propositional reasoning samples and then stratify the training data into several subsets of ascending complexity. Subsequently, we fine-tune the LLMs on these subsets, progressing from low to high reasoning complexity. Experimental results demonstrate that our approach outperforms single-stage fine-tuning baselines across diverse reasoning benchmarks.
What is the relation between the two main problems arising from vagueness, the Sorites Paradox and the Problem of the Many? This question seems to be neglected. In explaining this relation, this paper shows that the usual understanding of these problems is unsatisfactory and demonstrates what instead is fundamental to these problems. The usual understanding of the Sorites Paradox is that it is a problem arising for (apparently) vague concepts, while the Problem of the Many is understood to be a problem arising for ordinary objects. This paper, however, shows that both problems can arise for any kind of phenomenon of vagueness. Instead, what is fundamental to them is the number of boundary crossings that are involved, a novel notion introduced and further explained in this paper. Whereas the Problem of the Many arises for a collection of multiple boundary crossings, the Sorites Paradox arises for one single boundary crossing.
Identifying causal relationships rather than spurious correlations between words and class labels plays a crucial role in building robust text classifiers. Previous studies proposed using causal effects to distinguish words that are causally related to the sentiment, and then building robust text classifiers using words with high causal effects. However, we find that when a sentence has multiple causally related words simultaneously, the magnitude of causal effects will be significantly reduced, which limits the applicability of previous causal effect-based methods in distinguishing causally related words from spuriously correlated ones. To fill this gap, in this paper, we introduce both the probability of necessity (PN) and probability of sufficiency (PS), aiming to answer the counterfactual question that ‘if a sentence has a certain sentiment in the presence/absence of a word, would the sentiment change in the absence/presence of that word?’. Specifically, we first derive the identifiability of PN and PS under different sentiment monotonicities, and calibrate the estimation of PN and PS via the estimated average treatment effect. Finally, the robust text classifier is built by identifying the words with larger PN and PS as causally related words, and other words as spuriously correlated words, based on a contrastive learning approach name CPNS is proposed to achieve robust sentiment classification. Extensive experiments are conducted on public datasets to validate the effectiveness of our method.
Large language models (LLMs) have achieved remarkable successes on various tasks. However, recent studies have found that there are still significant challenges to the logical reasoning abilities of LLMs, which can be categorized into the following two aspects: (1) Logical question answering: LLMs often fail to generate the correct answer within a complex logical problem which requires sophisticated deductive, inductive or abductive reasoning given a collection of premises and constrains. (2) Logical consistency: LLMs are prone to producing responses contradicting themselves across different questions. For example, a state-of-the-art question-answering LLM Macaw, answers Yes to both questions Is a magpie a bird? and Does a bird have wings? but answers No to Does a magpie have wings?. To facilitate this research direction, we comprehensively investigate the most cutting-edge methods and propose a detailed taxonomy. Specifically, to accurately answer complex logic questions, previous methods can be categorized based on reliance on external solvers, prompts, and fine-tuning. To avoid logical contradictions, we discuss concepts and solutions of various logical consistencies, including implication, negation, transitivity, factuality consistencies, and their composites. In addition, we review commonly used benchmark datasets and evaluation metrics, and discuss promising research directions, such as extending to modal logic to account for uncertainty and developing efficient algorithms that simultaneously satisfy multiple logical consistencies.
While causal reasoning is a core facet of our cognitive abilities, its time-course has not received proper attention. As the duration of reasoning might prove crucial in understanding the underlying cognitive processes, we asked participants in two experiments to make probabilistic causal inferences while manipulating time pressure. We found that participants are less accurate under time pressure, a speed-accuracy-tradeoff, and that they respond more conservatively. Surprisingly, two other persistent reasoning errors—Markov violations and failures to explain away—appeared insensitive to time pressure. These observations seem related to confidence: Conservative inferences were associated with low confidence, whereas Markov violations and failures to explain were not. These findings challenge existing theories that predict an association between time pressure and all causal reasoning errors including conservatism. Our findings suggest that these errors should not be attributed to a single cognitive mechanism and emphasize that causal judgements are the result of multiple processes.
Generic sentences (e.g., “Dogs bark”) express generalizations about groups or individuals. Accounting for the meaning of generic sentences has been proven challenging, and there is still a very lively debate about which factors matter for whether or not we a willing to endorse a particular generic sentence. In this paper we study the effect of impact on the assertability of generic sentences, where impact refers to the dangerousity of the property the generic is ascribing to a group or individual. We run three preregistered experiments, testing assertability and endorsement of novel generic sentences with visual and textual stimuli. Employing Bayesian statistics we found that impact influences the assertability, and endorsement, of generic statements. However, we observed that the size of the effect impact value may have been previously overestimated by theoretical and experimental works alike. We also run an additional descriptive survey testing standard examples from the linguistic literature and found that at least for some of the examples endorsement appears to be lower than assumed. We end with exploring possible explanations for our results.
Recent scholarship on reasoning in LLMs has supplied evidence of impressive performance and flexible adaptation to machine generated or human feedback. Nonmonotonic reasoning, crucial to human cognition for navigating the real world, remains a challenging, yet understudied task. In this work, we study nonmonotonic reasoning capabilities of seven state-of-the-art LLMs in one abstract and one commonsense reasoning task featuring generics, such as 'Birds fly', and exceptions, 'Penguins don't fly' (see Fig. 1). While LLMs exhibit reasoning patterns in accordance with human nonmonotonic reasoning abilities, they fail to maintain stable beliefs on truth conditions of generics at the addition of supporting examples ('Owls fly') or unrelated information ('Lions have manes'). Our findings highlight pitfalls in attributing human reasoning behaviours to LLMs, as well as assessing general capabilities, while consistent reasoning remains elusive.
Conversational AI is a game-changer for science. Here's how to respond. Conversational AI is a game-changer for science. Here's how to respond.
Establish an independent scientific body to test and certify generative artificial intelligence, before the technology damages science and public trust.
The latest generation of LLMs can be prompted to achieve impressive zero-shot or few-shot performance in many NLP tasks. However, since performance is highly sensitive to the choice of prompts, considerable effort has been devoted to crowd-sourcing prompts or designing methods for prompt optimisation. Yet, we still lack a systematic understanding of how linguistic properties of prompts correlate with task performance. In this work, we investigate how LLMs of different sizes, pre-trained and instruction-tuned, perform on prompts that are semantically equivalent, but vary in linguistic structure. We investigate both grammatical properties such as mood, tense, aspect and modality, as well as lexico-semantic variation through the use of synonyms. Our findings contradict the common assumption that LLMs achieve optimal performance on lower perplexity prompts that reflect language use in pretraining or instruction-tuning data. Prompts transfer poorly between datasets or models, and performance cannot generally be explained by perplexity, word frequency, ambiguity or prompt length. Based on our results, we put forward a proposal for a more robust and comprehensive evaluation standard for prompting research.
In Suppose and Tell , Williamson makes a new and original attempt to defend the material conditional account of indicative conditionals. His overarching argument is that this account offers the best explanation of the data concerning how people evaluate and use such conditionals. We argue that Williamson overlooks several important alternative explanations, some of which appear to explain the relevant data at least as well as, or even better than, the material conditional account does. Along the way, we also show that Williamson errs at important junctures about what exactly the relevant data are.
The inferences of contraposition (A ⇒ C ∴ ¬C ⇒ ¬A), the hypothetical syllogism (A ⇒ B, B ⇒ C ∴ A ⇒ C), and others are widely seen as unacceptable for counterfactual conditionals. Adams convincingly argued, however, that these inferences are unacceptable for indicative conditionals as well. He argued that an indicative conditional of form A ⇒ C has assertability conditions instead of truth conditions, and that their assertability ‘goes with’ the conditional probability p(C|A). To account for inferences, Adams developed the notion of probabilistic entailment as an extension of classical entailment. This combined approach (correctly) predicts that contraposition and the hypothetical syllogism are invalid inferences. Perhaps less well-known, however, is that the approach also predicts that the unconditional counterparts of these inferences, e.g., modus tollens (A ⇒ C, ¬C ∴ ¬A), and iterated modus ponens (A ⇒ B, B ⇒ C, A ∴ C) are predicted to be valid. We will argue both by example and by calling to the results from a behavioral experiment (N = 159) that these latter predictions are incorrect if the unconditional premises in these inferences are seen as new information. Then we will discuss Adams’ (1998) dynamic probabilistic entailment relation, and argue that it is problematic. Finally, it will be shown how his dynamic entailment relation can be improved such that the incongruence predicted by Adams’ original system concerning conditionals and their unconditional counterparts are overcome. Finally, it will be argued that the idea behind this new notion of entailment is of more general relevance.
In this paper we argue that the antecedent of a (non-analytic) conditional is causally relevant to the consequent, ... at least if standard background conditions hold. Natural counterexamples to the causal relevance analysis are argued to be cases where the standardly assumed background condition(s) do not hold.
This paper explores the relations between two logical approaches to vagueness: on the one hand the fuzzy approach defended by [Smith, 2008], and on the other the strict-tolerant approach defended by [Cobreros et al., 2012]. Although the former approach uses continuum many values and the latter implicitly four, we show that both approaches can be subsumed under a common three-valued framework. In particular, we defend the claim that Smith’s continuum many values are not needed to solve what Smith calls ‘the jolt problem’, and we show that they are not needed for his account of logical consequence either. Not only are three values enough to satisfy Smith’s central desiderata, but they also allow us to internalize Smith’s closeness principle in the form of a tolerance principle at the object-language. The reduction, we argue, matters for the justification of many-valuedness in an adequate theory of vague language.
In this paper, we explain why the antecedent of a biscuit conditional is relevant to its consequent by extending Douvenʼs evidential support theory of conditionals making use of utilities. By this extension, we can also explain why a biscuit conditional gives rise to the inference that the consequence is (most likely) true. Finally, we account for the intuition that (indicative) biscuit sentences are false when the antecedent is false and allow for counterfactual biscuits.
In this paper, we investigate what types of stereotypical information are captured by pretrained language models. We present the first dataset comprising stereotypical attributes of a range of social groups and propose a method to elicit stereotypes encoded by pretrained language models in an unsupervised fashion. Moreover, we link the emergent stereotypes to their manifestation as basic emotions as a means to study their emotional effects in a more generalized manner. To demonstrate how our methods can be used to analyze emotion and stereotype shifts due to linguistic experience, we use fine-tuning on news sources as a case study. Our experiments expose how attitudes towards different social groups vary across models and how quickly emotions and stereotypes can shift at the fine-tuning stage.
AbstractIn this paper we argue that a typical member of a class, or category, is an extreme, rather than a central, member of this category. Making use of a formal notion of representativeness, we can say that a typical member of a category is a stereotype of this category. In the second part of the paper we show that this account of typicality can be given a rational motivation by providing a game-theoretical derivation.
In the seminal work by Awodey and Warren it was shown that the intensional identity types of Martin-Löf dependent type theory can be modelled categorically using weak factorisation systems. In this interpretation the dependent types are modelled by fibrations, i.e. the right maps of a weak factorisation system. This work inspired a lot of further research into such categorical models of identity types. Recently it was adapted by Gambino and Larrea to the setting of algebraic weak factorisation systems who added interpretations of the dependent sum and product types of said type theory. In their work the dependent types are interpreted using the algebras of the pointed endofunctor of the system, and in the present work we show that the same approach also works when we instead use the algebras for the monad of the system.
It is well known that in his Prior Analysis, Aristotle presents the system of syllogisms. Although many commentators consider Aristotle’s system of modal syllogisms almost impossible to understand from a modern point of view or even inconsistent, many philosophers still tried to account for these claims by looking for a consistent semantics of it. In this paper we will argue for a causal analysis of modal categorical sentences based on the notion of causal power. According to Cheng (1997), the causal power of A to produce B can be measured probabilistically. Based on Cheng’s hypothesis, we will derive a qualitative semantics for modal categorical sentences. We will argue that our approach fits well with Aristotle’s analysis of real definition in the Posterior Analytics, and that in this way we can account in a relatively straightforward way (using just Venn diagrams) for several puzzling aspects of Aristotle’s system of modal syllogisms.