I begin with an intuitively plausible principle: if proposition H retrodicts or predicts that proposition D is true, then Pr(D|H) > ½. I then consider the difference between forward-directed and backward-directed conditional probabilities. Until the 1980s, population genetics theory was dominated by the former; the situation changed with the appearance of coalescent theory, which deploys the latter. A forward-directed probability model cannot retrodict, and a backward-directed model cannot predict. A Bayesian model will do both if it deploys both types of conditional probability. I then describe ideas from model selection theory in statistics that throw doubt on the idea that true models of a process will always be more accurate predictors or retrodictors than false models of that process. I then discuss theorems in information theory that show how the passage of time degrades information in a causal chain. The rate of information loss can be reduced and even cancelled by the branching process in phylogenetic trees. The separation of within-lineage from across-lineage assessments of information loss is an instance of Simpson's paradox.
Suppose 50-year-old Sue now has lung cancer, due to the fact that C = c A = g (meaning that Sue smoked c cigarettes and inhaled g grams of asbestos over the previous 30 years), and that neither cause caused the other. Given this, a retrospective question arises – did one of those actual causes have a stronger influence than the other on her getting lung cancer? We propose a “zeroing-out” criterion for making sense of this question; it says that C = c was a stronger causal influence than A = g precisely when Pr(lung cancer | C = c A = g) – Pr(lung cancer | C = 0 A = g) > Pr(lung cancer | C = c A = g) – Pr(lung cancer | C = c A = 0). Zeroing-out uses Pr(lung cancer | C = c A = g) as a baseline and relates that baseline to two counterfactual probabilities. We discuss how zeroing-out applies to three evolutionary examples − the influences of selection and drift on the fixation of an allele, the influences of group and individual selection on the evolution of altruism, and the influences of stabilizing selection and ancestral influence (aka “phylogenetic inertia”) on the evolution of tetrapody in land vertebrates. Zeroing-out differs from an “adding-in” criterion, which uses Sue’s probability of having lung cancer at age 50, given her actual state at age 20 (at which time she was cancer free and C = 0 A = 0) as a baseline and asks whether her risk of having lung cancer at age 50 would be greater if C = c were true of the 30 years in between than it would be if A = g were true of those years. Zeroing-out and adding-in generate identical criteria for comparing causal influences in this example because the probabilities are related “monotonically” (a concept we define). We then describe examples in which monotonicity fails and the two criteria differ. We prove theorems that describe when the two criteria disagree and when they do not. We then consider how zeroing-out and adding-in are related to six quantitative measures of causal strength that have been proposed. Inter alia, we discuss how our framework is related to interventionism.
Ockham's razor, the principle of parsimony, states that simpler theories are better than theories that are more complex. It has a history dating back to Aristotle and it plays an important role in current physics, biology, and psychology. The razor also gets used outside of science - in everyday life and in philosophy. This book evaluates the principle and discusses its many applications. Fascinating examples from different domains provide a rich basis for contemplating the principle's promises and perils. It is obvious that simpler theories are beautiful and easy to understand; the hard problem is to figure out why the simplicity of a theory should be relevant to saying what the world is like. In this book, the ABCs of probability theory are succinctly developed and put to work to describe two 'parsimony paradigms' within which this problem can be solved.
In their 2013 paper, William Roche and Elliott Sober used Bayesian confirmation theory and the probabilistic concept of screening-off to pose problems for the epistemological theory of Inference to the Best Explanation. Several philosophers replied to those criticisms and Roche and Sober refined and extended their critique in subsequent papers. This article assesses where the debate now stands.
A proposition is derivationally robust precisely when it is a prediction of each model in an ensemble of models. We show that a recent and influential Bayesian defense of the epistemic merits of derivational robustness analysis faces substantial obstacles. The main reason for this is that a Bayesian characterization of derivational robustness requires one to condition on a logical or mathematical truth. Standardly, however, conditioning on a logical or mathematical truth cannot raise or lower the probability of any proposition, which means confirmation in such cases is impossible. What is required for a Bayesian defense of derivational robustness to be plausible is the adoption of a non-standard probability calculus, but the details of such an account have yet to be worked out in sufficient detail. We close by arguing against the view that agreement among multiple models has a confirmatory status that is the same or similar to agreement among multiple measurements of some quantity. It is straightforward to show that agreement among measurements has epistemic significance. This is not true for derivational robustness.
C. Lloyd Morgan's Canon says that a higher mental faculty should not be postulated to explain an organism's behavior if the behavior can be explained by the hypothesis that it has a lower faculty. This epistemological principle strongly influenced and continues to influence cognitive science. The focus of this paper is on Morgan's own argument for the Canon, which philosophers have generally held to be fatally flawed. We disagree. Morgan's good argument for his canon is grounded in Morgan's assumption that an organism that has a higher mental faculty must also have a lower one, together with his belief that a new mental faculty evolves only if it provides a new adaptive advantage. Our main goal here is not to defend Morgan's canon or his argument for it, but to identify the argument's logical location in Morgan's larger framework.