We compare three graphical methods for displaying evidence in a legal case: Wigmore Charts, Bayesian Networks, and Chain Event Graphs. We find that these methods are aimed at three distinct audiences, respectively, lawyers, forensic scientists and the police. The methods are illustrated using part of the evidence in the case of the murder of Meredith Kercher. More specifically, we focus on representing the list of propositions, evidence, testimony, and facts given in the first trial against Raffaele Sollecito and Amanda Knox with these graphical methodologies.
Because there are similarities between the evaluation of alternative stories in criminal trials and the evaluation of scientific theories, scholars have looked to literature in epistemology and the philosophy of science for insights into the evaluation of evidence in criminal trials. The philosophical literature is divided, however, on a key point-the epistemic value of novel predictive success. This article uses a Bayesian network analysis to explore, in the context of a criminal case, the circumstances in which "new evidence" discovered after a theory is propounded can provide stronger (or weaker) support for the theory than "old evidence" that was accommodated by the theory. It argues that insights from analysis of the strength of evidence in the criminal case can be applied more generally when assessing the relative merits of prediction and accommodation in scientific theory development, and are thus helpful in addressing the longstanding philosophic controversy over this issue.
We compare three graphical methods for displaying evidence in a legal case: Wigmore Charts, Bayesian Networks and Chain Event Graphs. We find that these methods are aimed at three distinct audiences, respectively lawyers, forensic scientists and the police. The methods are illustrated using part of the evidence in the case of the murder of Meredith Kercher. More specifically, we focus on representing the list of propositions, evidence, testimony and facts given in the first trial against Raffaele Sollecito and Amanda Knox with these graphical methodologies.
Justice systems are sometimes called upon to evaluate cases in which health-care professionals are suspected of killing their patients illegally. These cases are difficult to evaluate because they involve at least two levels of uncertainty. Commonly in a murder case it is clear that a homicide has occurred, and investigators must resolve uncertainty about who is responsible. In the cases we examine here there is also uncertainty about whether homicide has occurred. Investigators need to consider whether the deaths that prompted the investigation could plausibly have occurred for reasons other than homicide, in addition to considering whether, if homicide was indeed the cause, the person under suspicion is responsible. In this report, prepared under the auspices of the Royal Statistical Society, we provide advice and guidance on the investigation and evaluation of such cases. Our work was prompted by concerns about the statistical challenges such cases pose for the legal system.
Suspicions about medical murder sometimes arise due to a surprising or unexpected series of events, such as an apparently unusual number of deaths among patients under the care of a particular nurse. But also a single disturbing event might trigger suspicion about a particular nurse, and this might then lead to investigation of events which happened when she was thought to be present. In either case, there is a statistical challenge of distinguishing event clusters that arise from criminal acts from those that arise coincidentally from other causes. We show that an apparently striking association between a nurse's presence and a high rate of deaths in a hospital ward can easily be completely spurious. In short: in a medium-care hospital ward where many patients are suffering terminal illnesses, and deaths are frequent, most deaths occur in the morning. Most nurses are on duty in the morning, too. There are less deaths in the afternoon, and even less at night; correspondingly, less nurses are on duty in the afternoon, even less during the night. Consequently, a full time nurse works the most hours when the most deaths occur. The death rate is higher when she is present than when she is absent.
We give an overview of various topics tied to the expression of uncertainty about a variable or event by means of a probability distribution. We first consider methods used to evaluate a single probability forecaster, including scoring rules, calibration, resolution and refinement. We next revisit methods for combining several experts’ distributions, including the linear and logarithmic opinion pools. We discuss the model-based approach and the axiomatic approach to opinion pooling and describe the implications of imposing coherence constraints, based on a specific understanding of “expertise”. In the final part, we revisit from a statistical standpoint some results in the economics literature on prediction markets, where individuals sequentially place bets on the outcome of a future event. This leaves a trail of personal probabilities for the event, each conditional on the current individual’s private background knowledge and on the previously announced probabilities of other individuals. In particular, we consider the case of two individuals who start with the same probability distribution but have different private information and take turns in updating their probabilities. We note convergence of the announced probabilities to a limiting value, which may or may not be the same as that based on pooling their private information.
In both criminal cases and civil cases, there is an increasing demand for the analysis of DNA mixtures involving relationships. The goal might be, for example, to identify the contributors to a DNA mixture where the donors may be related, or to infer the relationship between individuals based on a mixture. This paper introduces an approach to modelling and computation for DNA mixtures involving contributors with arbitrarily complex relationships. It builds on an extension of Jacquard's condensed coefficients of identity, to specify and compute with joint relationships, not only pairwise ones, including the possibility of inbreeding. The methodology developed is applied to two casework examples involving a missing person, and simulation studies of performance, in which the ability of the methodology to recover complex relationship information from synthetic data with known ‘true’ family structure is examined. The methods used to analyse the examples are implemented in the new KinMix R package that extends the DNAmixtures package to allow for modelling DNA mixtures with related contributors.
In both criminal cases and civil cases there is an increasing demand for the analysis of DNA mixtures involving relationships. The goal might be, for example, to identify the contributors to a DNA mixture where the unknown donors may be related, or to infer the relationship between individuals based on a DNA mixture. This paper applies a new approach to modelling and computation for DNA mixtures involving contributors with arbitrarily complex relationships to two real cases from the Spanish Forensic Police.
Here, we illustrate how statistical methods can help extract information from mixed DNA profiles pertaining to an Italian case, referred to by the media as The murder of Yara Gambirasio. We base the analysis on a model for DNA mixtures that takes fully into account the peak heights and possible artefacts, like stutter and dropout that might occur in the DNA amplification process. We show how to combine the evidence from multiple samples and from different marker systems all within the model framework. The combined evidence is used for deconvolution, where the focus is to find likely profiles for the donors to the sample. We also show how a mixture can be used to establish familial relationships between a reference profile and a donor to the mixed DNA sample. We compare results based on a single mixed DNA profile, combination of replicates, combinations of different samples and combinations of different kits. Based on the Yara case, we discuss just a few of the plethora of possibilities of combining evidential information.
Here we present an Italian criminal case that shows how statistical methods can be used to extract information from a series of mixed DNA profiles. The case involves several different individuals and a set of different DNA traces. The case possibly involves persons of interest of a small population of Romani origin. First, a brief description of the case is provided. Secondly, we introduce some heuristic tools that can be used to evaluate the data and we also briefly outline the statistical model used for analysing DNA mixtures. Finally, we illustrate some of the findings on the case and discuss further directions of research. The results show how the use of different population database allele frequencies for analysing the DNA mixtures can lead to very different results, some seemingly inculpatory and some seemingly exculpatory. We also illustrate the results obtained from combining the evidence from different samples.
Forensic science has experienced a period of rapid change because of the tremendous evolution in DNA profiling. Problems of forensic identification from DNA evidence can become extremely challenging, both logically and computationally, in the presence of complicating features, such as in mixed DNA trace evidence. Additional complicating aspects are possible, such as missing data on individuals, heterogeneous populations, and kinship. In such cases, there is considerable uncertainty involved in determining whether or not the DNA of a given individual is actually present in the sample. We begin by giving a brief introduction to the genetic background needed for understanding forensic DNA mixtures, including the artifacts that commonly occur in the DNA amplification process. We then review different methods and software based on qualitative and quantitative information and give details on a quantitative method that uses Bayesian networks as a computational device for efficiently computing likelihoods. This method allows for the possibility of combining evidence from multiple samples to make inference about relationships from DNA mixtures and other more complex scenarios.
Bradley (Theory Decis 85:5–20, 2018) develops some theory of the linear opinion pool, in apparent contradiction to results of Dawid et al. (Test 4:263–314, 1995). We investigate the sources of these contradictions, and in particular identify a mathematical error in Bradley (2018) that invalidates his main result.
This article discusses the use of Bayesian analysis in the evaluation of temporal volatility and information flows in political campaigns. Using the 2004 US presidential election campaign as a case study, it demonstrates the utility of a model with two volatility regimes that simplifies the task of associating events with periods of high information. The article first explains why prediction markets are able to aggregate information such that the prices of future contracts are reflective of the event’s actual probability of occurring before analysing data from futures on ‘Bush wins the popular vote in 2004’, or the traded probability, of Bush winning the election. These data are used to build a measure of information flow. The results show that information flows increased as a result of the televised debates, and that these debates, along with the selection of the vice presidential candidate, increased prediction market volatility.
We present methods for inference about relationships between contributors to a DNA mixture and other individuals of known genotype: a basic example would be testing whether a contributor to a mixture is the father of a child of known genotype. The evidence for such a relationship is evaluated as the likelihood ratio for the specified relationship versus the alternative that there is no relationship. We analyse real casework examples from a criminal case and a disputed paternity case; in both examples part of the evidence was from a DNA mixture. DNA samples are of varying quality and therefore present challenging problems in interpretation. Our methods are based on a recent statistical model for DNA mixtures, in which a Bayesian network (BN) is used as a computational device; the present work builds on that approach, but makes more explicit use of the BN in the modelling. The R code for the analyses presented is freely available as supplementary material. We show how additional information of specific genotypes relevant to the relationship under analysis greatly strengthens the resulting inference. We find that taking full account of the uncertainty inherent in a DNA mixture can yield likelihood ratios very close to what one would obtain if we had a single source DNA profile. Furthermore, the methods can be readily extended to analyse different scenarios as our methods are not limited to the particular genotyping kits used in the examples, to the allele frequency databases used, to the numbers of contributors assumed, to the number of traces analysed simultaneously, nor to the specific hypotheses tested.
In a prediction market, individuals can sequentially place bets on the outcome of a future event. This leaves a trail of personal probabilities for the event, each being conditional on the current individual's private background knowledge and on the previously announced probabilities of other individuals, which give partial information about their private knowledge. By means of theory and examples, we revisit some results in this area. In particular, we consider the case of two individuals, who start with the same overall probability distribution but different private information, and then take turns in updating their probabilities. We note convergence of the announced probabilities to a limiting value, which may or may not be the same as that based on pooling their private information.
A statistical model for the quantitative peak information obtained from a forensic DNA mixture sample is illustrated on a real case example. We use the combined information from two DNA traces: to find likelihood ratios to quantify the strength of evidence; to deconvolve the mixtures for the purpose of finding likely profiles of unknown contributors to the traces; and to analyse the artefacts that might be present in the mixture after DNA amplification.
Here we analyse a complex disputed paternity case, where the DNA of the putative father was extracted from his corpse that had been inhumed for over 20 years. This DNA was contaminated and appears to be a mixture of at least two individuals. Furthermore, the mother's DNA was not available. The DNA mixture was analysed so as to predict the most probable genotypes of each contributor. The major contributor's profile was then used to compute the likelihood ratio for paternity. We also show how to take into account a dropout allele and the possibility of mutation in paternity testing.