
This study has been done in cooperation with the automotive supplier Valeo. In automotive industry, client needs evolve quickly in a competitiveness context, particularly, regarding the fan involved in the engine cooling module. The practitioners are asked to propose "optimal" new fans in short times. Unfortunately, each evaluation of the underlying computer code may be expensive whence the need of approximated models and specific, parsimonious, and efficient global optimization strategies. In this paper, we propose to use the Kriging interpolation combined with the expected improvement algorithm to provide new fan designs with high performances in terms of efficiency. As far as we know, such a use of Kriging interpolation together with the expected improvement methodology is unique in an industrial context and provide really promising results.
We study the control of the FamilyWise Error Rate (FWER) in the linear Gaussian model when the n × p design matrix is of rank p. Single step multiple testing procedures controlling the FWER are derived from hyperrectangular confidence regions. In this study, we aim to construct procedures derived from hyperrectangular confidence regions having a minimal volume. We show that minimizing the volume seems a fair criterion to improve the power of the multiple testing procedure. Numerical experiments demonstrate the performance of our approach when compared with the state-of-the-art single step and sequential procedures. We also provide an application to the detection of metabolites in metabolomics.
When clustering the nodes of a graph, a unique partition of the nodes is usually built, either the graph is undirected or directed. While this choice is pertinent for undirected graphs, it should be discussed for directed graphs because it implies that no difference is made between the clusters of source and target nodes. We examine this question in the context of probabilistic models with latent variables and compare the use of the Stochastic Block Model (SBM) and of the Latent Block Model (LBM). We analyze and discuss this comparison through simulated and real data sets and suggest some recommendation.
It is common in instrumental variable studies for instrument values to be missing, for example when the instrument is a genetic test in Mendelian randomization studies. In this paper we discuss two apparent paradoxes that arise in so-called single consent designs where there is one-sided noncompliance, i.e., where unencouraged units cannot access treatment. The first paradox is that, even under a missing completely at random assumption, a complete-case analysis is biased when knowledge of one-sided noncompliance is taken into account; this is not the case when such information is disregarded. This occurs because incorporating information about one-sided noncompliance induces a dependence between the missingness and treatment. The second paradox is that, although incorporating such information does not lead to efficiency gains without missing data, the story is different when instrument values are missing: there, incorporating such information changes the efficiency bound, allowing possible efficiency gains. This is because some of the missing values can be filled in, based on the fact that anyone who received treatment must have been encouraged by the instrument (since the unencouraged cannot access treatment).
Many questions in Data Science are fundamentally causal in that our objective is to learn the effect of some exposure, randomized or not, on an outcome interest. Even studies that are seemingly non-causal, such as those with the goal of prediction or prevalence estimation, have causal elements, including differential censoring or measurement. As a result, we, as Data Scientists, need to consider the underlying causal mechanisms that gave rise to the data, rather than simply the pattern or association observed in those data. In this work, we review the "Causal Roadmap" of Petersen and van der Laan (2014) to provide an introduction to some key concepts in causal inference. Similar to other causal frameworks, the steps of the Roadmap include clearly stating the scientific question, defining of the causal model, translating the scientific question into a causal parameter, assessing the assumptions needed to express the causal parameter as a statistical estimand, implementation of statistical estimators including parametric and semi-parametric methods, and interpretation of our findings. We believe that using such a framework in Data Science will help to ensure that our statistical analyses are guided by the scientific question driving our research, while avoiding over-interpreting our results. We focus on the effect of an exposure occurring at a single time point and highlight the use of targeted maximum likelihood estimation (TMLE) with Super Learner.
Preventive vaccines are an effective public health intervention for reducing the burden of infectious diseases, but have yet to be developed for several major infectious diseases. Vaccine sieve analysis studies whether and how the efficacy of a vaccine varies with the genetics of the infectious pathogen, which may help guide future vaccine development and deployment. A standard statistical approach to sieve analysis compares the effect of the vaccine to prevent infection and disease caused by pathogen types defined dichotomously as genetically near or far from a reference pathogen strain inside the vaccine construct. For example, near may be defined by amino acid identity at all amino acid positions considered in a multiple alignment and far defined by at least one amino acid difference. An alternative approach is to study the efficacy of the vaccine as a function of genetic distance from a pathogen to a reference vaccine strain where the distance cumulates over the set of amino acid positions. We propose a nonparametric method for estimating and testing the trend in the effect of a vaccine across genetic distance. We illustrate the operating characteristics of the estimator via simulation and apply the method to a recent preventive malaria vaccine efficacy trial.
Invited by Gilles Celeux, editor of the Journal de la Société Française de Statistique, to edit a special issue of the journal devoted to causality, we conceived the special issue as a structured collection of articles covering a large spectrum of this exciting topic. Renowned experts of the field contribute nine articles. The special issue unfolds in three parts. Part 1 discusses causality from an epistemological stance. Part 2 discusses the notions of causal models, causal quantities and identifiability. Part 3 discusses the inference of causal quantities. Résumé : Invités par Gilles Celeux, éditeur du Journal de la Société Française de Statistique, à orchestrer la préparation d’un numéro spécial dédié au thème de la causalité, nous l’avons conçu comme une collection structurée d’articles couvrant un large spectre de cet excitant domaine. Des experts de renommée internationale se sont associés au projet et ont offert neuf articles. Le numéro spécial se déploie en trois parties. La première partie discute de la causalité en adoptant un point de vue épistémologique. La seconde partie discute des notions de modèles causaux, de quantités causales et d’identifiabilité. La troisième partie aborde enfin le thème de l’inférence de quantités causales.
For the mathematically wary and unwary alike, Simpson's paradox may well function as a permanent invitation to error. We present Simpson's paradox and discuss its nature based on three examples. It appears that to run afoul of Simpson's paradox it suffices to (a) conflate an invalid probabilistic reasoning with a valid instance of unassailable causal reasoning, or (b) confuse the evidential concept of learning from observation, which for rational agents proceeds by conditioning on the evidence, with the causal concept of acting, represented in causal analysis by the operation of intervening in a causal graph.
Our ambition is to present a gentle introduction to the field of targeted learning. As an example, we consider statistical inference on a simple causal quantity that is ubiquitous in the causal literature. We use this exemplar parameter to introduce key concepts that can be applied to more complicated problems. The introduction weaves together two main threads, one theoretical and the other computational. It also contains exercises. The code is written in the programming language R, which is widely used among statisticians and data scientists to develop statistical software and data analysis. It uses tlrider, a package that we built specifically for this project.
This essay highlights some aspects, core themes and controversies regarding causality from a historical-philosophical perspective with special attention to their role in the AI-data science debate. Firstly, it outlines the contours of this debate and subsequently addresses the aporia of causality in statistics, AI and the philosophy and science. In view of the prevalent crisis some key themes and controversies are identified, and a frame of reference is proposed, that may clarify historical controversies and the current state of "agreeing to disagree" in science and philosophy. Secondly, the essay highlights the historical scope of the concept, outlines some early perspectives and "key moments", that involved main conceptual shifts. Thirdly, the essay outlines the rise of statistics and its role in attempting to defuse the crises by entering a sort of progressing liaison with causality. Finally, it is shown how research in AI has further shaped the concept and how and why causality is about to play a crucial role in the current quest for responsible, explainable and transparent AI and data science.
Targets of inference that establish causality are phrased in terms of counterfactual responses to interventions. These potential outcomes operationalize cause effect relationships by means of comparisons of cases and controls in hypothetical randomized controlled experiments. In many applied settings, data on such experiments is not directly available, necessitating assumptions linking the counterfactual target of inference with the factual observed data distribution. This link is provided by causal models. Originally defined on potential outcomes directly (Rubin, 1976), causal models have been extended to longitudinal settings (Robins, 1986), and reformulated as graphical models (Spirtes et al., 2001; Pearl, 2009). In settings where common causes of all observed variables are themselves observed, many causal inference targets are identified via variations of the expression referred to in the literature as the g-formula (Robins, 1986), the manipulated distribution (Spirtes et al., 2001), or the truncated factorization (Pearl, 2009). In settings where hidden variables are present, identification results become considerably more complicated. In this manuscript, we review identification theory in causal models with hidden variables for common targets that arise in causal inference applications, including causal effects, direct, indirect, and path-specific effects, and outcomes of dynamic treatment regimes. We will describe a simple formulation of this theory (Tian and Pearl, 2002; Shpitser and Pearl, 2006b,a; Tian, 2008; Shpitser, 2013) in terms of causal graphical models, and the fixing operator, a statistical analogue of the intervention operation (Richardson et al., 2017).
Suppose one wishes to estimate the effect of a binary treatment on a binary endpoint conditional on a post-randomization quantity in a counterfactual world in which all subjects received treatment. It is generally difficult to identify this parameter without strong, untestable assumptions. It has been shown that identifiability assumptions become much weaker under a crossover design in which subjects not receiving treatment are later given treatment. Under the assumption that the post-treatment biomarker observed in these crossover subjects is the same as would have been observed had they received treatment at the start of the study, one can identify the treatment effect with only mild additional assumptions. This remains true if the endpoint is absorbent, i.e. an endpoint such as death or HIV infection such that the post-crossover treatment biomarker is not meaningful if the endpoint has already occurred. In this work, we review identifiability results for a parameter of the distribution of the data observed under a crossover design with the principally stratified treatment effect of interest. We describe situations in which these assumptions would be falsifiable, and show that these assumptions are not otherwise falsifiable. We then provide a targeted minimum loss-based estimator for the setting that makes no assumptions on the distribution that generated the data. When the semiparametric efficiency bound is well defined, for which the primary condition is that the biomarker is discrete-valued, this estimator is efficient among all regular and asymptotically linear estimators. We also present a version of this estimator for situations in which the biomarker is continuous. Implications to closeout designs for vaccine trials are discussed.
We consider the estimation of the average treatment effect in the treated as a function of baseline covariates, where there is a valid (conditional) instrument. We describe two doubly robust (DR) estimators: a locally efficient g-estimator, and a targeted minimum loss-based estimator (TMLE). These two DR estimators can be viewed as generalisations of the two-stage least squares (TSLS) method to semi-parametric models that make weaker assumptions. We exploit recent theoretical results that extend to the g-estimator the use of data-adaptive fits for the nuisance parameters. A simulation study is used to compare standard TSLS with the two DR estimators' finite-sample performance, (1) when fitted using parametric nuisance models, and (2) using data-adaptive nuisance fits, obtained from the Super Learner, an ensemble machine learning method. Data-adaptive DR estimators have lower bias and improved coverage, when compared to incorrectly specified parametric DR estimators and TSLS. When the parametric model for the treatment effect curve is correctly specified, the g-estimator outperforms all others, but when this model is misspecified, TMLE performs best, while TSLS can result in large biases and zero coverage. Finally, we illustrate the methods by reanalysing the COPERS (COping with persistent Pain, Effectiveness Research in Self-management) trial to make inference about the causal effect of treatment actually received, and the extent to which this is modified by depression at baseline.
Birge and Massart proposed in 2001 the slope heuristics as a way to choose optimally from data an unknown multiplicative constant in front of a penalty. It is built upon the notion of minimal penalty, and it has been generalized since to some "minimal-penalty algorithms". This article reviews the theoretical results obtained for such algorithms, with a self-contained proof in the simplest framework, precise proof ideas for further generalizations, and a few new results. Explicit connections are made with residual-variance estimators -with an original contribution on this topic, showing that for this task the slope heuristics performs almost as well as a residual-based estimator with the best model choice- and some classical algorithms such as L-curve or elbow heuristics, Mallows' C-p, and Akaike's FPE. Practical issues are also addressed, including two new practical definitions of minimal-penalty algorithms that are compared on synthetic data to previously-proposed definitions. Finally, several conjectures and open problems are suggested as future research directions.
As mentioned in Sylvain’s long and thorough study, I worked with Pascal Massart on the subject of calibration of the penalty terms for model selection but that was actually many years ago. Since then I worked on a somewhat different subject and forgot a large part of this old work. Reading Sylvain’s paper reminded me of a few things and I also learned much from it in particular that a lot of progress has been made on the subject. Although I have been interested by another (but as we shall see not so different) type of problem, I occasionally thought about a particular case of this old stuff, namely complete variable selection. Let me first describe the mathematical framework that I was interested in. One observes n real random variables Y1, . . . ,Yn with the following structure: