Forecasting invasive-pathogen dynamics is paramount to anticipate eradication and containment strategies. Such predictions can be obtained using a model grounded on partial differential equations (PDE; often exploited to model invasions) and fitted to surveillance data. This framework allows the construction of phenomenological but concise models relying on mechanistic hypotheses and real observations. However, it may lead to models with overly rigid behavior and possible data-model mismatches. Hence, to avoid drawing a forecast grounded on a single PDE-based model that would be prone to errors, we propose to apply Bayesian model averaging (BMA), which allows us to account for both parameter and model uncertainties. Thus, we propose a set of different competing PDE-based models for representing the pathogen dynamics, we use an adaptive multiple importance sampling algorithm (AMIS) to estimate parameters of each competing model from surveillance data in a mechanistic-statistical framework, we evaluate the posterior probabilities of models by comparing different approaches proposed in the literature, and we apply BMA to draw posterior distributions of parameters and a posterior forecast of the pathogen dynamics. This approach is applied to predict the extent of Xylella fastidiosa in South Corsica, France, a phytopathogenic bacterium detected in situ in Europe less than 10 years ago (Italy 2013, France 2015). Separating data into training and validation sets, we show that the BMA forecast outperforms competing forecast approaches.
In this paper, we present a case study aimed at determining a billing plan that ensures customer loyalty and provides a profit for the energy company, whose point of view is taken in the paper. The energy provider promotes new contracts for residential buildings, in which customers pay a fixed rate chosen in advance, based on an overall energy consumption forecast. For such a purpose, we consider a practical Bayesian framework for the calibration and validation of a computer code used to forecast the energy consumption of a building. On the basis of power field measurements, collected from an experimental building cell in a given period of time, the code is calibrated, effectively reducing the epistemic uncertainty affecting the most relevant parameters of the code (albedo, thermal bridge factor, and convective coefficient). The validation is carried out by testing the goodness of fit of the code with respect to the field measurements, and then propagating the posterior parametric uncertainty through the code, obtaining probabilistic forecasts of the average electrical power delivered inside the cell in a given period of time. Finally, Bayesian decision-making methods are used to choose the optimal fixed rate (for the energy provider) of the contract, in order to balance short-term benefits with customer retention. We identify three significant contributions of the paper. First of all, the case study data were never analyzed from a Bayesian viewpoint, which is relevant here not only for estimating the parameters but also for properly assessing the uncertainty about the forecasts. Furthermore, the study of optimal policies for energy providers in this framework is new, to the best of our knowledge. Finally, we propose Bayesian posterior predictive p-value for validation.
The formulation of variable selection has been widely developed in the Bayesian literature by linking a random binary indicator to each variable. This Bayesian inference has the advantage of stochastically exploring the set of possible sub-models, whatever their dimension. Bayesian selection approaches, appropriate for categorical predictors, are generally beyond the scope of the standard Bayesian selection of regressors in the linear model since all levels of a categorical variable should be jointly handled in the selection procedure. For categorical covariates, new strategies have been developed to detect the effect of grouped covariates rather than the single effect of a quantitative regressor. In this paper, we review three Bayesian selection methods for categorical predictors: Bayesian Group Lasso with Spike and Slab priors, Bayesian Sparse Group Selection and Bayesian Effect Fusion using model-based clustering. The motivation behind this paper is to provide detailed information about the implementation of the three Bayesian selection methods mentioned above, appropriate for categorical predictors, using the JAGS software. Selection performance and sensitivity analysis of the hyperparameters tuning for prior specifications are assessed under various simulated scenarios. JAGS helps user implement these three Bayesian selection methods for more complex model structures such as hierarchical ones with latent layers.
Assessment of human exposure to atmospheric metals is a challenge, and mosses seem to be good biomonitors to help this purpose. Lacking roots, they are easy to collect and analyze. However, to our knowledge, no formal comparison was made between cadmium (Cd) measurements in Grimmia mosses and alternative forecasts of atmospheric Cd pollution as those produced by the CHIMERE chemistry transport model. This work aims at studying this link to improve further biomonitoring. We compare 128 Cd measurements in the cemetery mosses of Paris and Lyon metropolitan areas (France) to CHIMERE Cd atmospheric forecasts. The area to consider around the cemetery for the CHIMERE forecasts has been defined by Kendall rank correlations between both information sources—Cd in mosses and CHIMERE Cd forecasts—from different area sizes. Then, we fit linear models to those two data sets including step-by-step different sources of uncertainty. Finally, we calculate moss predictions to compare predictions and measurements in the two cities. The results show an apparent link between the Cd concentrations in mosses and CHIMERE Cd forecasts including in addition the same unique covariate, the moss support (grave or wall), in the two cities. However, this model cannot be directly transposed from region to region because the strength of the link appears to be regional.
Les méthodes de capture-marquage-recapture sont des méthodes astucieuses d’échantillonnage peu invasives pour évaluer le nombre d’individus dans une population. Utilisées principalement en écologie, elles trouvent aussi des applications de portée bien plus large dans divers domaines tels que la sociologie et la psychologie expérimentales. Du point de vue de la pédagogie, elles permettent d’illustrer de façon simple, pratique et vivante de nombreux points-clés du raisonnement probabiliste indispensables au statisticien-modélisateur. À l’aide d’une expérience ludique facile à effectuer en salle avec des gommettes, des haricots secs, une cuillère à soupe et un saladier, nous montrons comment aborder de façon simple et intéressante les points-clés suivants dans le cadre d’un problème d’estimation de la taille inconnue d’une population : – les ingrédients de base du problème de statistique inférentielle considéré, en particulier, inconnues versus observables ; – la construction d’un modèle probabiliste/ stochastique possible, fondé sur l’assemblage de plusieurs briques binomiales élémentaires, ainsi que les différentes décompositions possibles de la vraisemblance associée ; – la recherche d’estimateurs, leur étude théorique ainsi que la comparaison de leurs propriétés mathématiques par simulation numérique ; – les différences opérationnelles majeures entre approches statistiques fréquentielle et bayésienne. Cette expérience permet également d’illustrer en quoi le travail d’un statisticien-modélisateur ressemble bien souvent à celui d’un enquêteur de police...
Field experiments are often difficult and expensive to make. To bypass these issues, industrial companies have developed computational codes. These codes intend to be representative of the physical system, but come with a certain amount of problems. The code intends to be as close as possible to the physical system. It turns out that, despite continuous code development, the difference between the code outputs and experiments can remain significant. Two kinds of uncertainties are observed. The first one comes from the difference between the physical phenomenon and the values recorded experimentally. The second concerns the gap between the code and the physical system. To reduce this difference, often named model bias, discrepancy, or model error, computer codes are generally complexified in order to make them more realistic. These improvements lead to time consuming codes. Moreover, a code often depends on parameters to be set by the user to make the code as close as possible to field data. This estimation task is called calibration. This paper proposes a review of Bayesian calibration methods and is based on an application case which makes it possible to discuss the various methodological choices and to illustrate their divergences. This example is based on a code used to predict the power of a photovoltaic plant.
Invasion of new territories by alien organisms is of primary concern for environmental and health agencies and has been a core topic in mathematical modeling, in particular in the intents of reconstructing the past dynamics of the alien organisms and predicting their future spatial extents. Partial differential equations offer a rich and flexible modeling framework that has been applied to a large number of invasions. In this article, we are specifically interested in dating and localizing the introduction that led to an invasion using mathematical modeling, post-introduction data and an adequate statistical inference procedure. We adopt a mechanistic-statistical approach grounded on a coupled reaction–diffusion–absorption model representing the dynamics of an organism in an heterogeneous domain with respect to growth. Initial conditions (including the date and site of the introduction) and model parameters related to diffusion, reproduction and mortality are jointly estimated in the Bayesian framework by using an adaptive importance sampling algorithm. This framework is applied to the invasion of Xylella fastidiosa, a phytopathogenic bacterium detected in South Corsica in 2015, France.
When computer codes are used for modeling complex physical systems, their unknown parameters are tuned by calibration techniques. A discrepancy function may be added to the computer code in order to capture its discrepancy with the real physical process. By considering the validation question of a computer code as a Bayesian selection model problem, Damblin et al. (2016) have highlighted a possible confounding effect in certain configurations between the code discrepancy and a linear computer code by using a Bayesian testing procedure based on the intrinsic Bayes factor. In this paper, we investigate the issue of code error identifiability by applying another Bayesian model selection technique which has been recently developed by Kamary et al. (2014). By embedding the competing models within an encompassing mixture model, Kamary et al. (2014)'s method allows each observation to belong to a different mixing component, providing a more flexible inference, while remaining competitive in terms of computational cost with the intrinsic Bayesian approach. By using the technique of sharing parameters mentioned in Kamary et al. (2014), an improper non-informative prior can be used for some computer code parameters and we demonstrate that the resulting posterior distribution is proper. We then check the sensitivity of our posterior estimates to the choice of the parameter prior distributions. We illustrate that the value of the correlation length of the discrepancy Gaussian process prior impacts the Bayesian inference of the mixture model parameters and that the model discrepancy can be identified by applying the Kamary et al. (2014) method when the correlation length is not too small. Eventually, the proposed method is applied on a hydraulic code in an industrial context.
Air quality over cities is mainly monitored by in-situ surface measurements. However, these stations are too sparse to properly capture the inhomogeneity of pollutant concentrations over urban areas. The need for high-resolution concentration estimate has grown in recent years, together with the awareness of the harmful effects of air pollution. In this study, we develop a Bayesian scheme that combines the high-resolution (3 x 3 m(2)) Particulate Micro SWIFT SPRAY numerical air quality simulations (PMSS) with operational surface measurements. The goal is to improve NOX and PM10 PMSS concentrations estimates over monitoring stations and within their vicinity. For this purpose, we simulate pollutant concentrations over the city of Paris for ten days over the period of March 2016. The Bayesian model provides an enhanced estimate of pollutant concentration in space and time. At the monitoring stations location, these estimates are characterized by lower temporal dispersion compared to the simulated data. Within the vicinity of the monitor stations, enhanced concentration estimates are closer to observations. For NOX, the improvement is stronger and occurs in a larger area for urban background stations than for traffic stations. Overall, NOX improvement is higher than PM10 improvement. The initial PMSS model prediction is more biased for NOX than for PM10 due to large uncertainties in NOX emissions over the traffic network.
Meteorological ensemble members are a collection of scenarios for future weather issued by a meteorological center. Such ensembles nowadays form the main source of valuable information for probabilistic forecasting which aims at producing a predictive probability distribution of the quantity of interest instead of a single best guess point-wise estimate. Unfortunately, ensemble members cannot generally be considered as a sample from such a predictive probability distribution without a preliminary post-processing treatment to re-calibrate the ensemble. Two main families of post-processing methods, either competing such as the BMA or collaborative such as the EMOS, can be found in the literature. This paper proposes a mixed-effect model belonging to the collaborative family. The structure of the model is formally justified by Bruno de Finetti’s representation theorem which shows how to construct operational statistical models of ensemble based on judgments of invariance under the relabeling of the members. Its interesting specificities are as follows: (1) exchangeability contributes to parsimony, with an interpretation of the latent pivot of the ensemble in terms of a statistical synthesis of the essential meteorological features of the ensemble members, (2) a multiensemble implementation is straightforward, allowing to take advantage of various information so as to increase the sharpness of the forecasting procedure. Focus is cast onto normal statistical structures, first with a direct application for temperatures, then with its very convenient Tobit extension for precipitation. Inference is performed by expectation maximization (EM) algorithms with both steps leading to explicit analytic expressions in the Gaussian temperature case, and recourse is made to stochastic conditional simulations in the zero-inflated precipitation case. After checking its good behavior on artificial data, the proposed post-processing technique is applied to temperature and precipitation ensemble forecasts produced for lead times from 1 to 9 days over five river basins managed by Hydro-Québec, which ranks among the world’s largest electric companies. These ensemble forecasts, provided by three meteorological global forecast centers (Canadian, USA and European), were extracted from the THORPEX Interactive Grand Global Ensemble (TIGGE) database. The results indicate that post-processed ensembles are calibrated and generally sharper than the raw ensembles for the five watersheds under study. Supplementary materials accompanying this paper appear on-line.
Risk zoning for mountain hazards remains generally seen as a normative process resulting from transposition of a scheme elaborated for floods. In its hearth is the centennial hazard, a probabilistic reference difficult to properly define, unsuitable for destructive phenomena and little interpretable in terms of exposure. These shortcomings are sources of misunderstandings and they make questionable shortcuts as well as conservative field practices necessary. This article proposes a paradigm shift. Zoning is seen as a compromise between the losses due to the damaging phenomenon and the restrictions that society imposes to itself. The current scientific knowledge does not allow specifying a complete directive procedure, which is anyway not the responsibility of the technical sphere. However, individual risk mapping by combining the hazard model and the damage potential for different elements at risk and then zoning on the basis of acceptability thresholds integrates the multivariate nature of the hazard, allows consideration of uncertainties and authorizes the traceability of the whole decisional procedure. The choice of the individual risk values to which the zone limits correspond and the display of the residual risk after zoning materialize the chosen social compromise, which enables a re-appropriation of the zoning issue by the society. The proposed formal framework is compatible with a wide range of technical solutions and institutional guidelines. Finally, specific recommendations for practice and for future research are formulated. The paper is illustrated by a case study from the field of snow avalanches, but the purpose is transferable to all recurrent rapid mass movements.
Le carbone du sol est important non seulement pour assurer la securite alimentaire en maintenant la fertilite des sols, mais aussi pour limiter le rechauffement climatique en augmentant la sequestration du carbone dans le sol. Il est urgent de comprendre la reaction du carbone du sol face au rechauffement climatique et au changement des pratiques agricoles. Des modeles bio-physiques ont ete developpes depuis quelques decennies pour etudier la matiere organique du sol (SOM). Cependant, il existe encore une forte incertitude sur les mecanismes controlant la dynamique de la SOM, du niveau microbien aux echelles globales. Dans cet article, nous proposons une approche statistique bayesienne de selection de variables pour mieux cerner la dynamique du carbone du sol en examinant la variation en profondeur du radiocarbone pour 159 profils sous differentes conditions de climat (temperature, precipitations, ...) et d'environnement (type de sol, type d'usage du sol, ...). La recherche stochastique de selection de variables (SSVS) est appliquee au niveau des variables latentes d'un modele bayesien hierarchique. Ce modele decrit la variation du radiocarbone en fonction de la profondeur et en tenant compte des covariables explicatives potentielles tels que les facteurs climatiques et environnementaux. Cette approche nous permet d'avoir un jugement probabiliste sur la contribution conjointe du type de sol, du climat et de l'usage du sol a la dynamique verticale du carbone dans le sol. Nous discutons egalement de la performance pratique et des limitations de SSVS en presence de covariables categorielles et de la colinearite entre certaines covariables quand elles interviennent au niveau d'une couche latente d'un modele bayesien hierarchique.
Making good predictions of a physical system using a computer code requires the inputs to be carefully specified. Some of these inputs, called control variables, reproduce physical conditions, whereas other inputs, called parameters, are specific to the computer code and most often uncertain. The goal of statistical calibration consists in reducing their uncertainty with the help of a statistical model which links the code outputs with the field measurements. In a Bayesian setting, the posterior distribution of these parameters is typically sampled using Markov Chain Monte Carlo methods. However, they are impractical when the code runs are highly time-consuming. A way to circumvent this issue consists of replacing the computer code with a Gaussian process emulator, then sampling a surrogate posterior distribution based on it. Doing so, calibration is subject to an error which strongly depends on the numerical design of experiments used to fit the emulator. Under the assumption that there is no code discrepancy, we aim to reduce this error by constructing a sequential design by means of the expected improvement criterion. Numerical illustrations in several dimensions assess the efficiency of such sequential strategies.
Soil carbon is important not only to ensure food security via soil fertility, but also to potentially mitigate global warming via increasing soil carbon sequestration. There is an urgent need to understand the response of the soil carbon pool to climate change and agricultural practices. Biophysical models have been developed to study Soil Organic Matter (SOM) for some decades. However, there still remains considerable uncertainty about the mechanisms that affect SOM dynamics from the microbial level to global scales. In this paper, we propose a statistical Bayesian selection approach to study which forcing conditions influence soil carbon dynamics by looking at the depth distribution of radiocarbon content for 159 profiles under different conditions of climate (temperature, precipitation, etc.) and environment (soil type, land-use). Stochastic Search Variable Selection (SSVS) is here applied to latent variables in a hierarchical Bayesian model. The model describes variations of radiocarbon content as a function of depth and potential covariates such as climatic and environmental factors. SSVS provides a probabilistic judgment about the joint contribution of soil type, climate and land use on soil carbon dynamics. We also discuss the practical performance and limitations of SSVS in presence of categorical covariates and collinearity between covariates in the latent layers of the model.
In this article, we present a recently released R package for Bayesian calibration. Many industrial fields are facing unfeasible or costly field experiments. These experiments are replaced with numerical/computer experiments which are realized by running a numerical code. Bayesian calibration intends to estimate, through a posterior distribution, input parameters of the code in order to make the code outputs close to the available experimental data. The code can be time consuming while the Bayesian calibration implies a lot of code calls which makes studies too burdensome. A discrepancy might also appear between the numerical code and the physical system when facing incompatibility between experimental data and numerical code outputs. The package CaliCo deals with these issues through four statistical models which deal with a time consuming code or not and with discrepancy or not. A guideline for users is provided in order to illustrate the main functions and their arguments. Eventually, a toy example is detailed using CaliCo. This example (based on a real physical system) is in five dimensions and uses simulated data.
Probabilistic forecasting aims at producing a predictive distribution of the quantity of interest instead of a single best guess point-wise estimate. With regard to water flow forecasts, the two main sources of uncertainty stem from unknown future rainfall and temperature (input error, i.e., meteorological uncertainty) and from the inadequacy of the deterministic simulator mimicking the rainfall–runoff (RR) transformation (hydrological uncertainty or RR error). These two sources of uncertainty can be dealt with separately and only the latter will be considered here. Only hydrological uncertainty is at stake when recorded meteorological data (instead of meteorological forecasts) are used as inputs to feed the RR simulator (RRS) for probabilistic predictions. The predictive performance of the RRS may strongly depend on the hydrological regimes: rapid flood variations induce large errors of anticipation but a series of dry events will translate into a much more smoother sequence of river levels due to the easily predictable behavior of the soil reservoir emptying. Consequently, a model with several regimes adapted to different error structures appears as a solution to cope with the issue of unstationary predictive variance. The river regime is modeled as a latent variable, the distribution of which is based on additional outputs of the RRS to be selected. Inference is performed by the EM algorithm with both steps leading to explicit analytic expressions. Asymptotic confidence regions for the estimates are provided within the same EM framework. Model selection is also performed, including the length of the model memory as well as the choice of explanatory variables for the latent regimes. The model is applied to a series of water flow forecasts routinely issued by two hydroelectricity producers in France and in Québec and compared with their present operational forecasting methods.