Growing anthropogenic pressures have increased the need for robust predictive models. Meeting this demand requires approaches that can handle bigger data to yield forecasts that capture the variability and underlying uncertainty of ecological systems. Bayesian models are especially adept at this and are growing in use in ecology. Yet many ecologists today are not trained to take advantage of the bigger ecological data needed to generate more flexible robust models. Here we describe a broadly generalizable workflow for statistical analyses and show how it can enhance training in ecology. Building on the increasingly computational toolkit of many ecologists, this approach leverages simulation to integrate model building and testing for empirical data more fully with ecological theory. In turn this workflow can fit models that are more robust and well-suited to provide new ecological insights – allowing us to refine where to put resources for better estimates, better models, and better forecasts.
Shifts in phenology with climate change can lead to asynchrony between interacting species, with cascading impacts on ecosystem services. Previous meta-analyses have produced conflicting results on whether asynchrony has increased in recent decades, but the underlying data have also varied-including in species composition, interaction types and whether studies compared data grouped by trophic level or compared shifts in known interacting species pairs. Here, using updated data from previous studies and a Bayesian phylogenetic model, we found that species have advanced an average of 3.1 days per decade across 1,279 time series across 29 taxonomic classes. We found no evidence that shifts vary by trophic level: shifts were similar when grouped by trophic level, and for species pairs when grouped by their type of interaction-either as paired species known to interact or as randomly paired species. Phenology varied with phylogeny (λ = 0.4), suggesting that uneven sampling of species may affect estimates of phenology and potentially phenological shifts. These results could aid forecasting for well-sampled groups but suggest that climate change has not yet led to widespread increases in phenological asynchrony across interacting species, although substantial biases in current data make forecasting for most groups difficult.
Background Prospective malaria public health interventions are initially tested for entomological impact using standardised experimental hut trials. In some cases, data are collated as aggregated counts of potential outcomes from mosquito feeding attempts given the presence of an insecticidal intervention. Comprehensive data i.e. full breakdowns of probable outcomes of mosquito feeding attempts, are more rarely available. Bayesian evidence synthesis is a framework that explicitly combines data sources to enable the joint estimation of parameters and their uncertainties. The aggregated and comprehensive data can be combined using an evidence synthesis approach to enhance our inference about the potential impact of vector control products across different settings over time. Methods Aggregated and comprehensive data from a meta-analysis of the impact of Pirimiphos-methyl, an indoor residual spray (IRS) product active ingredient, used on wall surfaces to kill mosquitoes and reduce malaria transmission, were analysed using a series of statistical models to understand the benefits and limitations of each. Results Many more data are available in aggregated format ( N = 23 datasets, 4 studies) relative to comprehensive format ( N = 2 datasets, 1 study). The evidence synthesis model had the smallest uncertainty at predicting the probability of mosquitoes dying or surviving and blood-feeding. Generating odds ratios from the correlated Bernoulli random sample indicates that when mortality and blood-feeding are positively correlated, as exhibited in our data, the number of successfully fed mosquitoes will be under-estimated. Analysis of either dataset alone is problematic because aggregated data require an assumption of independence and there are few and variable data in the comprehensive format. Conclusions We developed an approach to combine sources from trials to maximise the inference that can be made from such data and that is applicable to other systems. Bayesian evidence synthesis enables inference from multiple datasets simultaneously to give a more informative result and highlight conflicts between sources. Advantages and limitations of these models are discussed.
Ecological communities change because of both natural and human factors. Distinguishing between the two is critical to ecology and conservation science. One of the most common approaches for modelling species composition changes is calculating beta diversity indices and then relating index changes to environmental changes. The main difficulty with these analyses is that beta diversity indices are paired comparisons, which means indices calculated with the same community are not independent. Mantel tests and generalised dissimilarity modelling (GDM) are two of the most commonly used statistical procedures for analysing such data, employing randomisation tests to consider the data’s dependence. Here, we introduce a Bayesian model-based approach called BetaBayes that explicitly incorporates the data dependence. This approach is based on the Bradley–Terry model, which is a widely used approach for modelling paired comparisons that involves building a standard regression model containing two varying intercepts, one for each community involved in the beta diversity index, that capture their respective contributions. We used BetaBayes to analyse a famous dataset collected in Panama that contains information on multiple 1 ha plots from the rain forests of Panama. We calculated the Bray–Curtis index between all pairs of plots, analysed the relationship between the index and two covariates (geographic distance and elevation), and compared the results of BetaBayes with those from the Mantel test and GDM. BetaBayes has two distinctive features. The first is its flexibility, which allows the user to quickly change it to fit the data structure; namely, by adding varying effects, incorporating spatial autocorrelation, and modelling complex nonlinear relationships. The second is that it provides a clear path for performing model validation and model improvement. BetaBayes avoids hypothesis testing, instead focusing on recreating the data generating process and quantifying all the model configurations that are consistent with the observed data.
Inferences about hypotheses are ubiquitous in the cognitive sciences. Bayes factors provide one general way to compare different hypotheses by their compatibility with the observed data. Those quantifications can then also be used to choose between hypotheses. While Bayes factors provide an immediate approach to hypothesis testing, they are highly sensitive to details of the data/model assumptions. Moreover it's not clear how straightforwardly this approach can be implemented in practice, and in particular how sensitive it is to the details of the computational implementation. Here, we investigate these questions for Bayes factor analyses in the cognitive sciences. We explain the statistics underlying Bayes factors as a tool for Bayesian inferences and discuss that utility functions are needed for principled decisions on hypotheses. Next, we study how Bayes factors misbehave under different conditions. This includes a study of errors in the estimation of Bayes factors. Importantly, it is unknown whether Bayes factor estimates based on bridge sampling are unbiased for complex analyses. We are the first to use simulation-based calibration as a tool to test the accuracy of Bayes factor estimates. Moreover, we study how stable Bayes factors are against different MCMC draws. We moreover study how Bayes factors depend on variation in the data. We also look at variability of decisions based on Bayes factors and how to optimize decisions using a utility function. We outline a Bayes factor workflow that researchers can use to study whether Bayes factors are robust for their individual analysis, and we illustrate this workflow using an example from the cognitive sciences. We hope that this study will provide a workflow to test the strengths and limitations of Bayes factors as a way to quantify evidence in support of scientific hypotheses. Reproducible code is available from https://osf.io/y354c/.
Pairwise comparison data are relatively common in ecology, with beta diversity indices being the most common. Mantel and partial-Mantel tests are the most widely used methods for analysing the relationship between changes in species composition as measured by beta diversity indices and changes in environmental covariates. However, recent studies have shown that these tests can produce invalid results and called for their replacement with more robust methods. In this work, we introduce a novel Bayesian approach for modelling pairwise comparisons that we apply to analyse changes in species composition. To analyse changes in species composition, we usually calculate community similarity indices (e.g. the Sorensen index) and assess the relationship between those indices and environmental covariates. The problem is that community similarity indices are paired comparisons, which means that indices calculated with the same community are not independent. To solve this issue, we followed a model-based approach to fit a regression model of beta diversity indices and covariates that contains two varying intercepts that capture the heterogeneity corresponding to the communities compared by the index. Additionally, our approach allows the relationship between beta diversity indices and covariates to change across data clusters. Moreover, it allows for different types of response variables (continuous or discrete) and provides a clear pathway for model validation. We demonstrate the benefits of our approach using both simulated and actual data on community similarity collected in 338 riparian plant communities. We used Sorensen indices to assess community similarity and analysed the effects of two covariates, network distance and precipitation difference. Our approach provides a robust and verifiable framework for analysing paired comparisons data that can be of particular interest to ecologists and evolutionary biologists, but also to researchers in other areas whenever pairwise comparisons are used.
The distance decay of community similarity (DDCS) is a pattern that is widely observed in terrestrial and aquatic environments. Niche-based theories argue that species are sorted in space according to their ability to adapt to new environmental conditions. The ecological neutral theory argues that community similarity decays due to ecological drift. The continuum hypothesis argues that niche and neutral factors are at the opposite ends of a continuum that ranges from competitive to stochastic exclusion. We assessed the association between niche-based and neutral factors and changes in community similarity measured by Sorensen’s index in riparian plant communities. We considered network distances and flow connection as neutral variables and Strahler order differences and precipitation differences as niche-based variables. We used a hierarchical Bayesian approach to determine which perspective is best supported by the results. We used a high-quality dataset composed of 338 vegetation censuses from eleven river basins in continental Portugal. We observed that changes in Sorensen indices were associated with all covariates but to different degrees. The results suggest that community similarity changes are associated with environmental and neutral factors, supporting the continuum hypothesis.
The distance decay of community similarity (DDCS) is a pattern that is widely observed in terrestrial and aquatic environments. Niche-based theories argue that species are sorted in space according to their ability to adapt to new environmental conditions. The ecological neutral theory argues that community similarity decays due to ecological drift. The continuum hypothesis provides an intermediate perspective between niche-based theories and the neutral theory, arguing that niche and neutral factors are at the opposite ends of a continuum that ranges from competitive to stochastic exclusion. We assessed the association between niche-based and neutral factors and changes in community similarity measured by Sorensen's index in riparian plant communities. We assessed the importance of neutral processes using network distances and flow connection and of niche-based processes using Strahler order differences and precipitation differences. We used a hierarchical Bayesian approach to determine which perspective is best supported by the results. We used dataset composed of 338 vegetation censuses from eleven river basins in continental Portugal. We observed that changes in Sorensen indices were associated with network distance, flow connection, Strahler order difference and precipitation difference but to different degrees. The results suggest that community similarity changes are associated with environmental and neutral factors, supporting the continuum hypothesis.
Experiments in research on memory, language, and in other areas of cognitive science are increasingly being analyzed using Bayesian methods. This has been facilitated by the development of probabilistic programming languages such as Stan, and easily accessible front-end packages such as brms. The utility of Bayesian methods, however, ultimately depends on the relevance of the Bayesian model, in particular whether or not it accurately captures the structure of the data and the data analyst's domain expertise. Even with powerful software, the analyst is responsible for verifying the utility of their model. To demonstrate this point, we introduce a principled Bayesian workflow (Betancourt, 2018) to cognitive science. Using a concrete working example, we describe basic questions one should ask about the model: prior predictive checks, computational faithfulness, model sensitivity, and posterior predictive checks. The running example for demonstrating the workflow is data on reading times with a linguistic manipulation of object versus subject relative clause sentences. This principled Bayesian workflow also demonstrates how to use domain knowledge to inform prior distributions. It provides guidelines and checks for valid data analysis, avoiding overfitting complex models to noise, and capturing relevant data structure in a probabilistic model. Given the increasing use of Bayesian methods, we aim to discuss how these methods can be properly employed to obtain robust answers to scientific questions. All data and code accompanying this article are available from https://osf.io/b2vx9/. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
v.2.17.0 (05 September 2017) New Features New algebraic solver! (stan-dev/stan#2023, #516) append_array now supports vectors of vectors Other C++11 (and some of 14) now supported; see https://github.com/stan-dev/stan/wiki/Supported-C---Compilers-and-Language-Features Updated to Boost 1.64.0 (#599) Makefile refactoring (#602, others) New forward-mode test kit (#0557, #0568) replace copy args with refs (#346)
A common strategy for inference in complex models is the relaxation of a simple model into the more complex target model, for example the prior into the posterior in Bayesian inference. Existing approaches that attempt to generate such transformations, however, are sensitive to the pathologies of complex distributions and can be difficult to implement in practice. Leveraging the geometry of thermodynamic processes I introduce a principled and robust approach to deforming measures that presents a powerful new tool for inference.
Although fundamental to the observable universe, the proton is not elementary. Rather the particle is a bound state of three valence quarks and the QCD vacuum that condenses around them, its properties an amalgamation of those underlying degrees of freedom. Naive expectations presume that contributions from the valence quarks dominate these properties, but the deep inelastic scattering (DIS) experiments which first investigated the proton structure in detail revealed the importance of the vacuum. In particular, polarized DIS measurements uncovered a surprisingly inadequate quark polarization, necessitating significant contributions to the proton spin from elsewhere. The total spin of the gluon field confining the quarks is one possibility, but a contribution only weakly constrained by the electromagnetic probes of DIS. An observable far more sensitive to contributions from the gluon field can be found in the collision of two polarized protons. By correlating the incident proton helicities with final-states originating from an initial-state gluon, the double-helicity asymmetry directly probes the underlying gluon polarization and provides much stronger experimental constraints. Asymmetries measured with hadronic final-states have already improved the understanding of the proton spin structure significantly, but with accumulating statistics these measurements will eventually be limited by systematic uncertainties. Although direct photons are rare in the hadronic environment, the simplicity of the resulting asymmetry ultimately promises a more precise probe of the gluon polarization. Located at the Relativistic Heavy Ion Collider (RHIC), the only facility in the world capable of accelerating and colliding polarized proton beams, the Solenoidal Tracker at RHIC (STAR) provides the large acceptance electromagnetic calorimetry and charged particle tracking critical for measuring direct photons and, subsequently, their asymmetry. Utilizing data from the 2009 running period with intricate simulation, state-of-the-art statistical methods have been developed to tease out the rare photon signal from an overwhelming hadronic background to enable the first direct photon measurements at STAR. This thesis details the construction of the unpolarized cross section and an initial double-helicity asymmetry, proving the feasibility of the direct photon program at the experiment. Thesis Supervisor: Robert Redwine Title: Professor of Physics iii