Organochlorine pesticides remain linked to adverse health outcomes, despite the widespread bans and regulations implemented during the 20th century. To better understand these links at a broader scale, researchers have leveraged meta-analyses to quantify pooled mean estimates and assess consistency across multiple primary studies. However, the rapid uptake of meta-analysis has created a diverse and largely fragmented secondary evidence base across many organochlorine pesticides and health outcomes. To consolidate and clarify this evidence base, we conducted a second-order synthesis of 40 meta-analyses, encompassing 129 meta-analytic model estimates. We examined the overall mean effect size and heterogeneity across the included meta-analyses. To ensure comparability, all effect sizes were converted to a common metric (odds ratio), and we used I2 as a common measure of heterogeneity. Our synthesis revealed that, across all included pesticides and adverse health outcomes, organochlorine pesticides increase the odds of an adverse health outcome by an average of 28% in organochlorine pesticide exposed groups compared to unexposed groups (OR = 1.279, 95% confidence interval, hereon CI = [1.16,1.41], 95% prediction interval, hereon PI = [0.660, 2.48]). Specifically, we found that DDE (OR = 1.41, CI = [1.09, 1.83], PI = [1.06, 1.08], number of meta-analyses, hereon (n): 12, number of meta-analytic model estimates, hereon (k): 17), and HCH (OR = 1.43. CI = [1.19, 1.7], PI = [1.07, 1.9]), n = 3, k = 3) exhibited the strongest associations with adverse health outcomes. Endocrine-related diseases showed the highest association with organochlorine pesticide exposure, with an average 52% increase in odds (OR = 1.52, CI = [1.18, 1.95], PI = [1,17, 1.97], n = 13, k = 31). We then revealed that on average the observed heterogeneity in each meta-analysis was moderately high across all outcomes and pesticides (I2within.MA.estimate.average = 54.8%, CI = [37.3, 67.4]). The organochlorine pesticide which exhibited the most consistent impacts was DDT (I2within.MA.estimate.average = 47.7%, CI = [22.0, 65.0]) and the most consistently impacted adverse health outcome were malignant neoplasms (I2within.MA.estimate.average = 19.2%, CI = [0.0, 50.4]. Together, our second-order synthesis quantifies the overall association between exposure to organochlorine pesticides and adverse health outcomes, and the consistency of that association across multiple pesticides and outcomes, providing valuable insights for decision makers and researchers.
Meta-analyses rely on comprehensive reporting of results and data in primary studies. They should therefore also strive for exemplary reporting and follow FAIR principles (Findable, Accessible, Interoperable, Reusable) for data sharing. Yet, despite growing calls for open, reliable, and transparent science, the quality, reporting and FAIRness of meta-analyses in ecology, evolution, and environmental sciences remain poorly characterised. Here, we evaluate features of methodological quality, reporting and data and code sharing practices of 81 meta-analyses published between 2016 and 2020 in ecology, evolution, and environmental sciences. Whilst we found data sharing was relatively common (68%), code sharing was rare (15%). We found low uptake of reporting guidelines (22%), a small proportion of unweighted approaches (17%) and no risk of bias assessments (0%). The FAIRness of data requires urgent attention. Most openly available meta-analytic datasets were Findable and Accessible, but less often Interoperable or Reusable. Low Interoperability and Reusability prevent verification and may explain why data reuse remains uncommon, and ultimately contributes to research waste. We found no evidence of higher FAIRness scores when studies reported having used guidelines, nor that FAIRness associated with citation rates, suggesting a lack of incentives for adopting FAIR principles. Lastly, we provide recommendations for readily implementable methodological improvements applicable to meta-analyses in any field.
Editorial standards are not administrative formalities; they function as scientific quality control mechanisms that directly shape the validity, credibility, integrity, and utility of published research. Since 2016, Environment International has been the first environmental health journal to implement specialist editorial policies for handling systematic review submissions. Over the past decade, Environment International has been committed to the continuous advancement and rigorous editorial standards to ensure publication of trustworthy, high-impact evidence-based research. Central to this effort is the CREST_Triage tool (https://osf.io/bv4en), which enables transparent and rigorous editorial assessment of methodological quality, reporting completeness and reproducibility. In this editorial, we describe the recent developments in our editorial policies and efforts undertaken to further develop and strengthen our standards for evidence syntheses and narrative reviews. A major component of this work has been the expansion of our editorial policies to scoping reviews, review of reviews, their respective protocols, and narrative reviews, which is reflected by the development of new triage instruments, guidance, and workflows, which we have implemented in April 2025. Environment International remains one of the very few journals that actively implement effective quality control measures and enforcement of editorial standards for the evidence syntheses it publishes. We believe that transparent and consistent editorial triage criteria and decisions are beneficial to our authors, peer-reviewers, and the field at large. Amidst the reproducibility crisis in science and increasing concern over the validity of evidence syntheses in environmental health, including those introduced by generative artificial intelligence, we call for stronger editorial stewardship and wider adoption of specialist editorial policies by other journals to increase the quality, transparency and reproducibility of evidence syntheses, leading to more comparable manuscript evaluations across journals. As the evidence synthesis toolkit and practice continues to evolve, our editorial policies and standards will adapt accordingly, ensuring that Environment International remains at the forefront, while upholding the philosophy and principles that have guided us from the beginning.
Meta-analytical models are typically formulated as a mixed-effects model where the sampling variances of the effect sizes are treated as known. In principle, such models could be fitted with standard mixed-modelling software such as the glmmTMB R package. This general-purpose package for generalized linear mixed models (GLMMs) provides flexibility in distributions and random effect covariance structures through the Template Model Builder (TMB). However, incorporating known sampling variances in the conventional inverse-variance formulation of meta-analysis was previously not easily accomplished in glmmTMB. Here, we introduce equalto, a new covariance structure in glmmTMB that allows users to supply a known sampling error variance-covariance matrix when fitting meta-analytic models. This enables explicit modelling of heteroscedasticity and dependence among sampling errors. The new implementation provides an alternative way to fit meta-analytic models, convenient for users already familiar with glmmTMB. Using simulations, we show that the new implementation produces model estimates identical to those from the established metafor package and illustrate its applicability with published meta-analyses in medicine, evolutionary ecology, and the social sciences. Further, this novel implementation in glmmTMB supports more flexible modelling of meta-analytical data, expanding the R toolkit available for evidence synthesis.
Data extraction in systematic reviews, maps, and meta-analyses is time-consuming and prone to human error or subjective judgment. Large Language Models offer the potential for saving time, yet their performance has been evaluated in a limited range of platforms, disciplines, and review types. We assessed the performance of the Elicit platform across diverse data extraction tasks using journal articles from seven systematic reviews in life and environmental sciences. Human-extracted data served as the gold standard. For each review, we used eight articles for prompt development and another eight for testing. Initial prompts were iteratively refined to exceed 87% accuracy or up to five rounds. We then tested extraction accuracy, reproducibility across user accounts, and the effect of Elicit's high-accuracy mode. Of 90 considered prompts, 70 exceeded the 87% accuracy when compared to gold standard, but tended to be lower when tested on a new set of articles. Repeating data extractions with different Elicit user accounts resulted in 90% agreement on extracted values, though supporting quotes and reasoning matched in only 46% and 30% of cases, respectively. In high-accuracy mode, value matches dropped to 77%, with just 10% quote matches and 0% reasoning matches. Extraction accuracy did not differ by data types. Elicit also helped identify eight (<1%) errors in the gold standard data. Our results show that Elicit can complement, but not replace, human data extractors. Elicit may be best used for sanity checks and to evaluate the clarity of data extraction protocols. Prompts must be fine-tuned and independently validated.
Classifying and ranking academic authorship lists is complex in practice, despite existing frameworks, and can lead to conflict. We propose Dragon Kill Points, adapted from multiplayer gaming, to track contributions throughout a project’s lifecycle. Dragon Kill Points is based on five principles: granularity, responsibility, equity, autonomy, and transparency (GREAT). These ensure detailed task records, clear criteria, equitable rules, contributor flexibility, and shared documentation. By applying Dragon Kill Points, teams can reduce disputes, promote inclusivity, and recognise all contributions, including middle authorship. This scalable system offers a practical solution for managing authorship in collaborative research.
Meta-analyses are regarded as the highest level in the hierarchy of evidence, yet standard models traditionally concentrated on estimating the mean effect size, often under restrictive assumptions about the underlying distribution, such as homogeneous variance, symmetric shapes. We introduce a distributional regression framework for meta-analysis that generalizes these conventional models by allowing all parameters of the effect size distribution, such as location, scale, and shape, to be modelled as functions of explanatory variables. This unified framework accommodates a wide range of existing models, including random-effects, multilevel, multivariate, location-scale, and outlier-robust meta-analyses, as special cases. We provide an illustrative example, using 67,393 meta-analyses from the Cochrane Database of Systematic Reviews, employing location-scale models to investigate whether smaller studies tend to report larger effect sizes (i.e., small-study effects) and exhibit greater heterogeneity. We discuss implementation strategies using existing software, considerations for model selection and pre-registration, and the need for further methodological development. By moving beyond the mean effect size, distributional regression enables researchers to explore systematic variation in distributional structure, facilitating the joint test of new hypotheses corresponding to multiple distributional parameters.
The Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guidelines are widely perceived as the gold standard for reporting evidence syntheses. However, the diversity of contributors, transparency of development processes and accessibility of PRISMA checklists have not been systematically examined. We systematically assessed 21 guidelines identified as PRISMA or PRISMA extensions, assessing equity, diversity and inclusion measures; transparency in guideline development; and the implementability and accessibility of their reporting checklists. We found that women were well represented among PRISMA authors (47%). Only 11% of PRISMA authors were affiliated with institutions in the Global South (0.01% excluding China), and 62% of extensions had no contributors from these regions. Transparency varied, most extensions reported following established methodological frameworks (72%) and seeking external feedback (62%), while only 24% provided summaries of consensus meetings and none reported repeatability or inter-rater reliability testing. Accessibility was similarly inconsistent. While 86% provided a reporting checklist, only 10% had openly accessible full texts. Our findings highlight the need for reform in the development of widely adopted reporting standards that underpin evidence synthesis and inform global decision-making. We provide practical recommendations to address these gaps and introduce an open-source reporting template based on current good practices.
Meta-analyses in ecology and evolution typically focus on population means via effect sizes such as the log response ratio. Recently, there has been interest in quantifying effects on variability using the log variability ratio and the log coefficient of variation ratio. Until now, testing for the effects on group means and variabilities has necessitated two separate models. We present a workflow for one integrated meta-analysis of mean and variation effects, or ‘IMAMV’. In a worked example, using data from the diet-mixing literature we show how the focal parameters from IMAMV match those from the equivalent two-model analysis. A common limitation to meta-analysis of variation, is unreported variance values in the primary literature. IMAMV can increase the power to detect effects on variation in meta-analytics datasets with missing variance values through ‘borrowing of strength’. We show, for example, that in a dataset with 20% missing variance values, IMAMV increased the precision of the meta-analytic estimate on the variation effect by 10% compared to the conventional two-model approach. IMAMV can be implemented in commonly used software and requires no additional data beyond that used in the analysis of group means.
Quantitative evidence synthesis method has become a central tool for integration of findings across multiple studies, multi-centre trials, and multi-source cohort data. However, the identification and interpretation of non-replicable, outlying, and influential studies remain insufficiently addressed in practice, despite their potential to substantially affect the robustness and credibility of meta-analytic conclusions. In this paper, we clarify the conceptual distinctions between non-replicability, statistical outlyingness, and study influence, emphasizing that these concepts are related but not interchangeable. We then review the standard principles and procedures of model diagnostics for detecting outlying and influential studies in meta-analysis, together with their underlying statistical rationale. Building on recent methodological developments, we further discuss several practical and methodological refinements, including approaches for handling imprecise and correlated sampling variances, robust diagnostic procedures, and graphical tools for facilitating the identification and interpretation of unusual studies. Finally, we summarize recent advances in outlier and influence diagnostics and provide recommendations for the cautious interpretation and evaluation of studies identified as potentially non-replicable, outlying, or influential within meta-analytic frameworks.
Learned societies, as professional bodies for scientists, are an integral part of the scientific system. However, their membership fees have the potential to be prohibitive to the most vulnerable members of the scientific community. To shed light on how membership fees are structured, we conducted a survey of 182 international learned societies relevant to researchers in ecology and evolution. We found that 83% of these societies offered fee concessions to students, but only 26% to postdoctoral researchers. An average regular membership fee-US$67.8, student fee-US$27.4 (42.7% of the regular fee) and postdoctoral fee-US$42.7 (52.9%). Other types of individual concessions, such as for emeritus, family or unemployed, were rare (2-20%). Of the surveyed societies, 43% had discounts for members from developing countries (Global South). Such discounts were more common among societies located in high-income countries. Societies with a publicly visible commitment to equity, diversity and inclusion were more likely to offer different types of concessions. Currently, fees may prevent researchers from vulnerable and underprivileged groups from accessing multiple professional benefits offered by learned societies in ecology and evolution. This includes postdoctoral researchers, who should receive more support. We recommend tangible actions towards making learned societies more affordable and accessible.
Understanding how both the mean (location) and variance (scale) of traits differ among species and lineages is fundamental to unveiling macroevolutionary patterns. Yet, traditional phylogenetic comparative methods primarily focus on modelling mean trait values, often overlooking variability and heteroscedasticity that can provide critical insights into evolutionary dynamics. Here, we introduce phylogenetic location-scale models (PLSMs), a novel framework that jointly analyses the evolution of trait means and variances. This dual approach captures heteroscedasticity and evolutionary changes in trait variability, allowing for the detection of clades with differing variances and revealing patterns of adaptation, diversification, and evolutionary constraints. Extending PLSMs to a multivariate context enables simultaneous analysis of multiple traits and their covariances, facilitating the testing of hypotheses about evolutionary trade-offs, pleiotropy and phenotypic integration. By modelling covariances between phylogenetic effects in both the location and scale parts, we can discern whether changes in one trait's mean or variance are associated with changes in another's, thereby offering deeper insights into the mechanisms driving trait co-evolution and co-divergence or 'contra-divergence'. We also describe how an extended version of PLSMs incorporating within-species variability can enhance our understanding of trait convergence and divergence arising from ecological and environmental factors. Our framework provides a powerful tool for exploring macroevolutionary patterns and can be used to reassess previously published comparative data, offering new insights into the mechanisms driving the diversity of life.
The production of chemical pesticides poses a critical threat to aquatic ecosystems worldwide, with sub-lethal impacts evident at even relatively low concentrations. Historically, ecotoxicologists have ignored an organism's social context when investigating the effects of pesticide exposure and, instead, have tended to focus on individual-level impacts. Recently, however, there has been a growing interest in understanding the impacts of pesticide exposure on social behaviour. Despite this shift, a holistic understanding of how pesticides impact conspecific interactions (i.e., social behaviour towards individuals of the same species) is lacking due to the multitude of behaviours, pesticides and species currently investigated. In this meta-analysis, we examine the effects of pesticide exposure on conspecific interactions in fish by using data collected from 37 studies on 31 pesticides and 11 species. Our results indicate that pesticide exposure generally reduces the expression of conspecific interactions, but it does not affect the variability of responses between individuals. Courtship behaviour was the most impaired, suggesting that pesticide exposure could weaken how matings are partitioned among individuals in a population. Triazoles and organochlorines were the most impactful pesticide classes for mean differences in behaviour, while triazoles and organophosphates had the greatest effects on response variability. These findings indicate that endocrine-disrupting and neurotoxic pesticides can impact fish conspecific interactions, regardless of their chemical class. Unfortunately, there is a large taxonomic bias in the literature, with most studies using zebrafish as a model, which, in turn, provides scope for studies using a broader range of fish species. We found little statistical evidence of publication biases in our dataset and our results were validated by sensitivity analyses. Overall, our synthesis suggests that pesticides broadly reduce the expression of social behaviours, though effects vary across behaviours, pesticide types, and fish species.
Dispersal is a keystone process shaping ecological and evolutionary dynamics, often assumed to be inherently costly. We synthesized 696 effect sizes from 206 studies across 148 animal species, spanning all continents and ecosystems, to test this assumption. Contrary to long-standing dogma, we found no overall effect of dispersal on fitness (mean effect size: -0.03, 95% CIs: -0.09 to 0.03). No tested biological or methodological moderators explained this variation. Instead, heterogeneity was highest within studies, suggesting that dispersal is highly context-dependent within studies and species. These findings align with game-theoretic expectations that dispersal and philopatry are alternative strategies maintained by balancing or frequency-dependent selection. Our findings are consistent with the view that dispersal involves a balancing act between strategies that yield equivalent long-term payoffs across variable conditions.
Meta-analyses, embedded in systematic reviews, are pivotal in today's scientific landscape for reconciling conflicting findings, increasing statistical power, and charting new research directions. However, poor reporting practices that conceal technical details and potential limitations often need to be revised to maintain their reliability. Despite existing reporting guidelines, a comprehensive tool has yet to be tailored to appraise the reporting quality of the quantitative aspects of meta-analysis in environmental sciences. To bridge this gap, we introduce the Meta-analysis Appraisal Tool for Environmental Sciences (MATES), a checklist of items to assess the reporting quality of meta-analyses. To develop MATES, we used an adapted Delphi process involving workshops (11-16 participants), a survey (193 participants), and validation (30 participants). This process resulted in a 14-item checklist, encompassing the environmental science communities view of important reporting elements. The validation, across 50 meta-analyses, indicated that the tool is repeatable (an average intra-class correlation of 88.97%) and time-efficient (17.00 ± 11.77 min) to implement. To enhance the accessibility and usability of MATES, we created an interactive web-based app that features training and implementation modules https://kylemorrisonisshiny99.shinyapps.io/MATES_shiny/. We also discuss how to interpret the MATES results, potential use cases of MATES and evaluate the development methodology. Overall, MATES provides authors, readers, reviewers, and editors with a reliable and user-friendly tool to assess the reporting quality of meta-analyses in the environmental sciences.
Meta-analyses are powerful synthesis tools that are popular in ecology and evolution owing to the rapidly growing literature of this field. Although the usefulness of meta-analyses depends on their reliability, such as the precision of individual and mean effect sizes, attempts to reproduce meta-analyses' results remain rare in ecology and evolution. Here, we assess the reliability of 41 meta-analyses on sexual signals by evaluating the reproducibility and replicability of their results. We attempted to: (i) reproduce meta-analyses' mean effect sizes using the datasets they provided; (ii) reproduce meta-analyses' effect sizes by re-extracting 5703 effect sizes from 246 primary studies they used as sources; (iii) assess the extent of relevant data missed by original meta-analyses; and (iv) replicate meta-analyses' mean effect sizes after incorporating re-extracted and relevant missing data. We found many discrepancies between meta-analyses' reported results and those generated by our analyses for all reproducibility and replicability attempts. Nonetheless, we argue that the meta-analyses we evaluated are largely reproducible and replicable because the differences we found were small in magnitude, leaving the original interpretation of these meta-analyses' results unchanged. Still, we highlight issues we observed in these meta-analyses that affected their reliability, providing recommendations to ameliorate them.
Rachel Carson's Silent Spring inspired a wave of research on the impacts of organochlorine pesticides, followed by a subsequent wave of meta-analyses. However, the methodological quality and content of these meta-analyses has not been evaluated. Here we systematically map and evaluate the methodological quality of 105 meta-analyses on organochlorine pesticides. We found that 83.4% of the evaluated methodological elements are low quality using the Collaboration for Environmental Evidence Synthesis Assessment Tool (CEESAT v2.1). We then reveal that 227 policy documents cited the included meta-analyses, and there is no difference in methodological quality between those that were cited in policy and those that were not. We also found a paucity of meta-analyses on wildlife despite ample primary evidence. Finally, we quantified the positive impact of using reporting guidelines and we provide recommendations for readily implementable methodological improvements.
Meta-analyses require an effect-size estimate and its corresponding sampling variance from primary studies. In some cases, estimators for the sampling variance of a given effect size statistic may not exist, necessitating the derivation of a new formula for sampling variance. Traditionally, sampling variance formulas are obtained via hand-derived Taylor expansions (the delta method), though this procedure can be challenging for non-statisticians. Building on the idea of single-fit parametric resampling, we introduce SAFE bootstrap: a Single-fit, Accurate, Fast, and Easy simulation recipe that replaces potentially complex algebra with four intuitive steps: fit, draw, transform, and summarise. In a unified framework, the SAFE bootstrap yields bias-corrected point estimates and standard errors for any effect size statistic, regardless of whether the outcome is continuous or discrete. SAFE bootstrapping works by drawing once from a simple sampling model (normal, binomial, etc.), converting each replicate into any effect size of interest and then calculating the bias and sampling variance from simulated data. We demonstrate how to implement the SAFE bootstrap for a simple example first, and then for common effect sizes, such as the standardised mean difference and log odds ratio, as well as for less common effect sizes. With some additional coding, SAFE can also handle zero values and small sample sizes. Our tutorial, with R code supplements, should not only enhance understanding of sampling variance for effect sizes, but also serve as an introduction to the power of simulation-based methods for deriving any effect size with bias correction and its associated sampling variance.