Structured elicitation protocols, such as the IDEA protocol, are used to elicit probabilistic judgements from multiple domain experts about uncertain events across fields including ecology, biosecurity risk assessment, and metascience. Individual expert judgements must subsequently be mathematically aggregated into a single group forecast. While the simplest case involves combining a set of point-estimates from multiple individuals, this process is further complicated when judgements include uncertainty bounds, or when elicitation is conducted across multiple rounds. This paper presents aggreCAT, an open-source R package that provides 29 aggregation methods for combining individual expert judgements into a single probabilistic estimate, accommodating designs ranging from single-round point estimates to multi-round three-point elicitation. The package follows tidy data principles, enabling straightforward integration with existing R workflows for application at scale. Methods range from unweighted arithmetic combinations to performance-weighted schemes and Bayesian models, with weights derived from uncertainty intervals, shifts in judgements between elicitation rounds, and breadth of expert reasoning. We provide worked examples illustrating the mechanics of representative aggregation methods, a general workflow for batch aggregation across multiple forecasts and methods, and built-in functions for evaluating and visualising forecast performance against known outcomes. aggreCAT fills a substantive gap in open software for mathematically aggregating expert judgement, and is intended to support researchers and decision analysts in rapidly and rigorously synthesising outputs from structured elicitation exercises.
Abstract Systematic reviews and meta‐analyses are key evidence synthesis methods for informing future research, interventions and policy. As the validity of their conclusions depends on the primary studies they synthesise, assessing the internal validity of the included studies is essential. In some fields, such as medicine, this is the norm and is commonly done using Risk of Bias assessment tools. Risk of Bias (RoB) assessment is however rare in ecology and evolutionary biology (EEB) even though several RoB tools have been developed in some ecological subfields and related fields. To identify potential reasons for a limited uptake, we conducted a survey of ecologists and evolutionary biologists with evidence synthesis experience and reviewed 275 journals that publish EEB research for guidelines on performing RoB or related assessments. Only 28 of 232 (12%) survey respondents had correct interpretation of the RoB concept, while 46 (40%) of 116 that had heard of RoB have confused RoB with publication bias. Just 10 (4%) had conducted a RoB assessment, most of whom found it challenging. Out of the 209 EEB journals that explicitly solicit evidence synthesis (N = 58) or reviews (N = 151), only five (2%) directly mentioned standards for conducting evidence synthesis, which include RoB or related (e.g. critical appraisal) assessments. An additional 45 (22%) journals indirectly linked to RoB or a related assessment via referring to the guidelines for reporting evidence synthesis (e.g. PRISMA), despite such reporting guidelines not providing information on how to conduct RoB assessments. To increase its uptake in EEB we recommend making RoB assessment: (1) known and recognised as an essential component of a reliable evidence synthesis by including it in training materials, and journals' and funders' guidelines and policies; (2) easy to perform by bringing the synthesis community together to determine the need for developing new or adjusting existing RoB tools; and (3) possible by further improving reporting standards for primary studies so that RoB assessment can be done on these studies. For those unfamiliar with the RoB assessment, we provide five key RoB questions that existing tools often cover. These questions can be considered to understand the basic composition of the evidence included in evidence synthesis.
As a community of researchers, behavioral ecologists work on a staggering diversity of animal systems and questions. Our goal is to build a series of generalizable hypotheses that can explain how and why animals do the things they do. But in order to fully delineate the boundaries and limitations of our hypotheses we need to replicate more. Most researchers generally agree that proper replication is necessary to ensure a firm foundation for our work, but lament that these studies can be hard to fund and even harder to publish. Here, we present a potential way forward: incorporating replication studies as part of standard research training for undergraduates and starting graduate students. We outline how these replication studies can be a positive pedagogical tool, providing training in research design, analysis and interpretation for early career researchers. Then we also introduce a new section of this journal, Replication Studies, which will publish close replications in short-format. Publishing more replication studies will shore up the foundations of our research and allow future evidence syntheses to uncover general patterns of animal behavior.
This paper reports two approaches to forecasting replicability for a corpus of 3000 published social science papers, using large-scale human assessments. Replication Markets used incentivized surveys and prediction markets where participants traded assets linked to replication outcomes and market prices were interpreted as replication predictions. The repliCATS project used a structured group deliberation protocol, including interactive discussion, and mathematically aggregated forecasts to generate replication predictions. Accuracy for both approaches was validated against a subset of (n=37) independent, high-power replication studies. The predictive accuracy (median [range]) achieved for AUC by Replication Markets was 0.76 [0.72-0.82], and 0.76 [0.68-0.81] by repliCATS. Replication Markets achieved a classification accuracy of 73% [68-76%] and repliCATS achieved 68% [59-78%]. These results place the performance of both teams within the accuracy range achieved in prior replication forecasting studies. We conclude that informative forecasts can be elicited by both methods, but there are trade-offs between scale and accuracy.
Meta-analyses in ecology and evolution sometimes target the magnitude of differences between groups rather than their direction, for example, when the question concerns deviation from a biological optimum or divergence from a reference state. A common practice is to convert signed effects into magnitudes by taking absolute values, but this can induce upward distortion and non-normal sampling distributions under standard meta-analytic models. Here we introduce lnM, a log-ratio effect size for the magnitude of difference between two groups, defined from one-way ANOVA components. lnM is asymptotically normal, supports standard multilevel meta-analysis and meta-regression and applies to both ratio- and interval-scale traits. Using theory, simulations and worked examples, we show when delta-method and single-fit bootstrap estimators perform well, and how lnM can be used to assess publication bias.
Meta-analyses, embedded in systematic reviews, are pivotal in today's scientific landscape for reconciling conflicting findings, increasing statistical power, and charting new research directions. However, poor reporting practices that conceal technical details and potential limitations often need to be revised to maintain their reliability. Despite existing reporting guidelines, a comprehensive tool has yet to be tailored to appraise the reporting quality of the quantitative aspects of meta-analysis in environmental sciences. To bridge this gap, we introduce the Meta-analysis Appraisal Tool for Environmental Sciences (MATES), a checklist of items to assess the reporting quality of meta-analyses. To develop MATES, we used an adapted Delphi process involving workshops (11-16 participants), a survey (193 participants), and validation (30 participants). This process resulted in a 14-item checklist, encompassing the environmental science communities view of important reporting elements. The validation, across 50 meta-analyses, indicated that the tool is repeatable (an average intra-class correlation of 88.97%) and time-efficient (17.00 ± 11.77 min) to implement. To enhance the accessibility and usability of MATES, we created an interactive web-based app that features training and implementation modules https://kylemorrisonisshiny99.shinyapps.io/MATES_shiny/. We also discuss how to interpret the MATES results, potential use cases of MATES and evaluate the development methodology. Overall, MATES provides authors, readers, reviewers, and editors with a reliable and user-friendly tool to assess the reporting quality of meta-analyses in the environmental sciences.
Populations must continuously respond to environmental change or risk extinction. These responses can be measured as phenotypic rates of change, which allow researchers to predict their contemporary evolutionary responses. In 1999, a database of phenotypic rates of change in wild populations was compiled. Since then, researchers have used (and expanded) this database to examine the phenotypic responses as a function of the features of the study system (i.e., the population or set of populations, of a given species, that experienced a specific driver or disturbance), the measured traits, and methodological approaches. Therefore, PROCEED (Phenotypic Rates of Change Evolutionary and Ecological Database) is an ongoing compilation of rates of phenotypic change, typically calculated as Haldanes and Darwins, published in peer-reviewed literature (but also including data from theses and technical reports). Studies in this database measure the intraspecific change in quantitative (continuous or discrete) traits and report either the time elapsed from the onset of environmental novelty, or reference a historical or biological event reported in other sources (e.g., a mine opening or a well-documented biological invasion). Included studies either follow a single population through time (allochronic design) or compare two or more populations that diverged at a known time (synchronic design). Some included studies account for the total phenotypic variability in the field (i.e., phenotypic studies), while others employed common-garden or other quantitative genetic approaches to account for the heritable component of the phenotypic change (i.e., genetic studies). PROCEED includes systems in both natural and experimental conditions, provided that reproduction was not manipulated (i.e., artificial selection experiments were excluded). In the included experimental systems, the environment of the focal populations was manipulated (e.g., an herbivory exclusion experiment, where the type and load of herbivory are manipulated) but the studies did not deliberately select for trait values in the study population (e.g., the plant height). PROCEED does not include systems where the phenotypic change is presumably due to interspecific hybridization, polyploidy, or other chromosomal alterations. Here, we present the most recently updated PROCEED (Version 6.1). This new, curated version has 9263 records (n) collated from 326 studies, 1801 systems, and 428 species. The database includes records belonging to mammals (n = 686), birds (n = 1475), reptiles (n = 96), amphibians (n = 23), fishes (n = 3671), invertebrates (n = 1141, mostly arthropods), and plants (n = 2171). The maximum elapsed time between the environmental change and the sampling is 500 years but is typically less than 100 years (third quartile 89.5; median 45 years). The database also includes a set of variables describing biological and methodological aspects of the study system and measured traits, along with features of the sampling design in the primary source of information. This new version of PROCEED also includes a time series dataset comprising a subset of records included in the general dataset. These are allochronic studies with three or more sampling times throughout the entire study period. The time series dataset contains 655 time series (s)-belonging to 61 studies, from 156 systems, and 77 species-including mammals (s = 140), birds (s = 77), reptiles (s = 4), amphibians (s = 8), fishes (s = 404), and plants (s = 22). The data are released under a Creative Commons CC0 1.0 Universal Public Domain Dedication license.
Data and code are essential for ensuring the credibility of scientific results and facilitating reproducibility, areas in which journal sharing policies play a crucial role. However, in ecology and evolution, we still do not know how widespread data- and code-sharing policies are, how accessible they are, and whether journals support data and code peer review. Here, we first assessed the clarity, strictness and timing of data- and code-sharing policies across 275 journals in ecology and evolution. Second, we assessed initial compliance to journal policies using submissions from two journals: Proceedings of the Royal Society B (Mar 2023-Feb 2024: n = 2340) and Ecology Letters (Jun 2021-Nov 2023: n = 571). Our results indicate the need for improvement: across 275 journals, 22.5% encouraged and 38.2% mandated data-sharing, while 26.6% encouraged and 26.9% mandated code-sharing. Journals that mandated data- or code-sharing typically required it for peer review (59.0% and 77.0%, respectively), which decreased when journals only encouraged sharing (40.3% and 24.7%, respectively). Our evaluation of policy compliance confirmed the important role of journals in increasing data- and code-sharing but also indicated the need for meaningful changes to enhance reproducibility. We provide seven recommendations to help improve data- and code-sharing, and policy compliance.
Publishing preprints is quickly becoming commonplace in ecology and evolutionary biology. Preprints can facilitate the rapid sharing of scientific knowledge establishing precedence and enabling feedback from the research community before peer review. Yet, significant barriers to preprint use exist, including language barriers, a lack of understanding about the benefits of preprints and a lack of diversity in the types of research outputs accepted (e.g. reports). Community-driven preprint initiatives can allow a research community to come together to break down these barriers to improve equity and coverage of global knowledge. Here, we explore the first preprints uploaded to EcoEvoRxiv (n = 1216), a community-driven preprint server for ecologists and evolutionary biologists, to characterize preprint use in ecology, evolution and conservation. Our perspective piece highlights some of the unique initiatives that EcoEvoRxiv has taken to break down barriers to scientific publishing by exploring the composition of articles, how gender and career stage influence preprint use, whether preprints are associated with greater open science practices (e.g. code and data sharing) and tracking preprint publication outcomes. Our analysis identifies areas that we still need to improve upon but highlights how community-driven initiatives, such as EcoEvoRxiv, can play a crucial role in shaping publishing practices in biology.
Meta-analysis is commonly a core component of systematic reviews and has become an important method to reconcile conflicting findings, increase statistical power, and chart new research directions. However, poor reporting practices make it challenging to evaluate the validity of meta-analyses. Despite the existence of reporting checklists, a specifically designed tool has yet to be developed to appraise the completeness with which a meta-analysis has been reported. To bridge this gap, we introduce the Meta-analysis Appraisal Tool for Environmental Sciences (MATES). To develop MATES, we adapted a Delphi process involving experts in meta-analysis methodologies, researchers with experience in guideline/appraisal tool development, and editors of relevant journals. The Delphi process had five steps, including three workshops (11-16 participants), a survey (193 participants), and a validation task (30 participants). This iterative development process resulted in a 14-item appraisal tool that reflects the environmental science and research syntheses community's consensus on essential elements to appraise the completeness with which a meta-analysis has been reported. Validation across 50 meta-analyses demonstrated that the tool is repeatable (average agreement rate: 88.97 %) and time-efficient to implement (17.00 ± 12.23 min). We also outline guidance for interpreting MATES results, describe its potential applications, and reflect on the development process. The authors provide practical implementation guidance for each MATES item, illustrated with real examples in the supplementary material. We also report an extended development methodology to support reproducibility. Finally, we built created a ShinyApp that includes both a training module and an application tool to enhance the usability of MATES (https://kylemorrisonisshiny99.shinyapps.io/MATES_shiny/). Overall, MATES provides authors, readers, stakeholders, and editors with a reliable and accessible tool for appraising the completeness with which a meta-analysis in environmental sciences has been reported.
Abstract Plant secondary metabolites (PSMs) are produced by plants to overcome environmental challenges, both biotic and abiotic. We were interested in characterizing how autumn seasonality in temperate and subtropical climates affects overall PSM production in comparison to herbivory. Herbivory is commonly measured between spring to summer when plants have high resource availability and prioritize growth and reproduction. However, autumn seasonality also challenges plants as they cope with limited resources and prepare survival for winter. This suggests a potential gap in our understanding of how herbivory affects PSM production in autumn compared to spring/summer. Using meta‐analysis, we recorded overall production of 22 different PSM subgroups from 58 published papers to calculate effect sizes from herbivory studies (absence to presence) and temperate to subtropical seasonal studies (summer to autumn), while considering other variables (e.g., plant type, increase in time since herbivory, temperature, and precipitation). We also compared production of five phenolic PSM subgroups – hydroxybenzoic acids, flavan‐3‐ols, flavonols, hydrolysable tannins, and condensed tannins. We wanted to detect a shared response across all PSMs and found that herbivory increased overall PSM production in herbaceous plants. Herbivory was also found to have a positive effect on individual PSM subgroups, such as flavonol production, while autumn seasonality was found to have a positive effect on flavan‐3‐ol and condensed tannin production. We discuss how these responses might stem from plants producing some PSMs constitutively, whereas others are induced only after herbivory, and how plants produce metabolites with higher costs only during seasons when other resources for growth and reproduction are less available, while other phenolic PSM subgroups serve more than one function for plants and such functions can be season dependent. The outcome of our meta‐analysis is that autumn seasonality changes some PSM production differently from herbivory, and we see value in further investigating seasonality–herbivory interactions with plant chemical defense.
When researchers collaboratively tackle challenging questions, such as those related to climate change, the impact of artificial intelligence, or the nature of consciousness, they often encounter disagreements that are difficult to resolve (1). When disagreements persist, what should the team members do? Options may seem limited to co-authoring a paper they disagree with, delaying in hopes that consensus will eventually be reached, or leaving without credit. However, there is a better approach: transparently documenting the disagreements.
Contributor Roles Taxonomy (CRediT) has recently changed how author contributions are acknowledged. To extend and complement CRediT, we propose MeRIT, a new way of writing the Methods section using the author’s initials to further clarify contributor roles for reproducibility and replicability. Lack of information on authors’ contribution to specific aspects of a study hampers reproducibility and replicability. Here, the authors propose a new, easily implemented reporting system to clarify contributor roles in the Methods section of an article.
Collaborative efforts to directly replicate empirical studies in the medical and social sciences have revealed alarmingly low rates of replicability, a phenomenon dubbed the 'replication crisis'. Poor replicability has spurred cultural changes targeted at improving reliability in these disciplines. Given the absence of equivalent replication projects in ecology and evolutionary biology, two inter-related indicators offer the opportunity to retrospectively assess replicability: publication bias and statistical power. This registered report assesses the prevalence and severity of small-study (i.e., smaller studies reporting larger effect sizes) and decline effects (i.e., effect sizes decreasing over time) across ecology and evolutionary biology using 87 meta-analyses comprising 4,250 primary studies and 17,638 effect sizes. Further, we estimate how publication bias might distort the estimation of effect sizes, statistical power, and errors in magnitude (Type M or exaggeration ratio) and sign (Type S). We show strong evidence for the pervasiveness of both small-study and decline effects in ecology and evolution. There was widespread prevalence of publication bias that resulted in meta-analytic means being over-estimated by (at least) 0.12 standard deviations. The prevalence of publication bias distorted confidence in meta-analytic results, with 66% of initially statistically significant meta-analytic means becoming non-significant after correcting for publication bias. Ecological and evolutionary studies consistently had low statistical power (15%) with a 4-fold exaggeration of effects on average (Type M error rates = 4.4). Notably, publication bias reduced power from 23% to 15% and increased type M error rates from 2.7 to 4.4 because it creates a non-random sample of effect size evidence. The sign errors of effect sizes (Type S error) increased from 5% to 8% because of publication bias. Our research provides clear evidence that many published ecological and evolutionary findings are inflated. Our results highlight the importance of designing high-power empirical studies (e.g., via collaborative team science), promoting and encouraging replication studies, testing and correcting for publication bias in meta-analyses, and adopting open and transparent research practices, such as (pre)registration, data- and code-sharing, and transparent reporting.
Most studies assessing rates of phenotypic change focus on population mean trait values, whereas a largely overlooked additional component is changes in population trait variation. Theoretically, eco-evolutionary dynamics mediated by such changes in trait variation could be as important as those mediated by changes in trait means. To date, however, no study has comprehensively summarised how phenotypic variation is changing in contemporary populations. Here, we explore four questions using a large database: How do changes in trait variances compare to changes in trait means? Do different human disturbances have different effects on trait variance? Do different trait types have different effects on changes in trait variance? Do studies that established a genetic basis for trait change show different patterns from those that did not? We find that changes in variation are typically small; yet we also see some very large changes associated with particular disturbances or trait types. We close by interpreting and discussing the implications of our findings in the context of eco-evolutionary studies.
1. Although meta-analysis has become an essential tool in ecology and evolution, reporting of meta-analytic results can still be much improved. To aid this, we have introduced the orchard plot, which presents not only overall estimates and their confidence intervals but also shows corresponding heterogeneity (as prediction intervals) and individual effect sizes. 2. Here, we have added significant enhancements by integrating many new functionalities as orchaRd 2.0. This updated version allows the visualisation of heteroscedasticity (different variances across levels of a categorical moderator), marginal estimates (e.g., marginalising out effects other than the one visualized), conditional estimates (i.e., estimates of different groups conditioned upon specific values of a continuous variable), and visualizations of all types of interactions between two categorical/continuous moderators.3. orchaRd 2.0 has additional functions which calculate key statistics from multilevel meta-analytic models such as I2 and R2. Importantly, orchaRd 2.0 contributes to better reporting by complying with PRISMA-EcoEvo (preferred reporting items for systematic reviews and meta-analyses in ecology and evolution). Taken together, orchaRd 2.0 can improve the presentation of meta-analytic results and facilitate the exploration of previously neglected patterns. 4. In addition, as a part of a literature survey, we found that graphical packages are rarely cited (~3%). We plea that researchers credit developers and maintainers of graphical packages, e.g., by citations in a figure legend, acknowledging the use of relevant packages.
Finding the optimal balance between survival and reproduction is a central puzzle in life-history theory. The terminal investment hypothesis predicts that when individuals encounter a survival threat that compromises future reproductive potential, they will increase immediate reproductive investment to maximise fitness. Despite decades of research on the terminal investment hypothesis, findings remain mixed. We examined the terminal investment hypothesis with a meta-analysis of studies that measured reproductive investment of multicellular iteroparous animals after a non-lethal immune challenge. We had two main aims. The first was to investigate whether individuals, on average, increase reproductive investment in response to an immune threat, as predicted by the terminal investment hypothesis. We also examined whether such responses vary adaptively on factors associated with the amount of reproductive opportunities left (residual reproductive value) in the individuals, as predicted by the terminal investment hypothesis. The second was to provide a quantitative test of a novel prediction based on the dynamic threshold model: that an immune threat increases between-individual variance in reproductive investment. Our results provided some support for our hypotheses. Older individuals, who are expected to have lower residual reproductive values, showed stronger mean terminal investment response than younger individuals. In terms of variance, individuals showed a divergence in responses, leading to an increase in variance. This increase in variance was especially amplified in longer-living species, which was consistent with our prediction that individuals in longer-living species should respond with greater individual variation due to increased phenotypic plasticity. We find little statistical evidence of publication bias. Together, our results highlight the need for a more nuanced view on the terminal investment hypothesis and a greater focus on the factors that drive individual responses.
Abstract The obesity epidemic, largely driven by the accessibility of ultra‐processed high‐energy foods, is one of the most pressing public health challenges of the 21st century. Consequently, there is increasing concern about the impacts of diet‐induced obesity on behavior and cognition. While research on this matter continues, to date, no study has explicitly investigated the effect of obesogenic diet on variance and covariance (correlation) in behavioral traits. Here, we examined how an obesogenic versus control diet impacts means and (co‐)variances of traits associated with body condition, behavior, and cognition in a laboratory population of ~160 adult zebrafish (Danio rerio). Overall, an obesogenic diet increased variation in several zebrafish traits. Zebrafish on an obesogenic diet were significantly heavier and displayed higher body weight variability; fasting blood glucose levels were similar between control and treatment zebrafish. During behavioral assays, zebrafish on the obesogenic diet displayed more exploratory behavior and were less reactive to video stimuli with conspecifics during a personality test, but these significant differences were sex‐specific. Zebrafish on an obesogenic diet also displayed repeatable responses in aversive learning tests whereas control zebrafish did not, suggesting an obesogenic diet resulted in more consistent, yet impaired, behavioral responses. Where behavioral syndromes existed (inter‐class correlations between personality traits), they did not differ between obesogenic and control zebrafish groups. By integrating a multifaceted, holistic approach that incorporates components of (co‐)variances, future studies will greatly benefit by quantifying neglected dimensions of obesogenic diets on behavioral changes.
Publication bias threatens the validity of quantitative evidence from meta‐analyses as it results in some findings being overrepresented in meta‐analytic datasets because they are published more frequently or sooner (e.g. ‘positive’ results). Unfortunately, methods to test for the presence of publication bias, or assess its impact on meta‐analytic results, are unsuitable for datasets with high heterogeneity and non‐independence, as is common in ecology and evolutionary biology. We first review both classic and emerging publication bias tests (e.g. funnel plots, Egger's regression, cumulative meta‐analysis, fail‐safe N, trim‐and‐fill tests, p‐curve and selection models), showing that some tests cannot handle heterogeneity, and, more importantly, none of the methods can deal with non‐independence. For each method, we estimate current usage in ecology and evolutionary biology, based on a representative sample of 102 meta‐analyses published in the last 10 years. Then, we propose a new method using multilevel meta‐regression, which can model both heterogeneity and non‐independence, by extending existing regression‐based methods (i.e. Egger's regression). We describe how our multilevel meta‐regression can test not only publication bias, but also time‐lag bias, and how it can be supplemented by residual funnel plots. Overall, we provide ecologists and evolutionary biologists with practical recommendations on which methods are appropriate to employ given independent and non‐independent effect sizes. No method is ideal, and more simulation studies are required to understand how Type 1 and Type 2 error rates are impacted by complex data structures. Still, the limitations of these methods do not justify ignoring publication bias in ecological and evolutionary meta‐analyses.