In recent years, concern has grown about the inappropriate application and interpretation of P values, especially the use of P<0.05 to denote “statistical significance” and the practice of P-hacking to produce results below this threshold and selectively reporting these in publications. Such behavior is said to be a major contributor to the large number of false and non-reproducible discoveries found in academic journals. In response, it has been proposed that the threshold for statistical significance be changed from 0.05 to 0.005. The aim of the current study was to use an evolutionary agent-based model comprised of researchers who test hypotheses and strive to increase their publication rates in order to explore the impact of a 0.005 P value threshold on P-hacking and published false positive rates. Three scenarios were examined, one in which researchers tested a single hypothesis, one in which they tested multiple hypotheses using a P<0.05 threshold, and one in which they tested multiple hypotheses using a P<0.005 threshold. Effects sizes were varied across models and output assessed in terms of researcher effort, number of hypotheses tested and number of publications, and the published false positive rate. The results supported the view that a more stringent P value threshold can serve to reduce the rate of published false positive results. Researchers still engaged in P-hacking with the new threshold, but the effort they expended increased substantially and their overall productivity was reduced, resulting in a decline in the published false positive rate. Compared to other proposed interventions to improve the academic publishing system, changing the P value threshold has the advantage of being relatively easy to implement and could be monitored and enforced with minimal effort by journal editors and peer reviewers.
Abstract Managing social‐ecological systems (SES) requires balancing the need to tailor actions to local heterogeneity and the need to work over large areas to accommodate the extent of SES. This balance is particularly challenging for policy since the level of government where the policy is being developed determines the extent and resolution of action. We make the case for a new research agenda focused on ecological federalism that seeks to address this challenge by capitalizing on the flexibility afforded by a federalist system of governance. Ecological federalism synthesizes the environmental federalism literature from law and economics with relevant ecological and biological literature to address a fundamental question: What aspects of SES should be managed by federal governments and which should be allocated to decentralized state governments? This new research agenda considers the bio‐geo‐physical processes that characterize state‐federal management tradeoffs for biodiversity conservation, resource management, infectious disease prevention, and invasive species control. Read the free Plain Language Summary for this article on the Journal blog.
Agent‐based models (ABMs) are increasing in popularity as tools to simulate and explore many biological systems. Successes in simulation lead to deeper investigations, from designing systems to optimizing performance. The typically stochastic, rule‐based structure of ABMs, however, does not lend itself to analytic and numerical techniques of optimization the way traditional dynamical systems models do. The goal of this work is to illustrate a technique for approximating ABMs with a partial differential equation (PDE) system to design some management strategies on the ABM. We propose a surrogate modeling approach, using differential equations that admit direct means of determining optimal controls, with a particular focus on environmental heterogeneity in the ABM. We implement this program with both PDE and ordinary differential equation (ODE) approximations on the well‐known rabbits and grass ABM, in which a pest population consumes a resource. The control problem addressed is the reduction of this pest population through an optimal control formulation. After fitting the ODE and PDE models to ABM simulation data in the absence of control, we compute optimal controls using the ODE and PDE models, which we them apply to the ABM. The results show promise for approximating ABMs with differential equations in this context.
The authority to manage natural capital often follows political boundaries rather than ecological. This mismatch can lead to unsustainable outcomes, as spillovers from one management area to the next may create adverse incentives for local decision making, even within a single country. At the same time, one-size-fits-all approaches of federal (centralized) authority can fail to respond to state (decentralized) heterogeneity and can result in inefficient economic or detrimental ecological outcomes. Here we utilize a spatially explicit coupled natural-human system model of a fishery to illuminate trade-offs posed by the choice between federal vs. state control of renewable resources. We solve for the dynamics of fishing effort and fish stocks that result from different approaches to federal management that vary in terms of flexibility. Adapting numerical methods from engineering, we also solve for the open-loop Nash equilibrium characterizing state management outcomes, where each state anticipates and responds to the choices of the others. We consider traditional federalism questions (state vs. federal management) as well as more contemporary questions about the economic and ecological impacts of shifting regulatory authority from one level to another. The key mechanisms behind the trade-offs include whether differences in local conditions are driven by biological or economic mechanisms; degree of flexibility embedded in the federal management; the spatial and temporal distribution of economic returns across states; and the status-quo management type. While simple rules-of-thumb are elusive, our analysis reveals the complex political economy dimensions of renewable resource federalism.
A gene regulatory network (GRN) is a set of transcription factors which regulate the level of expression of genes encoding other transcription factors. The dynamics of a GRN show how gene expression in the network changes over time. Microarray data were obtained from the Saccharomyces cerevisiae wild type strain and five transcription factor deletion strains (Δcin5, Δgln3, Δhap4, Δhmo1, Δzap1) before cold shock at 13°C and 15, 30, and 60 minutes after cold shock (NCBI GEO Series GSE83656). Genes that showed a significant change in expression were submitted to the YEASTRACT database to determine which transcription factors regulated them and to generate a candidate GRN of 15 nodes (transcription factors) and 28 edges (regulatory relationships). The edges of this intact network were then systematically deleted one‐at‐a‐time to create a family of 28 additional networks to determine the importance of each edge in the network. We used the open source software, GRNmap ( http://kdahlquist.github.io/GRNmap/), to model the dynamics of these GRNs. GRNmap models the change in expression for each gene as the production of mRNA minus its degradation using differential equations with a sigmoidal production function. Given published mRNA degradation rates (Neymotin et al. 2014; doi: 10.1261/rna.045104.114) and cold shock microarray data, GRNmap then estimated the production rates and expression thresholds for each gene, and regulatory weights for each edge, which denote the direction (activation or repression) and strength of the regulatory relationships. Edge weights were then visualized with the open source software GRNsight (Dahlquist et al. 2014; doi: 10.7717/peerj‐cs.85; https://dondi.github.io/GRNsight/). Figure 1 shows the intact GRN. Red edges indicate activation; blue edges indicate repression. Node stripes left to right indicate gene expression at 15, 30, and 60 minutes of cold shock, red is increased expression; blue is decreased. To evaluate the goodness of fit of the model, GRNmap reports the least squares error (LSE) between the model and the microarray data. The LSE was compared to the minimum theoretical least squares error (minLSE) achievable due to variation in the data, to determine the model performance between networks. In the 28 edge‐deletion networks, LSE:minLSE ratios indicated that five networks performed better than the intact network, ten networks performed about the same, and thirteen performed worse. The edge‐deletions involving the Hmo1, Msn2, and Cin5 transcription factors resulted in a poor fit of the model to data, indicating that those edges represent important regulatory relationships in the cold shock response. K‐means clustering was performed on the edge weight values from the intact network and edge‐deletion networks. An examination of the clusters showed that weight values deviated from those of the intact network for 6 of 9 edge‐deletions involving Msn2, 4 of 5 edge‐deletions involving Hmo1, and 3 of 6 edge‐deletions involving Cin5, further reinforcing their importance in the network controlling the cold shock response in yeast.Support or Funding InformationLoyola Marymount University Rains Research Assistant Program (A.K.F.)Figure 1
In recent years, serious concerns have arisen about reproducibility in science. Estimates of the cost of irreproducible preclinical studies range from 28 billion USD per year in the USA alone (Freedman et al. in PLoS Biol 13(6):e1002165, 2015) to over 200 billion USD per year worldwide (Chalmers and Glasziou in Lancet 374:86-89, 2009). The situation in the social sciences is not very different: Reproducibility in psychological research, for example, has been estimated to be below 50% as well (Open Science Collaboration in Science 349:6251, 2015). Less well studied is the issue of reproducibility of simulation research. A few replication studies of agent-based models, however, suggest the problem for computational modeling may be more severe than for laboratory experiments (Willensky and Rand in JASSS 10(4):2, 2007; Donkin et al. in Environ Model Softw 92:142-151, 2017; Bajracharya and Duboz in: Proceedings of the symposium on theory of modeling and simulationDEVS integrative M&S symposium, pp 6-11, 2013). In this perspective, we discuss problems of reproducibility in agent-based simulations of life and social science problems, drawing on best practices research in computer science and in wet-lab experiment design and execution to suggest some ways to improve simulation research practice.
Agent-based models (ABMs) have become an increasingly important mode of inquiry for the life sciences. They are particularly valuable for systems that are not understood well enough to build an equation-based model. These advantages, however, are counterbalanced by the difficulty of analyzing and using ABMs, due to the lack of the type of mathematical tools available for more traditional models, which leaves simulation as the primary approach. As models become large, simulation becomes challenging. This paper proposes a novel approach to two mathematical aspects of ABMs, optimization and control, and it presents a few first steps outlining how one might carry out this approach. Rather than viewing the ABM as a model, it is to be viewed as a surrogate for the actual system. For a given optimization or control problem (which may change over time), the surrogate system is modeled instead, using data from the ABM and a modeling framework for which ready-made mathematical tools exist, such as differential equations, or for which control strategies can explored more easily. Once the optimization problem is solved for the model of the surrogate, it is then lifted to the surrogate and tested. The final step is to lift the optimization solution from the surrogate system to the actual system. This program is illustrated with published work, using two relatively simple ABMs as a demonstration, Sugarscape and a consumer-resource ABM. Specific techniques discussed include dimension reduction and approximation of an ABM by difference equations as well systems of PDEs, related to certain specific control objectives. This demonstration illustrates the very challenging mathematical problems that need to be solved before this approach can be realistically applied to complex and large ABMs, current and future. The paper outlines a research program to address them.
A gene regulatory network (GRN) consists of genes, transcription factors, and the regulatory connections between them that govern the level of expression of mRNA and proteins from those genes. Over a period of several years, our group has developed a MATLAB software package, called GRNmap, that uses ordinary differential equations to model the dynamics of medium-scale GRNs. The program uses a penalized least squares approach (Dahlquist et al. 2015, DOI: 10.1007/s11538-015-0092-6) to estimate production rates, expression thresholds, and regulatory weights for each transcription factor in the network based on gene expression data, and then performs a forward simulation of the dynamics of the network. GRNmap has options for using a sigmoidal or Michaelis-Menten production function. The large number of developers and time span of development led to a code base that was difficult to revise and adjust. We therefore brought the code under version control in a GitHub repository and refactored the script-based software with global variables into a function-based package that uses an object to carry relevant information from function to function. This modular approach allows for cleaner, less ambiguous code and increased maintainability. We standardized the format of the input and output Excel workbooks, making them more readable. We also added an optimization diagnostics output worksheet which includes both the actual and theoretical minimum least squared error overall, and the mean squared errors for the individual genes. The MATLAB compiler was used to create an executable that can be run on any Windows machine without the need of a MATLAB license, increasing the accessibility of our program. Finally, we have implemented test-driven development, creating unit tests for all new features to speed up debugging and to prevent future code regressions. We are improving the test coverage of previous code. GRNsight is an open source web application for visualizing such models of gene regulatory networks. GRNsight accepts GRNmap- or user-generated spreadsheets containing an adjacency matrix representation of the GRN and automatically lays out the graph of the GRN model. It is written in JavaScript, with diagrams facilitated by D3.js. Node.js and the Express framework handle server-side functions. GRNsight’s diagrams are based on D3.js’s force graph layout algorithm, which was then extensively customized. GRNsight uses pointed and blunt arrowheads, and colors the edges and adjusts their thicknesses based on the sign (activation or repression) and magnitude of the GRNmap weight parameter. Visualizations can be modified through manual node dragging and sliders that adjust the force graph parameters. From the early stages, GRNsight has had a unit testing framework using Mocha and the Chai assertion library to perform test-driven development where unit tests are written before new functionality is coded. This framework consists of over 160 automated unit tests that examine over 450 test files to ensure that the program is running as expected. Error and warning messages inform the user what happened, the source of the problem, and possible solutions. Together, the life cycle of these two programs illustrate the differences between the cultures of mathematics and computing, the challenges and benefits of bringing an existing code base up to open development standards (GRNmap), and the advantages of starting a project using best practices from the beginning (GRNsight). Our goal is to facilitate reproducible research.
GRNsight is a web application and service for visualizing models of gene regulatory networks (GRNs). A gene regulatory network (GRN) consists of genes, transcription factors, and the regulatory connections between them which govern the level of expression of mRNA and protein from genes. The original motivation came from our efforts to perform parameter estimation and forward simulation of the dynamics of a differential equations model of a small GRN with 21 nodes and 31 edges. We wanted a quick and easy way to visualize the weight parameters from the model which represent the direction and magnitude of the influence of a transcription factor on its target gene, so we created GRNsight. GRNsight automatically lays out either an unweighted or weighted network graph based on an Excel spreadsheet containing an adjacency matrix where regulators are named in the columns and target genes in the rows, a Simple Interaction Format (SIF) text file, or a GraphML XML file. When a user uploads an input file specifying an unweighted network, GRNsight automatically lays out the graph using black lines and pointed arrowheads. For a weighted network, GRNsight uses pointed and blunt arrowheads, and colors the edges and adjusts their thicknesses based on the sign (positive for activation or negative for repression) and magnitude of the weight parameter. GRNsight is written in JavaScript, with diagrams facilitated by D3.js, a data visualization library. Node.js and the Express framework handle server-side functions. GRNsight's diagrams are based on D3.js's force graph layout algorithm, which was then extensively customized to support the specific needs of GRNs. Nodes are rectangular and support gene labels of up to 12 characters. The edges are arcs, which become straight lines when the nodes are close together. Self-regulatory edges are indicated by a loop. When a user mouses over an edge, the numerical value of the weight parameter is displayed. Visualizations can be modified by sliders that adjust the force graph layout parameters and through manual node dragging. GRNsight is best-suited for visualizing networks of fewer than 35 nodes and 70 edges, although it accepts networks of up to 75 nodes or 150 edges. GRNsight has general applicability for displaying any small, unweighted or weighted network with directed edges for systems biology or other application domains. GRNsight serves as an example of following and teaching best practices for scientific computing and complying with FAIR principles, using an open and test-driven development model with rigorous documentation of requirements and issues on GitHub. An exhaustive unit testing framework using Mocha and the Chai assertion library consists of around 160 automated unit tests that examine nearly 530 test files to ensure that the program is running as expected. The GRNsight application (http://dondi.github.io/GRNsight/) and code (https://github.com/dondi/GRNsight) are available under the open source BSD license.
A gene regulatory network (GRN) consists of a set of transcription factors that regulate the level of expression of genes encoding other transcription factors. The dynamics of a GRN show how gene expression in the network changes over time. While the transcriptional regulation of the response to the environmental stress of heat shock in budding yeast, Saccharomyces cerevisiae, is well‐understood, which transcription factors regulate the early response to cold shock is not fully understood. Thus, the focus of this study was to determine the GRN that controls the cold shock response in yeast and to model its dynamics. Microarray experiments were performed in the Dahlquist lab at various timepoints after cold shock (t=15, 30, and 60 minutes) to examine the gene expression patterns of both the wild type and mutant strains of yeast that had been deleted for the transcription factors Gln3 and Zap1. The microarray data were normalized using the limma package in R and a modified ANOVA was used to determine which genes had a log2 fold change significantly different than zero at any of the timepoints studied. The genes that met the significance criterion of an adjusted Benjamini and Hochberg p < 0.05 were submitted to the YEASTRACT database to determine which transcription factors potentially regulated those genes. Two GRNs, one derived from the Gln3 deletion strain data and one from the Zap1 deletion strain data were created. The GRNs and the corresponding microarray data were then input into GRNmap, a MATLAB software package that uses ordinary differential equations to model the dynamics of medium‐scale GRNs. The program estimated the production rates, expression thresholds, and regulatory weights for each transcription factor in the network based on the microarray data, and then performed a forward simulation of the dynamics of the network. The results of the modeling for both networks were similar. The best fit of model to data was for genes directly connected to the deleted transcription factor (either Gln3 or Zap1). However, the deletion of either transcription factor from their respective networks had a large effect on the dynamics of several genes in the network that were not accounted for by the network connections. This suggests that the network structure did not fully model the actual cellular conditions of cold shock. This may be because the regulatory networks in the YEASTRACT database are based on measurements taken during other types of growth conditions, not cold shock. Our future directions include expanding the number of GRNs that we model to determine which transcription factors are missing from the current model. Our working code is available on the GRNmap page (http://kdahlquist.github.io/GRNmap/), and visualization of the network is available on GRNsight (http://dondi.github.io/GRNsight/).Support or Funding InformationThis work was partially supported by NSF award 0921038 (B.G.F., K.D.D.), a Kadner‐Pitts Research Grant (K.D.D.) and the Loyola Marymount University Rains Research Assistant Program (T.A.M.).
Gene expression is regulated by proteins called transcription factors which can either repress or activate a gene's transcriptional output. A gene regulatory network (GRN) consists of a set of transcription factors that regulate the level of expression of genes encoding other transcription factors. The dynamics of a GRN show how gene expression in the network changes over time. Previously in the lab, a MATLAB software package called GRNmap was developed that uses ordinary differential equations to model the dynamics of medium‐scale GRNs from budding yeast, Saccharomyces cerevisiae. The program estimates production rates, expression thresholds, and regulatory weights for each transcription factor in the network based on DNA microarray data, and then performs a forward simulation of the dynamics of the network. DNA microarray data for 6189 yeast genes was obtained from the Dahlquist lab where they subjected yeast to cold shock at 13°C and measured gene expression at three time points (after 15, 30, and 60 minutes of cold shock). We performed LOESS normalization using the limma package in R and used a modified ANOVA to determine which genes had a log2 fold change significantly different than zero at any of the timepoints studied. Using GRNmap, we estimated the parameters of a GRN derived from the YEASTRACT database, consisting of 21 nodes (transcription factors) and 50 edges (regulatory relationships). To answer our fundamental question, whether the network accurately models what actually occurs during cold shock in the yeast cell, we analyzed the results of the modeling as follows. For each gene we evaluated how well the simulated data generated by the model fit the expression profile measured by the microarrays. Factors contributing to the goodness of fit included whether the transcripton factor itself exhibited significant dynamics (in the ANOVA test) and whether the regulators of that transcription actor also showed significant dynamics. This analysis showed that, while a few genes were modeled well, others were not, potentially because they are regulated by other transcription factors that were not present in the network modeled. To investigate this further, we compared the results of the modeling of the YEASTRACT‐derived network to 10 random networks that had the same nodes and and the same number of edges, but had randomized connections between nodes. We found that the random networks had larger values for the least squares error in the estimation and a different structure and degree‐distribution than the YEASTRACT‐derived network, suggesting that the YEASTRACT‐derived network modeled the transcriptional dynamics better. To improve the performance of the model and account for missing regulators, we are now evaluating a family of YEASTRACT‐derived networks by paring down an initial network of 35 nodes systematically down to 15 nodes by removing the next least significant transcription factor one by one from the network. From this analysis we expect to gain insight into the gene regulatory network that controls the cold shock response in yeast. Our working code is available on the GRNmap page (http://kdahlquist.github.io/GRNmap/), and visualization of the network is available on GRNsight (http://dondi.github.io/GRNsight/).Support or Funding InformationThis work was partially supported by NSF award 0921038 (B.G.F., K.D.D.), a Kadner‐Pitts Research Grant (K.D.D., K.G.J., N.E.W.), and a Loyola Marymount University Honors Summer Research Fellowship (K.G.J., N.E.W.).
A gene regulatory network (GRN) consists of genes, transcription factors, and the regulatory connections between them that govern the level of expression of mRNA and proteins from those genes. Over a period of several years, our group has developed a MATLAB software package, called GRNmap, that uses ordinary differential equations to model the dynamics of medium-scale GRNs. The program uses a penalized least squares approach (Dahlquist et al. 2015, DOI: 10.1007/s11538-015-0092-6) to estimate production rates, expression thresholds, and regulatory weights for each transcription factor in the network based on gene expression data, and then performs a forward simulation of the dynamics of the network. GRNmap has options for using a sigmoidal or Michaelis-Menten production function. The large number of developers and time span of development led to a code base that was difficult to revise and adjust. We therefore brought the code under version control in a GitHub repository and refactored the script-based software with global variables into a function-based package that uses an object to carry relevant information from function to function. This modular approach allows for cleaner, less ambiguous code and increased maintainability. We standardized the format of the input and output Excel workbooks, making them more readable. We also added an optimization diagnostics output worksheet which includes both the actual and theoretical minimum least squared error overall, and the mean squared errors for the individual genes. The MATLAB compiler was used to create an executable that can be run on any Windows machine without the need of a MATLAB license, increasing the accessibility of our program. Finally, we have implemented test-driven development, creating unit tests for all new features to speed up debugging and to prevent future code regressions. We are improving the test coverage of previous code. GRNsight is an open source web application for visualizing such models of gene regulatory networks. GRNsight accepts GRNmap- or user-generated spreadsheets containing an adjacency matrix representation of the GRN and automatically lays out the graph of the GRN model. It is written in JavaScript, with diagrams facilitated by D3.js. Node.js and the Express framework handle server-side functions. GRNsight’s diagrams are based on D3.js’s force graph layout algorithm, which was then extensively customized. GRNsight uses pointed and blunt arrowheads, and colors the edges and adjusts their thicknesses based on the sign (activation or repression) and magnitude of the GRNmap weight parameter. Visualizations can be modified through manual node dragging and sliders that adjust the force graph parameters. From the early stages, GRNsight has had a unit testing framework using Mocha and the Chai assertion library to perform test-driven development where unit tests are written before new functionality is coded. This framework consists of over 160 automated unit tests that examine over 450 test files to ensure that the program is running as expected. Error and warning messages inform the user what happened, the source of the problem, and possible solutions. Together, the life cycle of these two programs illustrate the differences between the cultures of mathematics and computing, the challenges and benefits of bringing an existing code base up to open development standards (GRNmap), and the advantages of starting a project using best practices from the beginning (GRNsight). Our goal is to facilitate reproducible research. Project Websites: http://kdahlquist.github.io/GRNmap/ and http://dondi.github.io/GRNsight/ Source Code: https://github.com/kdahlquist/GRNmap/ and https://github.com/dondi/GRNsight The video for this presentation at BOSC 2016 can be viewed here .
GRNsight is a web application and service for visualizing models of gene regulatory networks (GRNs). A gene regulatory network consists of genes, transcription factors, and the regulatory connections between them which govern the level of expression of mRNA and protein from genes. GRNmap, a MATLAB program that performs parameter estimation and forward simulation of a differential equations model of a GRN, can mathematically model the dynamics of GRNs. GRNsight automatically lays out the network graph based on GRNmap output spreadsheets. GRNsight uses pointed and blunt arrowheads, and colors the edges and adjusts their thicknesses based on the sign (activation or repression) and magnitude of the GRNmap weight parameter. Visualizations can be modified through manual node dragging and sliders that adjust the force graph parameters. We have now implemented an exhaustive unit testing framework using Mocha and the Chai assertion library to perform test‐driven development where unit tests are written before new functionality is coded. This framework consists of over 135 automated unit tests that examine about 450 test files to ensure that the program is running as expected. Error and warning messages have a three‐part framework that informs the user what happened, the source of the problem, and possible solutions. For example, GRNsight returns an error when the spreadsheet is formatted incorrectly or the maximum number of nodes or edges is exceeded. The completion of the testing framework marks the close of development for version 1 (the current release stands at version 1.12). In version 2.0 of GRNsight, a new feature will be implemented that colors the nodes (genes) based on the expression data. The user will have the choice to display the experimental data or the data produced by the forward simulation feature in GRNmap. GRNsight is available at http://dondi.github.io/GRNsight/.Support or Funding InformationThis work was partially supported by NSF award 0921038 (B.G.F., K.D.D.), a Kadner‐Pitts Research Grant (K.D.D.), the Loyola Marymount University Summer Undergraduate Research Program (A.V.) and the Loyola Marymount University Rains Research Assistant Program (N.A.A.).
College drinking is a problem with severe academic, health, and safety consequences. The underlying social processes that lead to increased drinking activity are not well understood. Social Norms Theory is an approach to analysis and intervention based on the notion that students' misperceptions about the drinking culture on campus lead to increases in alcohol use. In this paper we develop an agent-based simulation model, implemented in MATLAB, to examine college drinking. Students' drinking behaviors are governed by their identity (and how others perceive it) as well as peer influences, as they interact in small groups over the course of a drinking event. Our simulation results provide some insight into the potential effectiveness of interventions such as social norms marketing campaigns.