Many natural systems are observed as point patterns in time, space, or space and time. Examples include plant and cellular systems, animal colonies, earthquakes, and wildfires. In practice the locations of the points are not always observed correctly. However, in the point process literature, there has been relatively scant attention paid to the issue of errors in the location of points. In this paper, we discuss how the observed point pattern may deviate from the actual point pattern and review methods and models that exist to handle such deviations. The discussion is supplemented with several scientific illustrations.
Around the world, extreme sea levels are increasingly affecting low-lying and unprotected coasts. The rise in local mean sea level is widely assumed to be the dominant driver of changes in the frequency of coastal extremes, and most projections of future flood risk rely on the premise that extremes increase uniformly with mean sea level. Here, we show that this assumption does not always hold. Using non-stationary extreme-value models applied to long, high-quality tide-gauge records from the global GESLA-3 database, we find that roughly 15\% of sites exhibit statistically significant departures from this behaviour, with extreme sea levels rising faster or slower than local mean sea level. Regions such as the North Sea show surge amplification beyond what sea-level rise alone can explain, consistent with depth-sensitive wave dynamics in shallow basins. These findings highlight the need to incorporate non-stationary effects in coastal hazard assessments to support robust adaptation planning.
We introduce a class of proper scoring rules for evaluating spatial point process forecasts based on summary statistics. These scoring rules rely on Monte Carlo approximations of expectations and can therefore easily be evaluated for any point process model that can be simulated. In this regard, they are more flexible than the commonly used logarithmic score and other existing proper scores for point process predictions. The scoring rules allow for evaluating the calibration of a model to specific aspects of a point process, such as its spatial distribution or tendency toward clustering. Using simulations, we analyze the sensitivity of our scoring rules to different aspects of the forecasts and compare it to the logarithmic score. Applications to earthquake occurrences in northern California, United States and the spatial distribution of Pacific silver firs in Findley Lake Reserve in Washington highlight the usefulness of our scores for scientific model selection.
Abstract A new statistical approach to validating climate models is introduced. First, five observational estimates of global mean surface temperature with estimated standard errors are combined into one data product, latent observed annual global temperature anomalies for the years 1880–2014, using a Bayesian hierarchical statistical approach. Summarizing these observed anomalies, estimates of smooth trend, levels of warming, and residual dependence as summarized by the spectral density function, are provided with simultaneous 95% credible bands. Then, corresponding estimates of smooth trend, levels of warming, and residual dependence are produced for sixth Climate Model Intercomparison Project (CMIP6) historical simulations analyzed at the annual global temperature anomaly scale, and compared to these bands. Among our results, we find that 93 out of the 318 CMIP6 historical model runs contain trends fitting inside the simultaneous bands for the smooth trend constructed from the data products, and for residual temporal dependence 69 out of 318 model runs contain spectral density functions that are within the corresponding data‐product‐based‐bands. We estimate the mean global temperature increase from 1995–2014 relative to 1880–1899, from the data product, to be 0.896°C with a 95% credible interval of between 0.877 and 0.915. We find that 14 CMIP6 model runs agree with this interval, 197 model runs lead to a smaller temperature increase globally, and 107 model runs lead to a larger temperature increase.
The Pacific Northwest (PNW) of North America encompasses diverse tectonic settings that can produce damaging earthquakes near population centers. Seismicity in this region is often clustered into aftershock sequences and swarms, and their patterns and frequencies differ across subregions or tectonic regimes. Characterizing the seismicity of the PNW requires a catalog of observed earthquakes. Furthermore, applications with the catalog may require earthquake clusters to be identified and regarded separately. Unlike previous studies, we explicate how to overcome challenges when combining catalogs from different countries, particularly in accounting for duplicate events and other discrepancies. We apply this to merge authoritative catalogs for the United States and Canadian portions of the PNW, along with a third dataset with data quality measures. We also perform a window-based search for earthquake clusters, which then get labeled as possible or definite swarms or aftershock sequences. We further split the catalog into its two primary tectonic regimes. We then study the PNW catalog’s completeness, and the extent to which this varies between the northern and southern parts of the region. We provide a harmonized international PNW catalog with derived variables describing earthquake clustering and tectonic regimes. This entire processing pipeline has also been fully documented and is supported with software, enabling its use in other seismic regions.
Every statistical analysis must make a variety of modeling choices, each of which may be debatable. Our choice of B-splines to model the trend was perhaps the most frequent point of disagreement between us and the discussants. While the concept of trend is not always well defined in the literature, we often define trend to be continuous changes over longer scales that can be modeled independently from the rest of the series. SR suggest using a dynamic linear model (essentially a locally linear trend model with dependent errors) and show that this yields similar answers to our B-spline approach. This is perhaps not surprising considering the relationship between splines and random walk/stochastic differential equations (see, e.g., Fahrmeir & Lang, 2001; Speckman & Sun, 2003; Wahba, 1978; Yue et al., 2014). The trend increment estimate in the upper right panel of Figure 1 of SR either could be integrated to produce a trend estimate comparable to ours, or one could difference our trend estimate. Figure 1a shows our differenced trend estimate together with the estimate from SR. Essentially, our estimate looks like a smooth version of the estimate in SR, with the exception that our rate of growth has decreased more than their rate toward the end (this is partly an edge effect in our estimate). We put a horizontal line at our estimated growth rate since 1980. PS suggest either relating the mean to radiative forcings or to an ensemble mean of climate models. While there is no simple relationship between greenhouse gases and temperature, we find that there is a moderate positive association between a global time series of greenhouse gas emissions up to 2014 taken from Gütschow et al. (2016) and the posterior mean of our reconstruction. Including this greenhouse gas emissions series as an extra covariate in our model, a 95% credible interval for this regression parameter contains zero. Climate models, on the other hand, are an interesting suggestion, and could perhaps be used as a prior estimate of the trend function. We used a B-spline estimate of trend fitted to the ensemble mean of 58 CMIP6 models (Eyring et al., 2016)
See project description for general summary of data structure. The attached file is table reporting the estimates of the posterior daily wildfire size time series for the South Fork Complex in Oregon 2014.
Earthquake models can produce aftershock forecasts, which have recently been released to lay audiences. While visualization literature suggests that displaying forecast uncertainty can improve how forecast maps are used, research on uncertainty visualization is missing from earthquake science. We designed a pre-registered online experiment to test the effectiveness of three visualization techniques for displaying aftershock forecast maps and their uncertainty. These maps showed the forecasted number of aftershocks at each location for a week following a hypothetical mainshock, along with the uncertainty around each location's forecast. Three different uncertainty visualizations were produced: (1) forecast and uncertainty maps adjacent to one another; (2) the forecast map depicted in a color scheme, with the uncertainty shown by the transparency of the color; and (3) two maps that showed the lower and upper bounds of the forecast distribution at each location. We compared the three uncertainty visualizations using tasks that were specifically designed to address broadly applicable and user-generated communication goals. We compared task responses between participants using uncertainty visualizations and using the forecast map shown without its uncertainty (the current practice). Participants completed two map-reading tasks that targeted several dimensions of the readability of uncertainty visualizations. Participants then performed a Comparative Judgment task, which demonstrated whether a visualization was successful in reaching two key communication goals: indicating where many aftershocks and no aftershocks are likely (sure bets) and where the forecast is low but the uncertainty is high enough to imply potential risk (surprises). All visualizations performed equally well in the goal of communicating sure bet situations. But the visualization with lower and upper bounds was substantially better than the other designs at communicating surprises. These results have implications for the visual communication of forecast uncertainty both within and beyond earthquake science.
We focus on the discussion of modeling processes that are observed at fixed locations of a region (geostatistics). A standard approach is to assume that the process of interest follows a Gaussian Process with some mean and (valid) covariance functions. It is common to model the covariance function as the product between a variance parameter, and a correlation function which is a function of the Euclidean distance between locations. This implies that the distribution of the process is unchanged when the origin of the index set is translated, and the process is invariant under rotation about the origin; that is the process is stationary and isotropic or homogeneous. However, the assumption of stationarity and isotropy (homogeneity) rarely holds in practice. Commonly, the correlation structures of such processes are influenced by local characteristics resulting in different behaviors in neighborhoods of different spatial locations. We review models that allow for heterogeneous covariance structures and point to some avenues of future research.
We propose an interacting cluster model for the spatial distribution of epidermal nerve fibers (ENF). The model consists of a spatial process of parent points modeling the base points of nerve fiber bundles. To each base point there is an offspring point process of fibers end points associated. The parent process is, possibly, inhibited by fiber endings that belong to different base bundles. The fibers themselves have random length and spatial orientation. We consider a non-orphan process where we can connect each offspring to a parent. We examine the processes that can be described with our model, how coefficient estimation can be performed under the Bayesian paradigm via Markov chain Monte Carlo methods, and detail approaches for the implementation of efficient MCMC sampling schemes. We study the performance of the estimation procedure through a simulation study. An application to data from skin blister biopsy images of ENFs is presented.
See project description for general summary of data structure. The attached file is table reporting the estimates of the posterior daily wildfire size time series for the Beaver Fire in California 2014.
Single-cell lineage tracking strategies enabled by recent experimental technologies have produced significant insights into cell fate decisions, but lack the quantitative framework necessary for rigorous statistical analysis of mechanistic models describing cell division and differentiation. In this paper, we develop such a framework with corresponding moment-based parameter estimation techniques for continuous-time, multi-type branching processes. Such processes provide a probabilistic model of how cells divide and differentiate, and we apply our method to study hematopoiesis, the mechanism of blood cell production. We derive closed-form expressions for higher moments in a general class of such models. These analytical results allow us to efficiently estimate parameters of much richer statistical models of hematopoiesis than those used in previous statistical studies. After validating the methodology in simulation studies, we apply our estimator to hematopoietic lineage tracking data from rhesus macaques. Our analysis provides a more complete understanding of cell fate decisions during hematopoiesis in non-human primates, which may be more relevant to human biology and clinical strategies than previous findings from murine studies. For example, in addition to previously estimated hematopoietic stem cell self-renewal rate, we are able to estimate fate decision probabilities and to compare structurally distinct models of hematopoiesis using cross validation. These estimates of fate decision probabilities and our model selection results should help biologists compare competing hypotheses about how progenitor cells differentiate. The methodology is transferrable to a large class of stochastic compartmental models and multi-type branching models, commonly used in studies of cancer progression, epidemiology, and many other fields.
Three Nordic countries collaborate to build a suite of eScience tools to support long-term planning and decision-making in the face of a changing climate.
Statisticians are in general agreement that there are flaws in how science is currently practiced; there is less agreement in how to make repairs. Our prescription for a Post-p < 0.05 Era is to develop and teach courses that expand our view of what constitutes the domain of statistics and thereby bridge undergraduate statistics coursework and the graduate student experience of applying statistics in research. Such courses can speed up the process of gaining statistical wisdom by giving students insight into the human propensity to make statistical errors, the meaning of a single test within a research project, ways in which p-values work and don't work as expected, the role of statistics in the lifecycle of science, and best practices for statistical communication. The course we have developed follows the story of how we use data to understand the world, leveraging simulation-based approaches to perform customized analyses and evaluate the behavior of statistical procedures. We provide ideas for expanding beyond the traditional classroom, two example activities, and a course syllabus as well as the set of statistical best practices for creating and consuming scientific information that we develop during the course.
SummaryA separation between the academic subjects statistics and mathematical statistics has existed in Sweden almost as long as there have been statistics professors. The same distinction has not been maintained in other countries. Why has it been kept for so long in Sweden, and what consequences may it have had?In May 2015, it was 100 years since Mathematical Statistics was formally established as an academic discipline at a Swedish university where Statistics had existed since the turn of the century.We give an account of the debate in Lund and elsewhere about this division during the first decades after 1900 and present two of its leading personalities. The Lund University astronomer (and mathematical statistician) C. V. L. Charlier was a leading proponent for a position in mathematical statistics at the university. Charlier's adversary in the debate was Pontus Fahlbeck, professor in political science and statistics, who reserved the word statistics for ‘statistics as a social science’. Charlier not only secured the first academic position in Sweden in mathematical statistics for his former PhD student Sven Wicksell but also demonstrated that a mathematical statistician can be influential in matters of state, finance as well as in different natural sciences. Fahlbeck saw mathematical statistics as a set of tools that sometimes could be useful in his brand of statistics.After a summary of the organisational, educational and scientific growth of the statistical sciences in Sweden that has taken place during the last 50 years, we discuss what effects the Charlier–Fahlbeck divergence might have had on this development.
Non-stationary time series modelling is applied to long tidal records from Esbjerg, Denmark, and coupled to climate change projections for sea-level and storminess, to produce projections of likely future sea-level maxima. The model has several components: nonstationary models for mean sea-level, tides and extremes of residuals of sea level above tide level. The extreme value model (at least on an annual scale) has location-parameter dependent on mean sea level. Using the methodology of Bolin et al. (2015)and Guttorp et al. (2014) we calculate, using CMIP5 climate models, projections for mean sea level with attendant uncertainty. We simulate annual maxima in two ways - one method uses the empirically fitted non-stationary generalized extreme-value-distribution (GEV) of 20th-century annual maxima projected forward based on msl-projections, and the other has a stationary approach to extremes. We then consider return levels and the increase in these from AD 2000 to AD 2100. We find that the median of annual maxima with return period 100 years, taking into account all the nonstationarities, in year 2100 is 6.5m above current mean sea level (start of 21stC) levels
Stochastic simulation has played an important role in understanding hematopoiesis, but implementing and interpreting mathematical models requires a strong statistical background, often preventing their use by many clinical and translational researchers. Here, we introduce a user-friendly graphical interface with capabilities for visualizing hematopoiesis as a stochastic process, applicable to a variety of mammal systems and experimental designs. We describe the visualization tool and underlying mathematical model, and then use this to simulate serial transplantations in mice, human cord blood cell expansion, and clonal hematopoiesis of indeterminate potential. The outcomes of these virtual experiments challenge previous assumptions and provide examples of the flexible range of hypotheses easily testable via the visualization tool.