Implementation of many statistical methods for large, multivariate data sets requires one to solve a linear system that, depending on the method, is of the dimension of the number of observations or each individual data vector. This is often the limiting factor in scaling the method with data size and complexity. In this paper we illustrate the use of Krylov subspace methods to address this issue in a statistical solution to a source separation problem in cosmology where the data size is prohibitively large for direct solution of the required system. Two distinct approaches, adapted from techniques in the literature, are described: one that uses the method of conjugate gradients directly to the Kronecker-structured problem and another that reformulates the system as a Sylvester matrix equation. We show that both approaches produce an accurate solution within an acceptable computation time and with practical memory requirements for the data size that is currently available.
A method for conducting Bayesian elicitation and learning in risk assessment is presented. It assumes that the risk process can be described as a fault tree. This is viewed as a belief network, for which prior distributions on primary event probabilities are elicited by means of a pairwise comparison approach. A Bayesian updating procedure, following observation of some or all of the events in the fault tree, is described. The application is illustrated through the motivating example of risk assessment of spacecraft explosion during controlled re-entry.
This paper presents a new way to determine road profile and detect bridge damage using accelerations from a fleet of passing vehicles. Using off-bridge data, a Bayesian approach updates estimates of the road profile and vehicle properties. The profile elevations and vehicle properties are shown to be insensitive to random noise in acceleration measurements. On-bridge data, with recently updated vehicle properties, are used to estimate bridge damage. Bearing damage and local crack damage in a bridge are simulated. For bearing damage, the results show that this method can quantify the damage level of a bearing and infer other bridge properties. For local crack damage, the levels and the location of the damage are inferred from the simulated measurements.
Global estimates of the number of species of Fungi have ranged from 1.5 to 13.2 million, but have been based more on opinion and simple ratios than quantitative assessment. We analysed trends in the rate of description of fungal species over four centuries, noted the use of molecular methods in species delimitation, and used a statistical model designed for such data to predict future trends. A total of 144,035 fungal species were analysed, along with smaller species groups extracted from the core dataset that approximated biological and ecological traits. The groups explored included fungi of medical significance (728 spp), those associated with the marine environment (972 spp), rust and smut fungi (9,125 spp), arthropod ectoparasites of class Laboulbeniomycetes (2,376 spp), mushroom-forming fungi of class Agaricomycetes (37,717 spp), the budding yeasts of subphylum Saccharomycotina (1,165 spp), the class Dothideomycetes (30,912 spp), and lichenized fungi of classes Lecanoromycetes and Arthoniomycetes (12,154 spp). There was an acceleration in overall fungal description rates within the last two decades accompanied by the increased use of genetic data in new species descriptions. Mushroom-forming, lichenized, and plant-associated fungi were predicted to experience the greatest increase in new species. Increased description rates are supported by an increase in the number of authors describing species. However, the number of species described per author in a year has been declining since 1875. Because less than 10% of currently accepted fungal species have molecular data associated with corresponding type specimens, genetic data should not be used to discriminate new species without associated phenotypic information. An additional 68,750 species (48%) were predicted to be described this century, making Fungi the least well-described Kingdom assessed to date.
Taxonomic species are the best standardised metric of biodiversity. Therefore, there is broad scientific and public interest in how many species have already been named and how many more may exist. Crustaceans comprise about 6% of all named animal species and isopods about 15% of all crustaceans. Here, we review progress in the naming of isopods in relation to the number of people describing new species and estimate how many more species may yet be named by 2050 and 2100, respectively. In over two and a half centuries of discovery, 10,687 isopod species in 1,557 genera and 141 families have been described by 755 first authors. The number of authors has increased over time, especially since the 1950s, indicating increasing effort in the description of new species. Despite that the average number of species described per first author has declined since the 1910s, and the description rate has slowed down over the recent decades. Authors' publication lifetimes did not change considerably over time, and there was a distinct shift towards multi-authored publications in recent decades. Estimates from a non homogeneous renewal process model predict that an additional 660 isopod species will be described by 2100, assuming that the rate of description continues at its current pace.
Even though it is a global marine biodiversity hotspot, the contribution of Indonesia to marine species discovery has been disproportionately small, and despite the amount of biodiversity research conducted by local scientists, many species in this country remain undescribed. In this article, we used the discovery rate of Indonesian polychaete species as a case example to investigate the contributory factors leading to the slow rate of marine species discovery in the country. In addition, we evaluated ecological studies on Indonesian polychaetes and enumerated the number of local marine taxonomists along with the number of Indonesian species that they described. We found that throughout Indonesia's history, the country only had a few marine taxonomists and that past marine species discoveries have been largely dependent on overseas scientists. This has been the primary factor causing the slow rate of marine species discovery in the country and has led to limited taxonomic literature on local species, resulting in many species either remaining unidentified or incorrectly identified, as local researchers typically used identification keys written for other geographic regions. We further found that limited access to natural history collections, uneven distribution of research facilities, a lack of collaborative research and funding, and strict requirements for research permits for foreign researchers have also hampered discoveries of marine species new to science in Indonesia. Long term recruitment of local marine taxonomists through the job vacancies regularly offered (annually) by the Indonesian government along with taxonomic training by relevant governmental research institutions are the first immediate actions to address these problems. Moreover, collaborative work and the establishment of Indonesian marine reference collections, databases, and identification keys are other strategies to end taxonomic impediments in the nation. For polychaetes, provided that such actions are made, our model forecasts that between 120 and 270 more Indonesian species will be discovered by the end of this century. Any taxonomic investigations conducted in the Coral Triangle are likely to uncover greater numbers of undescribed marine species compared to any other location in Indonesia or the world.
Abstract This article illustrates why a fully Bayesian approach to Fault Tree Analysis can be useful, describes how it can be performed, and discusses one of the important considerations to be made regarding such analysis, namely, prior elicitation. It then describes an easy‐to‐use prior robustness approach for Fault Tree Analysis to quantify the sensitivity of the prior elicitation of the failure probability to misspecification of uncertainty in elementary events.
The paper discusses issues that surround decisions in risk and reliability, with a major emphasis on quantitative methods. We start with a brief history of quantitative methods in risk and reliability from the 17th century onwards. Then, we look at the principal concepts and methods in decision theory. Finally, we give several examples of their application to a wide variety of risk and reliability problems: software testing, preventive maintenance, portfolio selection, adversarial testing, and the defend-attack problem. These illustrate how the general framework of game and decision theory plays a relevant part in risk and reliability.
This work uses hierarchical logistic Gaussian processes to infer true redshift distributions of samples of galaxies, through their cross-correlations with spatially overlapping spectroscopic samples. We demonstrate that this method can accurately estimate these redshift distributions in a fully Bayesian manner jointly with galaxy-dark matter bias models. We forecast how systematic biases in the redshift-dependent galaxy-dark matter bias model affect redshift inference. Using published galaxy-dark matter bias measurements from the Illustris simulation, we compare these systematic biases with the statistical error budget from a forecasted weak gravitational lensing measurement. If the redshift-dependent galaxy-dark matter bias model is mis-specified, redshift inference can be biased. This can propagate into relative biases in the weak lensing convergence power spectrum on the 10-30 per cent level. We, therefore, showcase a methodology to detect these sources of error using Bayesian model selection techniques. Furthermore, we discuss the improvements that can be gained from incorporating prior information from Bayesian template fitting into the model, both in redshift prediction accuracy and in the detection of systematic modelling biases.
The number of species that exist on Earth has been an intriguing question in ecology and evolution. For marine species, previous works have analysed trends in the discovery of extant species, without comparison to the fossil record. Here, we compared the rate of description between extant and fossil species of the same group of marine invertebrates, Bryozoa. There are nearly 3 times as many described fossil species as there are extant species. This indicates that current biodiversity represents only a small proportion of Earth's past biodiversity, at least for Bryozoa. Despite these differences, our results showed similar trends in the description of new species between extant and fossil groups. There has been an increase in taxonomic effort during the past century, characterized by an increase in the number of taxonomists, but no change in their relative productivity (i.e. similar proportions of authors described most species). The 20th century had the most species described per author, reflecting increased effort in exploration and technological developments. Despite this progress, future projections in the discovery of bryozoan species predict that around 10 and 20% more fossil and extant species than named species, respectively, will be discovered by 2100, representing 2430 and 1350 more fossil and extant species, respectively. This highlights the continued need for both new species descriptions and taxonomic revisions, as well as ecological and biogeographical research, to better understand the biodiversity of Bryozoa.
Despite the availability of well-documented data, a comprehensive review of the discovery progress of polychaete worms (Annelida) has never been done. In the present study, we reviewed available data in the World Register of Marine Species, and found that 11,456 valid species of Recent polychaetes (1417 genera, 85 families) have been named by 835 first authors since 1758. Over this period, three discovery phases of the fauna were identified. That is, the initial phase (from 1758 to mid-nineteenth century) where nearly 500 species were described by few taxonomists, the second phase (from the 1850’s to mid-twentieth century) where almost 5000 species were largely described by some very productive taxonomists, and the third phase (from the 1950’s to modern times) in which about 6000 species were described by the most taxonomists ever. Six polychaete families with the most species were Syllidae (993 species), Polynoidae (876 species), Nereididae (687 species), Spionidae (612 species), Terebellidae (607 species) and Serpulidae (576 species). The increase in the number of first authors through time indicated greater taxonomic effort. By contrast, there was a decline in the number of polychaete species described in proportion to the number of first authors since around mid-nineteenth century. This suggested that it has been getting more difficult to find new polychaete species. According to our modelling, we predict that 5200 more species will be discovered between now and the year 2100. The total number of polychaete species of the world by the end of this century is thus anticipated to be about 16,700 species.
Starting in the late 80s Bayesian methods have gained increasing attention in the reliability literature. The focus of most of the earlier Bayesian work in reliability involved statistical inference and thus the main emphasis was on modeling and analysis. Advances in Bayesian computing after the 90's have significantly contributed not only to the use of Bayesian inference and prediction but also to the implementation of Bayesian decision-theoretic approaches in reliability problems. In this review we present an overview of Bayesian methods to solve decision problems in reliability some of which involve two or more decision makers with conflicting objectives. We consider problems in areas such as design, life testing, preventive maintenance, reliability certification, or warranty policies. In doing so, we present key aspects of the decision problems, give a brief review of earlier methods and finally discuss recent advances in Bayesian approaches to solve them. (C) 2019 Elsevier B.V. All rights reserved.
In this paper we present a Bayesian competing risk proportional hazards model to describe mortgage defaults and prepayments. We develop Bayesian inference for the model using Markov chain Monte Carlo methods. Implementation of the model is illustrated using actual default/prepayment data and additional insights that can be obtained from the Bayesian analysis are discussed.
The symposia on Games and Decisions in Reliability and Risk (GDRR) have been running since 2009, when the first meeting was held at George Washington University in Washington, DC. They have been held every 2 years since then, with the 6th symposium of the series making a return to its first venue at the end of May 2019. The focus of the meetings is to present research on the application of game and decision theory methods to reliability and risk analysis across diverse disciplines such as economics, engineering, finance, medical sciences, and transportation. This special issue of Applied Stochastic Models in Business and Industry has arisen out of the 5th workshop that was held in 2017 in Madrid, Spain. The 2-day meeting held application-focussed sessions on cybersecurity, environmental risks, fraud, and health risk, and more methodologically focussed sessions on decision theory, graphical models, and reliability theory. There are 5 papers in this special issue that cover several of the topics covered at the meeting. Balbás and Charron look at the important idea of the value at risk, a risk measure that is seeing increasing use in financial regulation. They present some theory on its properties in “ambiguous” settings, where lack of data or measurement errors mean that one is deriving value at risk with imperfect knowledge. Gallego et al look at an application of decision theory to allocate expenditure on advertising so as to maximize sales. This makes use of structural time series models, fitted via a Bayesian learning approach and has a general risk-based application to investment allocation. Two papers focus on risk management issues in cybersecurity. Miaoui et al adopt a strategic approach describing a model that supports the investment decisions of an organization in security controls, cyberinsurance, and forensic tools. Naveiro et al adopt an operational approach, looking at questions of network safety and monitoring, and how to make forecasts of potential issues in a network at large scale, using variants of dynamic linear models. They describe a framework for time series forecasting across a large number of nodes in a network that is entirely automatic. Finally, Najem and Coolen examine the effect of swapping components of the same type, upon failure of one of them, on the survival signature of a system, as a means of improving its resilience. They also analyze component importance. This strategy is an alternative to other approaches such as increasing redundancy, component reliability, or maintenance activities. We would like to thank the special issue authors for their interesting work and look forward to continuing the GDRR symposium series tradition into the future.
This paper reports on an innovative human–machine interaction methodology adopted to assess the case, role and requirements for a new ground collision awareness technology. Specifically, this paper reports on the analysis of ground collision incident data and the subsequent advancement of user scenarios and bow-ties based on this data analysis, for the purpose of generating preliminary user and design requirements for this technology. In so doing, the requirements elicitation and validation methods used in this research are framed from an epistemological perspective. Accordingly, the particular methods adopted are presented and discussed in terms of concepts of evidence, bearing witness and the distinction between facts and values. As such, this paper promotes thinking about evidence-based design practices. Overall, this evidence-based approach aims to improve the development of scenarios and associated problem solving around technology cases, user requirements and user interface design features. The proposed method is useful in terms of bridging the gap from data analysis to design, and validating design decisions. In this regard, it is argued that the generation of user scenarios based on the analysis of incident data (i.e. data coding and statistical analysis), and the reframing of such scenarios in terms of bow-ties for the purpose of requirements/design envisionment, extends existing scenario-based design approaches. Although the use of bow-ties is not new, the advancement of bow-ties from data-driven scenarios is. Specifically, the bow-tie method was applied in a design context, to support problem solving around design decisions, as opposed to formal risk analysis.
A method for sequential Bayesian inference of the static parameters of a dynamic state space model is proposed. The method is able to use any valid approximation to the filtering and prediction densities of the state process. It computes the posterior distribution of the static parameters on a discrete grid that tracks the support dynamically. For inference of the state process, the Kalman filter and its extensions as well as cubature filtering have been used. It is illustrated with several examples including the stochastic volatility model and the challenging Kitagawa model and is compared to both online and offline methods. It is shown to provide a good trade off between speed and performance.
At present, amphipod crustaceans comprise 9,980 species, 1,664 genera, 444 subfamilies, and 221 families. Of these, 1,940 species (almost 20%) have been discovered within the last decade, including 18 fossil records for amphipods, which mostly occurred in Miocene amber and are probably all freshwater species. There have been more authors describing species since the 1950s and fewer species described per author since the 1860s, implying greater taxonomic effort and that it might be harder to find new amphipod species, respectively. There was no evidence of any change in papers per author or publication life-times of taxonomists over time that might have biased apparent effort. Using a nonhomogeneous renewal process model, we predicted that by the year 2100, 5,600 to 6,600 new amphipod species will be discovered. This indicates that about two-thirds of amphipods remain to be discovered which is twice the proportion than for species overall. Amphipods thus rank amongst the least well described taxa. To increase the prospect of discovering new amphipod species, studying undersampled areas and benthic microhabitats are recommended.
We propose a prior robustness approach for the Bayesian implementation of the fault tree analysis (FTA). FTA is often used to evaluate risk in large, safety critical systems but has limitations due to its static structure. Bayesian approaches have been proposed as a superior alternative to it, however, this involves prior elicitation, which is not straightforward. We show that minor misspecification of priors for elementary events can result in a significant prior misspecification for the top event. A large amount of data is required to correctly update a misspecified prior and such data may not be available for many complex, safety critical systems. In such cases, prior misspecification equals posterior misspecification. Therefore, there is a need to develop a robustness approach for FTA, which can quantify the effects of prior misspecification on the posterior analysis. Here, we propose the first prior robustness approach specifically developed for FTA. We not only prove a few important mathematical properties of this approach, but also develop easy to use Monte Carlo sampling algorithms to implement this approach on any given fault tree with and and/or or gates. We then implement this Bayesian robustness approach on two real-life examples: a spacecraft re-entry example and a feeding control system example. We also provide a step-by-step illustration of how this approach can be applied to a real-life problem.