We introduce Simulation-Based Imaging (SBI), a framework for non-destructive acoustic imaging in which machine learning models trained entirely on simulated data serve as real-time solvers for the acoustic inverse problem. A high-fidelity nodal Discontinuous Galerkin forward solver generates large training datasets by randomizing inclusion geometry within a unit-cube domain; a 2D convolutional neural network then learns a direct mapping from boundary pressure measurements to a 32 by 32 by 32 voxel reconstruction of the interior. The trained model reliably recovers inclusion position and size from 144 boundary sensors with no prior knowledge of inclusion count or geometry. Reconstruction error degrades by only 13
Abstract. Fire affects soil and vegetation, which in turn can promote the initiation and growth of runoff-generated debris flows in steep watersheds. Postfire hazard assessments often focus on identifying the most likely watersheds to produce debris flows, quantifying rainfall intensity-duration thresholds for debris flow initiation, and estimating the volume of potential debris flows. This work seeks to expand on such analyses and forecast downstream debris flow runout and peak flow depth. Here, we report on a high-fidelity computational framework that enables debris flow simulation over two watersheds and the downstream alluvial fan, although at significant computational cost. We then develop a Gaussian Process surrogate model, allowing for rapid prediction of simulator outputs for untested scenarios. With a modest training of debris flow simulations, this surrogate is able to approximate peak flow depth with a mean squared error that is generally in the range of 0.1–0.2 m. We utilize this framework to explore model sensitivity to rainfall intensity and sediment availability as well as parameters associated with saturated hydraulic conductivity, hydraulic roughness, grain size, and sediment entrainment. Simulation results are most sensitive to hydraulic roughness and grain size. Further, we use this approach to examine variations in debris flow inundation patterns at different stages of postfire recovery, and we find that the area inundated by postfire flows decreases substantially over a time period as short as 9 months. In this case, we also see that temporal changes in hydraulic roughness and grain size following fire would be particularly beneficial for forecasting debris flow runout throughout the postfire recovery period. The emulator methodology presented here also provides a means to compute the probability of a debris flow inundating a specific downstream region, consequent to a forecast or design rainstorm. This workflow could be employed in prefire scenario-based planning or postfire hazard assessments.
Fire affects soil and vegetation, which in turn can promote the initiation and growth of runoff-generated debris flows in steep watersheds. Postfire hazard assessments often focus on identifying the most likely watersheds to produce debris flows, quantifying rainfall intensity-duration thresholds for debris flow initiation, and estimating the volume of potential debris flows. This work seeks to expand on such analyses and forecast downstream debris flow runout and peak flow depth. Here, we report on a high fidelity computational framework that enables debris flow simulation over two watersheds and the downstream alluvial fan, although at significant computational cost. We also develop a Gaussian Process surrogate model, allowing for rapid prediction of simulator outputs for untested scenarios. We utilize this framework to explore model sensitivity to rainfall intensity and sediment availability as well as parameters associated with saturated hydraulic conductivity, hydraulic roughness, grain size, and sediment entrainment. Simulation results are most sensitive to peak rainfall intensity and hydraulic roughness. We further use this approach to examine variations in debris flow inundation patterns at different stages of postfire recovery. Sensitivity analysis indicates that constraints on temporal changes in hydraulic roughness, saturated hydraulic conductivity, and grain size following fire would be particularly beneficial for forecasting debris flow runout throughout the postfire recovery period. The emulator methodology presented here also provides a means to compute the probability of a debris flow inundating a specific downstream region, consequent to a forecast or design rainstorm. This workflow could be employed in scenario-based planning for postfire hazard mitigation.
Inference, optimization, and inverse problems are but three examples of mathematical operations that require the repeated solution of a complex system of mathematical equations. To this end, surrogates are often used to approximate the output of these large computer simulations, providing fast and cheap approximation solutions. Statistical emulators are surrogates that, in addition to predicting the mean behavior of the system, provide an estimate of the error in that prediction. Classical Gaussian stochastic process emulators predict scalar outputs based on a modest number of input parameters. Making predictions across a space-time field of input variables is not feasible using classical Gaussian process methods. Parallel partial emulation is a new statistical emulator methodology that predicts afield of outputs based on the input parameters. Parallel partial emulation is constructed as a Gaussian process in parameter space, but no correlation among space or time points is assumed. Thus the computational work of parallel partial emulation scales as the cube of the number of input parameters (as traditional Gaussian Process emulation) and linearly with a space-time grid. The numerical methods used in numerical simulations are often designed to exploit properties of the equations to be solved. For example, modern solvers for hyperbolic conservation laws satisfy conservation at each time step, insuring overall conservation of the physical variables. Similarly, symplectic methods are used to solve Hamiltonian problems in physics. It is of interest, then, to study whether parallel partial emulation predictions inherit properties possessed by the simulation outputs. Does an emulated solution of a conservation law preserve the conserved quantities? Does an emulator of a Hamiltonian system preserve the energy? This paper investigates the properties of emulator predictions, in the context of systems of partial differential equations. We study conservation properties for three different kinds of equations-conservation laws, reaction-diffusion systems, and a Hamiltonian system. We also investigate the effective convergence, in parameter space, of the predicted solution of a highly nonlinear system modeling shape memory alloys.
Ongoing resurgence affects Campi Flegrei caldera (Italy) via bradyseism, i.e. a series of ground deformation episodes accompanied by increases in shallow seismicity. In this study, we perform a mathematical analysis of the GPS and seismic data in the instrumental catalogs from 2000 to 2020, and a comparison of them to the preceding data from 1983 to 1999. We clearly identify and characterize two overlying trends, i.e. a decennial-like acceleration and cyclic oscillations with various periods. In particular, we show that all the signals have been accelerating since 2005, and 90–97% of their increase has occurred since 2011, 40–80% since 2018. Nevertheless, the seismic and ground deformation signals evolved differently—the seismic count increased faster than the GPS data since 2011, and even more so since 2015, growing faster than an exponential function The ground deformation has a linearized rate slope, i.e. acceleration, of 0.6 cm/yr 2 and 0.3 cm/yr 2 from 2000 to 2020, respectively for the vertical (RITE GPS) and the horizontal (ACAE GPS) components. In addition, all annual rates show alternating speed-ups and slow-downs, consistent between the signals. We find seven major rate maxima since 2000, one every 2.8–3.5 years, with secondary maxima at fractions of the intervals. A cycle with longer period of 6.5–9 years is also identified. Finally, we apply the probabilistic failure forecast method, a nonlinear regression that calculates the theoretical time limit of the signals going to infinity (interpreted here as a critical state potentially reached by the volcano), conditional on the continuation of the observed nonlinear accelerations. Since 2000, we perform a retrospective analysis of the temporal evolution of these forecasts which highlight the periods of more intense acceleration. The failure forecast method applied on the seismic count from 2001 to 2020 produces upper time limits of [0, 3, 11] years (corresponding to the 5th, 50th and 95th percentiles, respectively), significantly shorter than those based on the GPS data, e.g. [0, 6, 21] years. Such estimates, only valid under the model assumption of continuation of the ongoing decennial-like acceleration, warn to keep the guard up on the future evolution of Campi Flegrei caldera.
Using the failure forecast method we describe a first assessment of failure time on present-day unrest signals at Campi Flegrei caldera (Italy) based on the horizontal deformation data collected in [2011, 2020] at eleven GPS stations.
In this work we couple the Metropolis-Hastings algorithm with the volcanic ash transport model Tephra2, and present the coupled algorithm as a new method to estimate the Eruption Source Parameters of volcanic eruptions based on mass per unit area or thickness measurements of tephra fall deposits. Outputs of the algorithm are presented as sample posterior distributions for variables of interest. Basic elements in the algorithm and how to implement it are introduced. Experiments are done with synthetic datasets. These experiments are designed to demonstrate that the algorithm works from different perspectives, and to show how inputs affect its performance. Advantages of the algorithm are that it has the ability to i) incorporate prior knowledge; ii) quantify the uncertainty; iii) capture correlations between variables of interest in the estimated Eruption Source Parameters; and iv) no simplification is assumed in sampling from the posterior probability distribution. A limitation is that some of the inputs need to be specified subjectively, which is designed intentionally such that the full capacity of the Bayes’ rule can be explored by users. How and why inputs of the algorithm affect its performance and how to specify them properly are explained and listed. Correlation between variables of interest in the posterior distributions exists in many of our experiments. They can be well-explained by the physics of tephra transport. We point out that in tephra deposit inversion, caution is needed in attempting to estimate Eruption Source Parameters and wind direction and speed at each elevation level, because this could be unnecessary or would increase the number of variables to be estimated, and these variables could be highly correlated. The algorithm is applied to a mass per unit area dataset of the tephra deposit from the 2011 Kirishima-Shinmoedake eruption. Simulation results from Tephra2 using posterior means from the algorithm are consistent with field observations, suggesting that this approach reliably reconstructs Eruption Source Parameters and wind conditions from deposits.
In metal catalytic design, there is a well-established linear scaling relationship between reaction and adsorption energies. However, owing to the challenges of performing experimental and/or computational experiments, there is a paucity of empirical data regarding these systems. In particular, there is little experimental evidence suggesting how the linear scaling law might be overcome in order to discover catalysts with more desirable properties. In this paper, we employ machine-learning techniques in order to predict reaction and adsorption energies for 300 hypothetical binary compounds. We then apply outlier detection methods to identify which of these predicted compounds do not follow the known scaling law. These outlier compounds, which would not have been identified through traditional design rules, are the most likely to have unexpected and potentially transformative catalytic behavior. Thus, this paper proposes a data-driven screening methodology to identify those metallic compounds (as a function of gaseous environment) which are most likely to have targeted catalytic behavior.
In addition to student assessment, curriculum assessment is a critical element to any pedagogy. It helps the educator assess the teaching of concepts, determine what may be lacking, and make changes for continual improvement. Meaningful assessment can be complicated when disciplines converge or when new approaches are implemented. To facilitate this, we present a network-based visualization schema to represent a materials informatics curriculum that combines materials science and data science concepts. We analyze the curriculum using network representations and relevant concepts from graph theory. This reveals established connections, linkages between materials science and data science, and the extent to which different concepts are connected. We also describe how some materials science topics are introduced from a data perspective, and present an illustrative case study from the curriculum.
In this work we couple the Metropolis-Hastings algorithm with the volcanic ash transport model TEPHRA2, and present the coupled algorithm as a new method to estimate the Eruption Source Parameters of volcanic eruptions based on mass per unit area or thickness measurements of tephra fall deposits. Basic elements in the algorithm and how to implement it are introduced. Experiments are done with synthetic datasets. These experiments are designed to demonstrate that the algorithm works, and to show how inputs affect its performance. Results are presented as sample posterior distribution estimates for variables of interest. Advantages of the algorithm are that it has the ability to i) incorporate prior knowledge; ii) quantify the uncertainty; and iii) capture correlations between variables of interest in the estimated Eruption Source Parameters. A limitation is that some of the inputs need to be specifed subjectively. How and why such inputs affect the performance of the algorithm and how to specify them properly are explained and listed. Correlation between variables of interest are well-explained by the physics of tephra transport. We point out that in tephra deposit inversion, caution is needed in attempting to estimate Eruption Source Parameters, and wind direction and speed at each elevation level, as this increases the number of variables to be estimated. The algorithm is applied to a mass per unit area dataset of the tephra deposit from the 2011 Kirishima-Shinmoedake eruption. Simulation results from TEPHRA2 using posterior means from the algorithm are consistent with field observations, suggesting that this approach reliably reconstructs Eruption Source Parameters and wind conditions from the deposit.
We advocate here a methodology for characterizing models of geophysical flows and the modeling assumptions they represent, using a statistical approach over the full range of applicability of the models. Such a characterization may then be used to decide the appropriateness of a model and modeling assumption for use. We present our method by comparing three different models arising from different rheology assumptions, and the output data show unambiguously the performance of the models across a wide range of possible flow regimes. This comparison is facilitated by the recent development of the new release of our TITAN2D mass flow code that allows choice of multiple rheologies. The quantitative and probabilistic analysis of contributions from different modeling assumptions in the models is particularly illustrative of the impact of the assumptions. Knowledge of which assumptions dominate, and, by how much, is illustrated in the topography on the SW slope of Volcán de Colima (MX). A simple model performance evaluation completes the presentation.
Statistical emulators are a key tool for rapidly producing probabilistic hazard analysis of geophysical processes. Given output data computed for a relatively small number of parameter inputs, an emulator interpolates the data, providing the expected value of the output at untried inputs and an estimate of error at that point. In this work, we propose to fit Gaussian Process emulators to the output from a volcanic ash transport model, Ash3d. Our goal is to predict the simulated volcanic ash thickness from Ash3d at a location of interest using the emulator. Our approach is motivated by two challenges to fitting emulators-characterizing the input wind field and interactions between that wind field and variable grain sizes. We resolve these challenges by using physical knowledge on tephra dispersal. We propose new physically motivated variables as inputs and use normalized output as the response for fitting the emulator. Subsetting based on the initial conditions is also critical in our emulator construction. Simulation studies characterize the accuracy and efficiency of our emulator construction and also reveal its current limitations. Our work represents the first emulator construction for volcanic ash transport models with considerations of the simulated physical process.
Intra-tumor and inter-patient heterogeneity are two challenges in developing mathematical models for precision medicine diagnostics. Here we review several techniques that can be used to aid the mathematical modeller in inferring and quantifying both sources of heterogeneity from patient data. These techniques include virtual populations, nonlinear mixed effects modeling, non-parametric estimation, Bayesian techniques, and machine learning. We create simulated virtual populations in this study and then apply the four remaining methods to these datasets to highlight the strengths and weak-nesses of each technique. We provide all code used in this review at https://github.com/jtnardin/Tumor-Heterogeneity/ so that this study may serve as a tutorial for the mathematical modelling community. This review article was a product of a Tumor Heterogeneity Working Group as part of the 2018-2019 Program on Statistical, Mathematical, and Computational Methods for Precision Medicine which took place at the Statistical and Applied Mathematical Sciences Institute.
Effective volcanic hazard management in regions where populations live in close proximity to persistent volcanic activity involves understanding the dynamic nature of hazards, and associated risk. Emphasis until now has been placed on identification and forecasting of the escalation phase of activity, in order to provide adequate warning of what might be to come. However, understanding eruption hiatus and post-eruption unrest hazards, or how to quantify residual hazard after the end of an eruption, is also important and often key to timely post-eruption recovery. Unfortunately, in many cases when the level of activity lessens, the hazards, although reduced, do not necessarily cease altogether. This is due to both the imprecise nature of determination of the “end” of an eruptive phase as well as to the possibility that post-eruption hazardous processes may continue to occur. An example of the latter is continued dome collapse hazard from lava domes which have ceased to grow, or sector collapse of parts of volcanic edifices, including lava dome complexes. We present a new probabilistic model for forecasting pyroclastic density currents (PDCs) from lava dome collapse that takes into account the heavy-tailed distribution of the lengths of eruptive phases, the periods of quiescence, and the forecast window of interest. In the hazard analysis, we also consider probabilistic scenario models describing the flow’s volume and initial direction. Further, with the use of statistical emulators, we combine these models with physics-based simulations of PDCs at Soufrière Hills Volcano to produce a series of probabilistic hazard maps for flow inundation over 5, 10, and 20 year periods. The development and application of this assessment approach is the first of its kind for the quantification of periods of diminished volcanic activity. As such, it offers evidence-based guidance for dome collapse hazards that can be used to inform decision-making around provisions of access and reoccupation in areas around volcanoes that are becoming less active over time.
Iltra-tumor and inter patient heterogeneity are two challenges in developing mathematical models for precision medicine diagnostics. Here we review several techniques that can he used to aid the mathematical modeller in inferring and quantifying both sources of heterogeneity from patient data. These techniques include virtual populations, nonlinear mixed effects modeling, non -parametric estimation, Bayesian techniques, and machine learning. We create simulated virtual populations in this study and then apply the four remaining methods to these datasets to highlight the strengths and weaknesses of each technique. We provide all code used in this review at https://github.com/jtnardin/TumorHeterogeneity/ so that this study may serve as a tutorial for the mathematical modelling community. This review article was a product of a Tumor Heterogeneity Working Group as part of the 2018-2019 Program on Statistical, Mathematical, and Computational Methods for Precision Medicine which took place at the Statistical and Applied Mathematical Sciences Institute.
We present two models using monitoring data in the production of volcanic eruption forecasts. The first model enhances the well-established failure forecast method introducing an SDE in its formula...