The reliability analysis of complex systems is crucial for cost-effective evaluations, particularly when using deterministic black-box models. This study examines system performance under uncertainty, where the input vector x is an element of X subset of Rd defines both system and environmental conditions, and failure is characterized by F = {x is an element of X divided by y(x) s star}, with y the variable of interest and s star is an element of R a given threshold. Since x is uncertain, a probabilistic analysis is required to ensure robust safety assessments. Such an analysis typically involves two key steps: first, estimating the system's probability of failure (noted pf), and then, evaluating it against safety standards or expert knowledge. While considerable effort has been invested in proposing efficient methods for estimating pf, little attention has been paid to the decision phase, which should take into account the uncertainties. This work focuses on the definition and use of safety margins in system reliability analysis with a final decision making purpose, especially when the knowledge of the input vectors x is limited to a finite set of n observations. A key distinction is made between cases where n is large or small relative to 1/pf. The main contributions of the paper focus on scenarios with small n and propose two approaches for defining reasonable safety margins. The first estimates the probability distribution of x, while the second, based on extreme value theory, directly assesses the tail behavior of the output distribution. The proposed framework is validated through numerical case studies.
Numerical simulations are crucial for modeling complex systems, but calibrating them becomes challenging when data are noisy or incomplete and likelihood evaluations are computationally expensive. Bayesian calibration offers an interesting way to handle uncertainty, yet computing the posterior distribution remains a major challenge under such conditions. To address this, we propose a sequential surrogate-based approach that incrementally improves the approximation of the log-likelihood using Gaussian Process Regression. Starting from limited evaluations, the surrogate and its gradient are refined step by step. At each iteration, new evaluations of the expensive likelihood are added only at informative locations, that is to say where the surrogate is most uncertain and where the potential impact on the posterior is greatest. The surrogate is then coupled with the Metropolis-Adjusted Langevin Algorithm, which uses gradient information to efficiently explore the posterior. This approach accelerates convergence, handles relatively high-dimensional settings, and keeps computational costs low. We demonstrate its effectiveness on both a synthetic benchmark and an industrial application involving the calibration of high-speed train parameters from incomplete sensor data.
Calibration of sensors in partially observed and correlated environments raises fundamental challenges for variable selection and model interpretation. When observations are noisy and influenced by unmeasured factors, classical selection criteria based on regression coefficients, cross-validation errors, or sensitivity indices may fail to identify the variables that most effectively reduce uncertainty in the target quantity. This paper introduces a probabilistic framework for variable selection based on variance reduction and prediction stability. The proposed approach relies on the systematic evaluation of models built from different subsets of observed variables and on a criterion that minimizes conditional prediction variance under a parsimony constraint. This criterion naturally penalizes spurious correlations and distinguishes variables that contribute to uncertainty reduction from those that merely compensate for missing information. The framework is illustrated through analytical examples and numerical experiments on simulated data, which highlight its robustness to noise and unmeasured confounders. Although motivated by calibration problems in environmental sensing, the proposed methodology is general and applicable to a wide range of regression and inference tasks involving correlated inputs and unobserved factors.
Explaining the outcome of dynamical systems is non-trivial due to the temporal nature and correlation of the variables. In this work, we propose a novel framework of history-aware sensitivity analysis for stationary time-series to quantify different memory effects and clarify their roles. For this purpose, we decompose the output time series into non-correlated components, namely the instantaneous component and the memory components. The latter are sorted in decreasing order of variance to reflect the importance of the variables. We highlight the compensation phenomena between the resulting components and illustrate them in the case of independent variables in a linear setting. To enable history-aware explanations, variance-based sensitivity indices are derived from the obtained decomposition. We demonstrate the effectiveness of our methodology in providing insights to explain output time-series in both synthetic and real-world cases.
A sequential Bayesian inversion algorithm is proposed for the estimation of unknown parameters in computationally demanding physical models, a frequent challenge in industrial applications. The approach aims to accurately infer the posterior distribution of the parameters, ensuring that it is concentrated as possible around the true values while minimizing the number of direct model evaluations. To achieve this, a multi-fidelity meta-modeling strategy is employed within a reduced-order space, leveraging recursive Gaussian process regression to approximate the projection functions efficiently. The meta-model and the projection matrix are iteratively refined in a sequential way, by incorporating new training data selected based on the evolving posterior distribution, ensuring that only the most informative simulations are performed. This approach enables precise parameter estimation while controlling computational costs. Numerical and experimental case studies on the thermal modeling of a single-layer wall illustrates the method’s effectiveness in identifying thermal properties from four functional outputs. A comparative analysis with a fixed-cost identification strategy highlights the robustness of the proposed sequential method, demonstrating robust parameter estimates across iterations and a progressive reduction of uncertainty. By adaptively targeting the most valuable simulations, this algorithm efficiently balances computational cost and estimation accuracy, making it particularly well-suited for industrial applications involving expensive simulation codes.
Numerical simulations are widely used to predict the behavior of physical systems, with Bayesian approaches being particularly well suited for this purpose. However, experimental observations are necessary to calibrate certain simu-lator parameters for the prediction. In this work, we use a multi-output simulator to predict all its outputs, including those that have never been experimentally observed. This situation is referred to as the transposition context. To ac-curately quantify the discrepancy between model outputs and real data in this context, conventional methods cannot be applied, and the Bayesian calibration must be augmented by incorporating a joint model error across all outputs. To achieve this, the proposed method is to consider additional input parameters within a hierarchical Bayesian model, which includes hyperparameters for the prior distribution of the calibration variables. This approach is applied to a computer code with three outputs that models the Taylor cylinder impact test with a small number of observations. The outputs are considered as the observed variables one at a time, to work with three different transposition situations. The proposed method is compared with other approaches that embed model errors to demonstrate the significance of the hierarchical formulation.
Air and water pollution are major threats to public health, highlighting the need for reliable environmental monitoring. Low-cost multisensor systems are promising but suffer from limited selectivity, because their responses are influenced by non-target variables (interferents) such as temperature and humidity. This complicates pollutant detection, especially in data-driven models with noisy, correlated inputs. We propose a method for selecting the most relevant interferents for sensor calibration, balancing performance and cost. Including too many variables can lead to overfitting, while omitting key variables reduces accuracy. Our approach evaluates numerous models using a bias-variance trade-off and variance analysis. The method is first validated on simulated data to assess strengths and limitations, then applied to a carbon nanotube-based sensor array deployed outdoors to characterize its sensitivity to air pollutants.
Bayesian optimization algorithms form an important class of methods to minimize functions that are costly to evaluate, which is a very common situation. These algorithms iteratively infer Gaussian processes from past observations of the function and decide where new observations should be made through the maximization of an acquisition criterion. Often, the objective function is defined on a compact set such as in a hyper-rectangle of the $d$-dimensional real space, and the bounds are chosen wide enough so that the optimum is inside the search domain. In this situation, this work provides a way to integrate in the acquisition criterion the \textit{a priori} information that these functions, once modeled as GP trajectories, should be evaluated at their minima, and not at any point as usual acquisition criteria do. We propose an adaptation of the widely used Expected Improvement acquisition criterion that accounts only for GP trajectories where the first order partial derivatives are zero and the Hessian matrix is positive definite. The new acquisition criterion keeps an analytical, computationally efficient, expression. This new acquisition criterion is found to improve Bayesian optimization on a test bed of functions made of Gaussian process trajectories in low dimension problems. The addition of first and second order derivative information is particularly useful for multimodal functions.
While carbon nanotube sensor arrays are highly sensitive to multiple gases, they typically have moderate selectivity, major interferents and significant measurement noise. This makes the calibration of such systems deployed in uncontrolled environments a challenging task. This paper presents a Bayesian framework that addresses these difficulties. It is used to demonstrate the joint ozone and carbon monoxide prediction capability of a fully integrated 10x2 carbon nanotube-based chemistor array deployed for several weeks in outdoor conditions. For ozone, mean absolute error (MAE) was found to be 4ppb over the range 15 to 83 ppb with a detection limit at 5.1 ppb; for carbon monoxide, 22 ppb MAE over the range 7 to 8 ppm with a detection limit at 2 ppm.
In the world of connected automated objects, increasingly rich and structured data are collected daily (positions, environmental variables, etc.). In this work, we are interested in the characterization of the variability of the trajectories of one of these objects (robot, drone, or delivery droid for example) along a particular path from irregularly sampled data in time and space. To do so, we model the position of the considered object by a random field indexed in time, whose distribution we try to estimate (for risk analysis for example). This distribution being by construction concentrated on an unknown curve, two phases are proposed for its reconstruction: a phase of identification of this curve, by clustering and polynomial smoothing techniques, then a phase of statistical inference of the random field orthogonal to this curve, by spectral methods and kernel reconstructions. The efficiency of the proposed approach, both in terms of computation time and reconstruction quality, is illustrated on several numerical applications.
The railway world is undergoing major changes. The advent of new technologies allows us to rethink the train system and face new challenges, but one must not forget all the ecological constraints that are now accentuated by the increase in energy costs. This paper focuses on the optimization of the driver commands to limit the energy consumption of the trains under punctuality and security constraints. This problem falls within the framework of control optimization problems for nonlinear dynamic mechanical systems in the presence of constraints and uncertainties. A four-step approach is then proposed in this paper to solve this problem: (1) the introduction of simplified and fast-to-evaluate models to model the nonlinear dynamic behavior of the train and its energy consumption; (2) the identification in a Bayesian formalism of the parameters on which these models depend from on-track measurements on commercial trains; (3) the reformulation of the optimization problem so that it integrates the uncertainties related to an imperfect knowledge of these estimated parameters; (4) the resolution of the optimization problem using evolutionary algorithms. The main specificity of this work lies in the fact that not only the objective function to be minimized, here the energy consumed by the train, is impacted by the uncertainties, but also the admissibility constraints of the solution, here punctuality and operating safety. The integration of the uncertainties in the search for the control function is thus not trivial and requires several original adaptations in order to make the final optimization problem well posed.
In this study, a radio frequency (RF) humidity sensor for monitoring outdoor relative humidity (RH) was developed using polyethyleneimine as a sensing material. The sensor's performance was assessed in laboratory and outdoor conditions. The primary sensor demonstrated an exponential sensitivity in frequency and magnitude, approximately -3.65 MHz/%RH and -7.69 MHz/%RH, -0.051 dB/%RH, and -0.12 dB/%RH, across the RH ranges of 30%RH-50%RH and 50%RH-70%RH, respectively, with consistent results from 25 degree celsius to 45degree celsius. The sensor exhibited a response and recovery time of 22/44 s, with minimal hysteresis and marginal cross-sensitivity to temperature. Following calibration using different calibration methods (analytical and machine-learning-based techniques), the RH prediction performance was gauged on a separate dataset. The exponential calibration law achieved mean absolute errors (MAEs) as low as 0.8%RH, while other methods, such as Gaussian regression process, k-nearest neighbor (KNN), and generalized linear regression (GLR), resulted in MAE of 0.9%RH, 1%RH, and 1.1%RH. A secondary sensor, also exhibiting exponential sensitivity, underwent initial laboratory testing before transitioning to outdoor conditions. Applying laboratory calibration to outdoor data resulted in an MAE of 4.4%RH, showing remarkable transferability with only 0.8%RH increased MAE. Cross-sensitivity analysis revealed no dependency with temperature, NO2, NO, CO, CO2, and O-3 in outdoor conditions. Finally, we confirmed the transferability of calibration between the two sensors under laboratory and outdoor conditions, utilizing output standardization with the slope-bias correction algorithm. This research underscores the potential of our RF humidity sensor for accurate outdoor RH monitoring.
The presence of systematic interferents and high measurement uncertainties strongly diminishes performances of environmental sensor array when deployed in real environment. The present paper shows how to account for such perturbations in a Bayesian framework and demonstrates the increase in performances - up to 50% on simulated and 12% on experimental data.
Societal demands in the field of air and water pollution monitoring require the ability to simultaneously detect a wide variety of chemical elements at very low concentrations in complex environments using compact and low-cost sensor devices. Although nanomaterial-based sensors have long been proposed as a solution to these exacting requirements, their detection accuracy is generally degraded in real-world conditions compared to laboratory conditions due to the effects of various interferents. To manage the related uncertainties and to associate confidence to the estimations, it seems natural to formulate the calibration-estimation problem in a probabilistic framework. This probabilistic formulation and its successful application to the monitoring of pH and active chlorine in drinking water are the main contributions of this work. While other tested calibration methods only allowed the monitoring of active chlorine, our solution enables the monitoring of both active chlorine and pH with quantified and reasonable uncertainties. Its success relies mainly on two adaptations of standard calibration methods: the consideration of the sensor inputs measurement errors (and not only the errors associated with the sensor outputs as it is generally done), and the introduction of two sources of model error (ME), one accounting for the approximate character of the calibration model, the other for unmeasured—and possibly unknown—interferents in the calibration environment.
Xavier Descombes合作论文数INRIA Sophia Antipolis Mediterranee15