The research Chair OQUAIDO - for Optimisation et QUAntification d'Incertitudes pour les Donnees Onereuses in French - gathers academic and technological research partners to work on statistical learning problems involving scarce and error-prone data. This Chair, was created in January 2016 for a period of 5 years. It has tackled problems where small data is described by statistical models that, in turn, serve to characterize uncertainties, calibrate computer codes and search for optimal configurations. Many of the investigated approaches rely on Gaussian processes and confront mathematical challenges such as high dimension (even functional inputs / outputs), mixed continuous and categorical inputs, specific constraints and medium data. This activity report highlights noticeable scientific contributions of OQUAIDO, provides bibliography indicators and summarizes the events that have marked the research Chair life. It is concluded by a few lessons on collaborative projects at the intersection between mathematics and engineering that were learnt during these 5 years.
Gaussian processes (GP) are widely used as a metamodel for emulating time-consuming computer codes. We focus on problems involving categorical inputs, with a potentially large number L of levels (typically several tens), partitioned in G << L groups of various sizes. Parsimonious covariance functions, or kernels, can then be defined by block covariance matrices T with constant covariances between pairs of blocks and within blocks. We study the positive definiteness of such matrices to encourage their practical use. The hierarchical group/level structure, equivalent to a nested Bayesian linear model, provides a parameterization of valid block matrices T. The same model can then be used when the assumption within blocks is relaxed, giving a flexible parametric family of valid covariance matrices with constant covariances between pairs of blocks. The positive definiteness of T is equivalent to the positive definiteness of a smaller matrix of size G, obtained by averaging each block. The model is applied to a problem in nuclear waste analysis, where one of the categorical inputs is atomic number, which has more than 90 levels.