Efficient Bayesian model selection relies on the model evidence or marginal likelihood, whose computation often requires evaluating an intractable integral. The harmonic mean estimator (HME) has long been a standard method of approximating the evidence. While computationally simple, the version introduced by Newton and Raftery (1994) potentially suffers from infinite variance. To overcome this issue, Gelfand and Dey (1994) defined a standardized representation of the estimator based on an instrumental function and Robert and Wraith (2009) later proposed to use higher posterior density (HPD) indicators as instrumental functions. Following this approach, a practical method is proposed, based on an elliptical covering of the HPD region with non-overlapping ellipsoids. The resulting estimator, called the Elliptical Covering Marginal Likelihood Estimator (ECMLE), not only eliminates the infinite-variance issue of the original HME and allows exact volume computations, but is also able to be used in multimodal settings. Through several examples, we illustrate that ECMLE outperforms other recent methods such as THAMES and its improved version (Metodiev et al. 2025). Moreover, ECMLE demonstrates lower variance-a key challenge that subsequent HME variants have sought to address-and provides more stable evidence approximations, even in challenging settings.
Kernel quadrature is widely used to approximate integrals of smooth functions, with worst-case error typically decaying at the minimax rate n^-α/d for smoothness α in dimension d. Existing rate-optimal methods often depend on deterministic point sets tailored to a specific kernel, making them sensitive to misspecification and less robust in practice. In this work, we study randomized quadrature methods with a focus on robustness rather than kernel-specific optimality. We construct an explicit, n-dependent sampling distribution that achieves minimax rates for worst-case error over smoothness classes without requiring knowledge of the kernel. This kernel-agnostic design improves robustness while retaining optimal rates. Our analysis includes unbounded sampling measures such as Gaussian and Student-t distributions, extending beyond compact domains. The results provide both theoretical guarantees and a practical recipe for robust, rate-optimal randomized quadrature.
Noninformative priors constructed for estimation purposes are usually not appropriate for model selection and testing. The methodology of integral priors was developed to get prior distributions for Bayesian model selection when comparing two models, modifying initial improper reference priors. We propose a generalisation of this methodology to more than two models. Our approach adds an artificial copy of each model under comparison by compactifying the parametric space and creating an ergodic Markov chain across all models that returns the integral priors as marginals of the stationary distribution. Besides the guarantee of their existence and the lack of paradoxes attached to estimation reference priors, an additional advantage of this methodology is that the simulation of this Markov chain is straightforward as it only requires simulations of imaginary training samples for all models and from the corresponding posterior distributions. We present some examples, including situations where other methodologies need specific adjustments or do not produce a satisfactory answer.
Two by-now folkloric results in the theory of risk sharing are that (i) any feasible allocation is convex dominated by a comonotonic allocation; and (ii) an allocation is Pareto optimal for the convex order only if it is comonotonic. Here, comonotonicity corresponds to the so-called no-sabotage condition, which the interests of all parties involved. Several proofs of these two results have been provided in the literature, based on a version of the comonotonic improvement algorithm of Landsberger and Meilijson (1994) and argument based on the Martingale Convergence Theorem. However, no proof of (i) is explicit enough for an easy algorithmic implementation in practice; and no proof of (ii) provides a closed-form characterization of Pareto optima. In addition, while all of the existing proofs of (i) are provided only for the case of a two economy with the observation that they can be easily extended beyond two agents, such an extension is being trivial in the context of the algorithm of Landsberger and Meilijson (1994) and it has never been explicitly implemented. In this paper, we provide novel proofs of these foundational results. Our proof of (i) is based theory of majorization and an extension of a result of Lorentz and Shimogaki (1968), which allows us to an explicit algorithmic construction that can be easily implemented beyond the case of two agents. In our proof of (ii) leads to a crisp closed-form characterization of Pareto-optimal allocations in terms of alpha-quantiles (mixed quantiles). An application to peer-to-peer insurance, or collaborative insurance, illustrates the relevance of these results.
Enriching Brownian motion with regenerations from a fixed regeneration distribution $\mu$ at a particular regeneration rate $\kappa$ results in a Markov process that has a target distribution $\pi$ as its invariant distribution. For the purpose of Monte Carlo inference, implementing such a scheme requires firstly selection of regeneration distribution $\mu$, and secondly computation of a specific constant $C$. Both of these tasks can be very difficult in practice for good performance. We introduce a method for adapting the regeneration distribution, by adding point masses to it. This allows the process to be simulated with as few regenerations as possible and obviates the need to find said constant $C$. Moreover, the choice of fixed $\mu$ is replaced with the choice of the initial regeneration distribution, which is considerably less difficult. We establish convergence of this resulting self-reinforcing process and explore its effectiveness at sampling from a number of target distributions. The examples show that adapting the regeneration distribution guards against poor choices of fixed regeneration distribution and can reduce the error of Monte Carlo estimates of expectations of interest, especially when $\pi$ is skewed.
The financial viability of renewable energy projects is challenged by the variability and unpredictability of production due to weather fluctuations. This paper proposes a novel risk management framework combining parametric insurance and peer-to-peer (P2P) risk sharing to address production uncertainty in solar electricity generation. We first design a weather-based parametric insurance scheme to protect against forecast errors, recalibrated at the site level to mitigate geographical basis risk. To handle residual mismatches between insurance payouts and actual losses, we introduce a complementary P2P mechanism that redistributes the remaining basis risk among participants. The method leverages physically based simulation models to reconstruct day-ahead forecasts and realized productions, integrating climate data and solar farm characteristics. A second-order theoretical approximation of the conditional expectation of the individual loss given the weather index links heterogeneous local models to a shared weather index, making risk sharing operationally feasible. In an empirical application to 50 German solar farms, our approach reduces the volatility of production losses by 55
Approximate Bayesian Computation (ABC) methods have become essential tools for performing inference when likelihood functions are intractable or computationally prohibitive. However, their scalability remains a major challenge in hierarchical or high-dimensional models. In this paper, we introduce permABC, a new ABC framework designed for settings with both global and local parameters, where observations are grouped into exchangeable compartments. Building upon the Sequential Monte Carlo ABC (ABC-SMC) framework, permABC exploits the exchangeability of compartments through permutation-based matching, significantly improving computational efficiency. We then develop two further, complementary sequential strategies: Over Sampling, which facilitates early-stage acceptance by temporarily increasing the number of simulated compartments, and Under Matching, which relaxes the acceptance condition by matching only subsets of the data. These techniques allow for robust and scalable inference even in high-dimensional regimes. Through synthetic and real-world experiments – including a hierarchical Susceptible-Infectious-Recover model of the early COVID-19 epidemic across 94 French departments – we demonstrate the practical gains in accuracy and efficiency achieved by our approach.
This paper introduces a new stochastic process with values in the set Z of integers with sign. The increments of process are Poisson differences and the dynamics has an autoregressive structure. We study the properties of the process and exploit the thinning representation to derive stationarity conditions and the stationary distribution of the process. We provide a Bayesian inference method and an efficient posterior approximation procedure based on Monte Carlo. Numerical illustrations on both simulated and real data show the effectiveness of the proposed inference.
In some applied scenarios, the availability of complete data is restricted, often due to privacy concerns; only aggregated, robust and inefficient statistics derived from the data are made accessible. These robust statistics are not sufficient, but they demonstrate reduced sensitivity to outliers and offer enhanced data protection due to their higher breakdown point. We consider a parametric framework and propose a method to sample from the posterior distribution of parameters conditioned on various robust and inefficient statistics: specifically, the pairs (median, MAD) or (median, IQR), or a collection of quantiles. Our approach leverages a Gibbs sampler and simulates latent augmented data, which facilitates simulation from the posterior distribution of parameters belonging to specific families of distributions. A by-product of these samples from the joint posterior distribution of parameters and data given the observed statistics is that we can estimate Bayes factors based on observed statistics via bridge sampling. We validate and outline the limitations of the proposed methods through toy examples and an application to real-world income data.
Generalized additive models (GAMs) are a leading model class for interpretable machine learning. GAMs were originally defined with smooth shape functions of the predictor variables and trained using smoothing splines. Recently, tree-based GAMs where shape functions are gradient-boosted ensembles of bagged trees were proposed, leaving the door open for the estimation of a broader class of shape functions (e.g. Explainable Boosting Machine (EBM)). In this paper, we introduce a competing three-step GAM learning approach where we combine (i) the knowledge of the way to split the covariates space brought by an additive tree model (ATM), (ii) an ensemble of predictive linear scores derived from generalized linear models (GLMs) using a binning strategy based on the ATM, and (iii) a final GLM to have a prediction model that ensures auto-calibration. Numerical experiments illustrate the competitive performances of our approach on several datasets compared to GAM with splines, EBM, or GLM with binarsity penalization. A case study in trade credit insurance is also provided.
This paper considers risk-sharing schemes, aiming to demonstrate that it is possible to enforce equal compensations by charging actuarially fair contributions. Specifically, consider a group of individuals exposed to the occurrence of a predefined event with adverse financial consequences such as death, survival or being diagnosed with a critical illness for instance. All members of the group agree to contribute in advance a fixed amount to a pool constituted over a reference period with the understanding that the sum of the contributions is shared in arrear among those participants having experienced the predefined event. This allocation is either uniform among claiming participants, each one receiving an equal share of total contributions, or participants are offered the choice to select a desired protection level. In the latter case, participants are free to subscribe one or several units of protection from the fund, and the total amount collected in advance is shared in arrear, equally among all units held by those participants who experienced the predefined event. Endowment contingency funds aim to provide participants with a cheap and effective protection compared to commercial insurance. The reason is that the proposed system is fully funded so that there is no risk borne by the organizer. The benefits in case the event occurs are therefore random but the volatility of the terminal payouts turns out to be limited when the number of participants gets large enough. Under independence, insurance at fair price is recovered at the limit, within infinitely large pools. As an application, the paper considers mutual aid funds and survivor funds. A comparison with takaful insurance is performed.
This paper considers a risk sharing scheme of independent discrete losses that combines risk retention at individual level, risk transfer for too expensive losses and risk pooling for the middle layer. This ensures that pooled losses can be considered as being uniformly bounded. We study the no-sabotage requirement and diversification effects when the conditional mean risk-sharing rule is applied to allocate pooled losses. The no-sabotage requirement is equivalent to Efron’s monotonicity property for conditional expectations, which is known to hold under log-concavity. Elementary proofs of this result for discrete losses are provided for finite population pools. The no-sabotage requirement and diversification effects are then examined within large pools. It is shown that Efron’s monotonicity property holds asymptotically and that risk can be eliminated under fairly general conditions which are fulfilled in applications.
The 21st century has seen an enormous growth in the development and use of approximate Bayesian methods. Such methods produce computational solutions to certain intractable statistical problems that challenge exact methods like Markov chain Monte Carlo: for instance, models with unavailable likelihoods, high-dimensional models, and models featuring large data sets. These approximate methods are the subject of this review. The aim is to help new researchers in particular -- and more generally those interested in adopting a Bayesian approach to empirical work -- distinguish between different approximate techniques; understand the sense in which they are approximate; appreciate when and why particular methods are useful; and see the ways in which they can can be combined.
D. M. Titterington合作论文数Department of Statistics8