This short study presents an opportunistic approach to a (more) reliable validation method for prediction uncertainty average calibration. Considering that variance-based calibration metrics (ZMS, NLL, RCE...) are quite sensitive to the presence of heavy tails in the uncertainty and error distributions, a shift is proposed to an interval-based metric, the Prediction Interval Coverage Probability (PICP). It is shown on a large ensemble of molecular properties datasets that (1) sets of z-scores are well represented by Student's-t(ν) distributions, ν being the number of degrees of freedom; (2) accurate estimation of 95 % prediction intervals can be obtained by the simple 2σ rule for ν>3; and (3) the resulting PICPs are more quickly and reliably tested than variance-based calibration metrics. Overall, this method enables to test 20 % more datasets than ZMS testing. Conditional calibration is also assessed using the PICP approach.
Average calibration of the uncertainties of machine learning regression tasks can be tested in two ways. One way is to estimate the calibration error (CE) as the difference between the mean absolute error (MSE) and the mean variance (MV) or mean squared uncertainty. The alternative is to compare the mean squared z-scores or scaled errors (ZMS) to 1. Both approaches might lead to different conclusion, as illustrated on an ensemble of datasets from the recent machine learning uncertainty quantification literature. It is shown here that the CE is very sensitive to the distribution of uncertainties, and notably to the presence of outlying uncertainties, and that it cannot be used reliably for calibration testing. By contrast, the ZMS statistic does not present this sensitivity issue and offers the most reliable approach in this context. Implications for the validation of conditional calibration are discussed.
Average calibration of the (variance-based) prediction uncertainties of machine learning regression tasks can be tested in two ways: one is to estimate the calibration error (CE) as the difference between the mean absolute error (MSE) and the mean variance (MV); the alternative is to compare the mean squared z-scores (ZMS) to 1. The problem is that both approaches might lead to different conclusions, as illustrated in this study for an ensemble of datasets from the recent machine learning uncertainty quantification (ML-UQ) literature. It is shown that the estimation of MV, MSE and their confidence intervals becomes unreliable for heavy-tailed uncertainty and error distributions, which seems to be a frequent feature of ML-UQ datasets. By contrast, the ZMS statistic is less sensitive and offers the most reliable approach in this context, still acknowledging that datasets with heavy-tailed z-scores distributions should be considered with great care. Unfortunately, the same problem is expected to affect also conditional calibrations statistics, such as the popular ENCE, and very likely post-hoc calibration methods based on similar statistics. Several solutions to circumvent the outlined problems are proposed.
Some popular Machine Learning Uncertainty Quantification (ML-UQ) calibration statistics do not have predefined reference values and are mostly used in comparative studies. In consequence, calibration is almost never validated and the diagnostic is left to the appreciation of the reader. Simulated reference values, based on synthetic calibrated datasets derived from actual uncertainties, have been proposed to palliate this problem. As the generative probability distribution for the simulation of synthetic errors is often not constrained, the sensitivity of simulated reference values to the choice of generative distribution might be problematic, shedding a doubt on the calibration diagnostic. This study explores various facets of this problem, and shows that some statistics are excessively sensitive to the choice of generative distribution to be used for validation when the generative distribution is unknown. This is the case, for instance, of the correlation coefficient between absolute errors and uncertainties (CC) and of the expected normalized calibration error (ENCE). A robust validation workflow to deal with simulated reference values is proposed.
Reliable uncertainty quantification (UQ) in machine learning (ML) regression tasks is becoming the focus of many studies in materials and chemical science. It is now well understood that average calibration is insufficient, and most studies implement additional methods for testing the conditional calibration with respect to uncertainty, i.e., consistency. Consistency is assessed mostly by so-called reliability diagrams. There exists, however, another way beyond average calibration, which is conditional calibration with respect to input features, i.e., adaptivity. In practice, adaptivity is the main concern of the final users of the ML-UQ method, seeking the reliability of predictions and uncertainties for any point in the feature space. This article aims to show that consistency and adaptivity are complementary validation targets and that good consistency does not imply good adaptivity. An integrated validation framework is proposed and illustrated with a representative example.
Light-induced charge accumulation is at the heart of biomimetic systems aiming at solar fuel production in the realm of artificial photosynthesis. Understanding the mechanisms upon which these processes operate is a necessary condition to drive down the rational catalyst design road. We have built a nanosecond pump-pump-probe resonance Raman setup to witness the sequential charge accumulation process while probing vibrational features of different charge-separated states. By employing a reversible model system featuring methyl viologen (MV) as a dual electron acceptor, we have been able to watch the photosensitized production of its neutral form, MV0, resulting from two sequential electron transfer reactions. We have found that, upon double excitation, a fingerprint vibrational mode corresponding to the doubly reduced species appears at 992 cm-1 and peaks at 30 μs after the second excitation. This has been further confirmed by simulated resonance Raman spectra which fully support our experimental findings in this unprecedented buildup of charge seen by a resonance Raman probe.
The Expected Normalized Calibration Error (ENCE) is a popular calibration statistic used in Machine Learning to assess the quality of prediction uncertainties for regression problems. Estimation of the ENCE is based on the binning of calibration data. In this short note, I illustrate an annoying property of the ENCE, i.e. its proportionality to the square root of the number of bins for well calibrated or nearly calibrated datasets. A similar behavior affects the calibration error based on the variance of z-scores (ZVE), and in both cases this property is a consequence of the use of a Mean Absolute Deviation (MAD) statistic to estimate calibration errors. Hence, the question arises of which number of bins to choose for a reliable estimation of calibration error statistics. A solution is proposed to infer ENCE and ZVE values that do not depend on the number of bins for datasets assumed to be calibrated, providing simultaneously a statistical calibration test. It is also shown that the ZVE is less sensitive than the ENCE to outstanding errors or uncertainties.
Abstract Post hoc recalibration of prediction uncertainties of machine learning regression problems by isotonic regression might present a problem for bin-based calibration error statistics (e.g. ENCE). Isotonic regression often produces stratified uncertainties, i.e. subsets of uncertainties with identical numerical values. Partitioning of the resulting data into equal-sized bins introduces an aleatoric component to the estimation of bin-based calibration statistics. The partitioning of stratified data into bins depends on the order of the data, which is typically an uncontrolled property of calibration test/validation sets. The tie-braking method of the ordering algorithm used for binning might also introduce an aleatoric component. I show on an example how this might significantly affect the calibration diagnostics.
The practice of uncertainty quantification (UQ) validation, notably in machine learning for the physico-chemical sciences, rests on several graphical methods (scattering plots, calibration curves, reliability diagrams and confidence curves) which explore complementary aspects of calibration, without covering all the desirable ones. For instance, none of these methods deals with the reliability of UQ metrics across the range of input features (adaptivity). Based on the complementary concepts of consistency and adaptivity, the toolbox of common validation methods for variance- and intervals- based UQ metrics is revisited with the aim to provide a better grasp on their capabilities. This study is conceived as an introduction to UQ validation, and all methods are derived from a few basic rules. The methods are illustrated and tested on synthetic datasets and representative examples extracted from the recent physico-chemical machine learning UQ literature.
Modelling the chemical composition of Titan’s ionosphere is a very challenging issue. Latest works perform either inversion of CASSINI’s INMS mass spectra (neutral[1] or ion[2]), or design coupled ion-neutral chemistry models[3]. Coupling ionic and neutral chemistry has been reported to be an essential feature of accurate modelling[3]. Electron Dissociative Recombination (EDR), where free electrons recombine with positive ions to produce neutral species, is a key component of ion-neutral coupling. Experimental databases of EDR are incomplete. Indeed, for heavy hydrocarbon ions (CnH y , n ≥ 3), final neutral molecules distribution is unknown. Data are available only over the carbon composition of fragments, which leaves a deterministic approach irrelevant. Thanks to a novel stochastic description of the EDR chemical reactions, we were able to calculate neutral species fluxes du to EDR processes and compare them to fluxes du to neutral chemistry processes. Some species’ fluxes, like nitrile molecules, are mainly described by EDR processes (Fig. 1). Although hydrocarbon fluxes are dominated by the neutral chemistry contribution, EDR contribution grows with the molecules’ mass (Fig. 2) and is therefore of importance for heavy molecules coupling.
We describe an automated algorithm allowing extraction of quantitative corneal transparency parameters with clinical Spectral-Domain Optical Coherence Tomography (SD-OCT). Our algorithm employs a novel pre-processing procedure to standardize SD-OCT image analysis and to numerically correct common instrumental artifacts before extracting mean intensity stromal-depth (z) profiles over a 6-mm-wide corneal area. The z-profiles are analyzed using our previously developed objective method deriving quantitative transparency parameters which are directly related to the physics of light propagation in tissues. Tissular heterogeneity is quantified by the Birge ratio, Br; for homogeneous tissues (i.e., Br~1), the photon mean-free path (ls) may be determined. Images of 83 normal corneas (ages 22–50 years) from a standard SD-OCT device (RTVue-XR Avanti, Optovue Inc.) were processed to establish a normative dataset of transparency values. After confirming stromal homogeneity (Br⪅10), we measured a median ls of 570 μm (interdecile range: 270–2400 μm). Considering corneal thicknesses, this may be translated into a median fraction of transmitted (coherent) light Tcoh(stroma) of 51% (interdecile range: 22–83%). Excluding images with central saturation artifact raised our median Tcoh(stroma) to 73% (inter-decile range: 34–84%). These transparency values are slightly lower than previously reported, which we attribute to the detection configuration of SD-OCT with a relatively small and selective acceptance angle. No statistically significant correlation between transparency and age or thickness was found. Our algorithm provides robust and quantitative measurements of corneal transparency from standard SD-OCT images with sufficient quality and addresses the demand for such an objective means in the clinical setting.
Binwise Variance Scaling (BVS) has recently been proposed as a post hoc recalibration method for prediction uncertainties of machine learning regression problems that is able of more efficient corrections than uniform variance (or temperature) scaling. The original version of BVS uses uncertainty-based binning, which is aimed to improve calibration conditionally on uncertainty, i.e. consistency. I explore here several adaptations of BVS, in particular with alternative loss functions and a binning scheme based on an input-feature (X) in order to improve adaptivity, i.e. calibration conditional on X. The performances of BVS and its proposed variants are tested on a benchmark dataset for the prediction of atomization energies and compared to the results of isotonic regression.
Corneal transparency is essential to provide a clear view into and out of the eye, yet clinical means to assess such transparency are extremely limited and usually involve a subjective grading of visible opacities by means of slit-lamp biomicroscopy. Here, we describe an automated algorithm allowing extraction of quantitative corneal transparency parameters with standard clinical spectral-domain optical coherence tomography (SD-OCT). Our algorithm employs a novel pre-processing procedure to standardize SD-OCT image analysis and to numerically correct common instrumental artifacts before extracting mean intensity stromal-depth ( z ) profiles over a 6-mm-wide corneal area. The z -profiles are analyzed using our previously developed objective method that derives quantitative transparency parameters directly related to the physics of light propagation in tissues. Tissular heterogeneity is quantified by the Birge ratio B r and the photon mean-free path ( l s ) is determined for homogeneous tissues (i.e., B r ~1 ). SD-OCT images of 83 normal corneas (ages 22–50 years) from a standard SD-OCT device (RTVue-XR Avanti, Optovue Inc.) were processed to establish a normative dataset of transparency values. After confirming stromal homogeneity ( B r <10), we measured a median l s of 570 μm (interdecile range: 270–2400 μm). By also considering corneal thicknesses, this may be translated into a median fraction of transmitted (coherent) light T coh(stroma) of 51% (interdecile range: 22–83%). Excluding images with central saturation artifact raised our median T coh(stroma) to 73% (interdecile range: 34–84%). These transparency values are slightly lower than those previously reported, which we attribute to the detection configuration of SD-OCT with a relatively small and selective acceptance angle. No statistically significant correlation between transparency and age or thickness was found. In conclusion, our algorithm provides robust and quantitative measurements of corneal transparency from standard SD-OCT images with sufficient quality (such as ‘Line’ and ‘CrossLine’ B-scan modes without central saturation artifact) and addresses the demand for such an objective means in the clinical setting.
Adjusting the band gap energy of metal halide perovskite by anion exchange (CsPbBr 3− y X y : X = Cl, Br, I) leads to optimal interfacial electron transfer from CsPbBr 3− y X y to TiO 2 , and thus to improved photocatalytic hydrogen generation.
Context. The chemical building blocks of life contain a large proportion of nitrogen, an essential element. Titan, the largest moon of Saturn, with its dense atmosphere of molecular nitrogen and methane, offers an exceptional opportunity to explore how this element is incorporated into carbon chains through atmospheric chemistry in our Solar System. A brownish dense haze is consistently produced in the atmosphere and accumulates on the surface on the moon. This solid material is nitrogen-rich and may contain prebiotic molecules carrying nitrogen. Aims. To date, our knowledge of the processes leading to the incorporation of nitrogen into organic chains has been rather limited. In the present work, we investigate the formation of nitrogen-bearing ions in an experiment simulating Titan’s upper atmosphere, with strong implications for the incorporation of nitrogen into organic matter on Titan. Methods. By combining experiments and theoretical calculations, we show that the abundant N 2 + ion, produced at high altitude by extreme-ultraviolet solar radiation, is able to form nitrogen-rich organic species. Results. An unexpected and important formation of CH 3 N 2 + and CH 2 N 2 + diazo-ions is experimentally observed when exposing a gas mixture composed of molecular nitrogen and methane to extreme-ultraviolet radiation. Our theoretical calculations show that these diazo-ions are mainly produced by the reaction of N 2 + with CH 3 radicals. These small nitrogen-rich diazo-ions, with a N/C ratio of two, appear to be a missing link that could explain the high nitrogen content in Titan’s organic matter. More generally, this work highlights the importance of reactions between ions and radicals, which have rarely been studied thus far, opening up new perspectives in astrochemistry.
In this paper, the history, present status, and future of density-functional theory (DFT) is informally reviewed and discussed by 70 workers in the field, including molecular scientists, materials scientists, method developers and practitioners. The format of the paper is that of a roundtable discussion, in which the participants express and exchange views on DFT in the form of 300 individual contributions, formulated as responses to a preset list of 26 questions. Supported by a bibliography of 776 entries, the paper represents a broad snapshot of DFT, anno 2022.
This work shows that S atom substitution in phosphate controls the directionality of hole transfer processes between the base and sugar-phosphate backbone in DNA systems. The investigation combines synthesis, electron spin resonance (ESR) studies in supercooled homogeneous solution, pulse radiolysis in aqueous solution at ambient temperature, and density functional theory (DFT) calculations of in-house synthesized model compound dimethylphosphorothioate (DMTP(O-)═S) and nucleotide (5'-O-methoxyphosphorothioyl-2'-deoxyguanosine (G-P(O-)═S)). ESR investigations show that DMTP(O-)═S reacts with Cl2•- to form the σ2σ*1 adduct radical -P-S[Formula: see text]Cl, which subsequently reacts with DMTP(O-)═S to produce [-P-S[Formula: see text]S-P-]-. -P-S[Formula: see text]Cl in G-P(O-)═S undergoes hole transfer to Gua, forming the cation radical (G•+) via thermally activated hopping. However, pulse radiolysis measurements show that DMTP(O-)═S forms the thiyl radical (-P-S•) by one-electron oxidation, which did not produce [-P-S[Formula: see text]S-P-]-. Gua in G-P(O-)═S is oxidized unimolecularly by the -P-S• intermediate in the sub-picosecond range. DFT thermochemical calculations explain the differences in ESR and pulse radiolysis results obtained at different temperatures.
Confidence curves are used in uncertainty validation to assess how large uncertainties ($u_{E}$) are associated with large errors ($E$). An oracle curve is commonly used as reference to estimate the quality of the tested datasets. The oracle is a perfect, deterministic, error predictor, such as $|E|=\pm u_{E}$, which corresponds to a very unlikely error distribution in a probabilistic framework and is unable unable to inform us on the calibration of $u_{E}$. I propose here to replace the oracle by a probabilistic reference curve, deriving from the more realistic scenario where errors should be random draws from a distribution with standard deviation $u_{E}$. The probabilistic curve and its confidence interval enable a direct test of the quality of a confidence curve. Paired with the probabilistic reference, a confidence curve can be used to check the calibration and tightness of prediction uncertainties.
Validation of prediction uncertainty (PU) is becoming an essential task for modern computational chemistry. Designed to quantify the reliability of predictions in meteorology, the calibration-sharpness (CS) framework is now widely used to optimize and validate uncertainty-aware machine learning (ML) methods. However, its application is not limited to ML and it can serve as a principled framework for any PU validation. The present article is intended as a step-by-step introduction to the concepts and techniques of PU validation in the CS framework, adapted to the specifics of computational chemistry. The presented methods range from elementary graphical checks to more sophisticated ones based on local calibration statistics. The concept of tightness, is introduced. The methods are illustrated on synthetic datasets and applied to uncertainty quantification data issued from the computational chemistry literature.