Explainable AI (XAI) methods like SHAP and LIME produce numerical feature attributions that remain inaccessible to non expert users. Prior work has shown that Large Language Models (LLMs) can transform these outputs into natural language explanations (NLEs), but it remains unclear which factors contribute to high-quality explanations. We present a systematic factorial study investigating how Forecasting model choice, XAI method, LLM selection, and prompting strategy affect NLE quality. Our design spans four models (XGBoost (XGB), Random Forest (RF), Multilayer Perceptron (MLP), and SARIMAX - comparing black-box Machine-Learning (ML) against classical time-series approaches), three XAI conditions (SHAP, LIME, and a no-XAI baseline), three LLMs (GPT-4o, Llama-3-8B, DeepSeek-R1), and eight prompting strategies. Using G-Eval, an LLM-as-a-judge evaluation method, with dual LLM judges and four evaluation criteria, we evaluate 660 explanations for time-series forecasting. Our results suggest that: (1) XAI provides only small improvements over no-XAI baselines, and only for expert audiences; (2) LLM choice dominates all other factors, with DeepSeek-R1 outperforming GPT-4o and Llama-3; (3) we observe an interpretability paradox: in our setting, SARIMAX yielded lower NLE quality than ML models despite higher prediction accuracy; (4) zero-shot prompting is competitive with self-consistency at 7-times lower cost; and (5) chain-of-thought hurts rather than helps.
Clinical prediction models play a crucial role in advancing personalized care for mental health disorders, providing essential insights for diagnosis, prognosis and intervention planning. This work examines the current methodological approaches used to develop such models, emphasizing their application to mental health problems, including depression. To illustrate these concepts, we used data on prenatal depression from a multinational observational study of 5,372 pregnant women. The goal is to develop an individual prognostic model for depressive symptoms that can be used already at the beginning of pregnancy. Our analysis explores variable selection strategies, validation methodologies and the integration of clinical expertise with data-driven approaches. Particular attention is given to addressing challenges such as population heterogeneity, overfitting and the importance of external validation for generalizability across diverse settings. We distinguish between statistical regression models and machine learning techniques, discussing their respective strengths and limitations in terms of interpretability, predictive accuracy and clinical usability. This work offers practical guidance for researchers and clinicians, focusing on the critical steps for model development and implementation. We highlight best practices to avoid common pitfalls, advocate for interdisciplinary collaboration and address challenges of integrating advanced statistical and machine learning tools into clinical practice. By providing practical guidance and addressing these issues, our aim is to support the development of robust and clinically relevant prediction models.
Distributional regression (DR) refers to regression methods that model the entire conditional probability distribution of a response variable given a set of explanatory variables. The generalized additive model for location, scale and shape (GAMLSS) is a common framework of DR, extending traditional regression models (such as linear models, generalized linear models and generalized additive models) by modelling the mean, variance, skewness and tail behaviour as functions of explanatory variables. The influence of explanatory variables on variability, quantiles and exceedance probabilities can thus be directly measured, providing deeper insight into underlying data-generating process. DR is ideal for applications in which predicting uncertainty and extreme events — and hence risk assessment — is critical. In this Primer, we provide an overview of GAMLSS-based DR, including the theoretical background and guidelines for practical implementation and model checking. Case studies across disciplines demonstrate that GAMLSS captures intrinsic distributional features often missed by classical mean-based regression and machine learning approaches. Finally, we highlight future directions, particularly in combination with machine learning. Distributional regression models entire conditional response distributions. In this Primer, Merder et al. discuss how distributional regression is used to predict uncertainty and extreme conditions in ways often missed by classical regression and machine learning approaches.
This study compares two regional flood frequency analysis (RFFA) approaches—the parameter regression technique (PRT) and the quantile regression technique (QRT)—in the context of retention-based hydrologic applications, which require consistent design flood estimates across durations. Using a moving-window approach to estimate average streamflow over fixed durations, we assess whether predicting distribution parameters (PRT) or quantiles directly (QRT) affects the accuracy and duration consistency of flood estimates at ungauged locations. Both methods show similar predictive accuracy at return periods up to 50 years. However, the PRT more reliably preserves duration consistency: at 32 stations, the QRT produced inconsistent estimates at return periods beyond 2 years, while the PRT did not. These inconsistencies were not explained by record length or catchment characteristics. Additionally, at 64 of 239 stations, out-of-sample estimation of the median (index) flood was duration-inconsistent, suggesting difficulties in differentiating between the order of magnitude of floods across durations. Such inconsistencies were not observed in the spread and shape parameters estimated by the PRT. A quantile-based reparameterization of the generalized extreme value distribution was used to align both methods around the median flood, making it easier to compare how the methods behave across durations.
Reliable parameter estimation for the two-parameter inverse Weibull distribution is critical for accurate lifetime predictions and risk assessment in reliability engineering. However, the lack of a thorough proof of a global maximizer of the maximum likelihood estimation for this specific probability density function is an important issue because it theoretically guarantees that a possible optimization algorithm possesses numerical robustness. This implies practically that misinterpretations are avoided by not getting stuck in local extrema in optimization algorithms. Hence, in this work, we provide a comprehensive, thorough proof of existence and uniqueness for global maximizers of parameter estimation for this distribution with respect to the parameter space as our main contribution. At first, we prove local existence and uniqueness by Schauder’s fixed-point theorem, monotonicity arguments and local concavity of the Hessian matrix. Thus, we are capable of showing existence and uniqueness of a global maximizer of maximum log-likelihood parameter estimation for the two-parameter inverse Weibull distribution by considering all limiting cases with respect to the parameter space. Conclusively, we strengthen our theoretical findings by numerical experiments which include real-world data sets from reliability engineering to stress this probability distribution’s usefulness in real-world applications.
In recent years, budburst, the timing of leaf emergence, has advanced less than expected despite continued spring warming, suggesting counteracting ecological forces. One of these forces might be increased and earlier herbivory on young leaves under climate warming. Here using 5 years of satellite radar data from 27,500 pixels (10 ×10 m2) across 60 temperate oak forest sites under experimental manipulation of insect herbivore loads in Central Europe, we show that prior-year leaf herbivory delayed budburst by 3 days, cancelling the phenological advance observed during a decade of warming. This delay reduced subsequent herbivory by 55%, exceeding the effects of parasitoids or pathogens, and persisted even during pest outbreaks. Across landscapes, the delay was strongest where it probably provided the highest benefit, that is, where a given amount of delay most effectively reduced following herbivory, which suggests an adaptive tree defence. Ultimately, trees may be trapped between responding to two opposing consequences of global change: warming selects for earlier budburst, whereas herbivory selects for delay. Our results underscore the need to consider not only climate, but also plant-herbivore interactions and adaptive evolution to predict tree responses to a changing world.
A central challenge in distributional regression is to allow the shape of the conditional distribution of the response variable to vary flexibly with covariates while retaining directly interpretable effects on its mean and standard deviation. We extend the penalized transformation model (PTM) family into a conditional-shape PTM, which assigns separate structured additive predictors to the conditional mean, standard deviation, and standardized distributional shape beyond location and scale. A covariate-dependent monotone transformation maps the standardized response to a fixed reference distribution, while affine standardization enforces mean zero and variance one for the induced standardized distribution. Thus, the first two predictors remain exactly the conditional mean and standard deviation. The shape predictor accommodates selected linear, nonlinear, group-specific, spatial, and interaction effects; regularization shrinks unsupported departures toward a reference-family location-scale model. We fit the PTM using mini-batched stochastic variational inference with model-aligned Gaussian blocks and staged optimization. In simulations, the PTM recovers smooth mean and standard-deviation effects and a covariate-dependent transition from skewness to bimodality while suppressing unnecessary shape effects. Under a deliberately misspecified design, it remains competitive with a structured additive Dirichlet-process mixture in test-set density and distribution-function accuracy, although all models show undercoverage and both flexible methods miss fine features. Applications to 13,425 Norwegian water-conductivity observations and 1,182,514 German daily-temperature observations demonstrate selective group-specific, seasonal, and spatial shape variation. Predictive performance criteria favor the conditional-shape PTM over a fixed-shape PTM and a Gaussian location-scale model.
Counts in epidemiology often deviate from equidispersion and exhibit spatial, temporal, and nonlinear structure that the Poisson model cannot accommodate. We introduce a gamma-count structured additive regression model that strategically integrates penalized complexity priors in two critical aspects: (i) a principled penalized complexity prior on the dispersion parameter of the gamma-count distribution, which naturally shrinks toward the base Poisson model when the data support equidispersion, and (ii) scale-dependent penalized complexity hyperpriors on the smoothing variances for nonlinear, spatial, and temporal effects. By formulating the model within a latent Gaussian framework, we enable efficient approximate Bayesian inference through integrated nested Laplace approximations. Simulation studies across under-, equi-, and over-dispersed regimes show that the penalized complexity prior for dispersion parameter combined with scale-dependent hyperpriors yields accurate estimation of dispersion and smooth effects, favorable predictive scores, and robust inference. In empirical applications to larynx cancer mortality in Germany, COVID-19 incidence in Georgia (USA), and lung and bronchus cancer in Iowa (USA), the gamma-count structured additive regression model exhibits competitive or enhanced fit relative to Poisson and negative binomial counterparts, while elucidating interpretable nonlinear and spatial structures. This framework delivers robust, spatially resolved estimates of disease burden in the presence of non-equidispersion, thereby facilitating evidence-based resource allocation, epidemiological surveillance, and monitoring of health disparities, contributing to Sustainable Development Goals (SDGs) 3 (Good Health and Well-Being) and 10 (Reduced Inequalities). For geographically targeted analyses, it further supports informed decision-making in urban and community planning, aligning with SDG 11 (Sustainable Cities and Communities).
We introduce CAFE (Compound-AI Factorial Evaluation), an open-source platform that brings design of experiments to the evaluation of compound AI systems (CAIS). Such systems expose many interchangeable choices - e.g. which retriever, model, or prompt - and practitioners rarely know which of them most affects answer quality. With CAFE, a practitioner registers each swappable component of a pipeline as a factor to build a factorial design over the chosen factors, run the resulting configurations, and score the answers on a shared rubric using a configurable LLM judge together with human raters. From these ratings it attributes answer-quality variance to the components and their interactions with mixed-effects models and reports effect sizes, significance, the best configuration, cost and latency trade-offs, and judge-human reliability. Whereas existing tools mostly either search for a good configuration or score outputs in isolation, CAFE also explains which component drives quality and whether an observed difference is significant. We validate CAFE on a retrieval-augmented question-answering (QA) pipeline over the HotpotQA benchmark dataset, where it recovers planted factor effects and stays calibrated under a permutation null. CAFE is released as a Python package and as a Web application.
This paper introduces a novel changepoint detection framework that combines ensemble statistical methods with Large Language Models (LLMs) to enhance both detection accuracy and the interpretability of regime changes in time series data. Two critical limitations in the field are addressed. First, individual detection methods exhibit complementary strengths and weaknesses depending on data characteristics, making method selection non-trivial and prone to suboptimal results. Second, automated, contextual explanations for detected changes are largely absent. The proposed ensemble method aggregates results from ten distinct changepoint detection algorithms, achieving superior performance and robustness compared to individual methods. Additionally, an LLM-powered explanation pipeline automatically generates contextual narratives, linking detected changepoints to potential real-world historical events. For private or domain-specific data, a Retrieval-Augmented Generation (RAG) solution enables explanations grounded in user-provided documents. The open source Python framework demonstrates practical utility in diverse domains, including finance, political science, and environmental science, transforming raw statistical output into actionable insights for analysts and decision-makers.
ABSTRACT In spatial regression models, unmeasured spatial variables, represented by spatial random effects, are typically not independent of observed covariates and can induce significant bias in covariate effect estimates. This fundamental problem, known as spatial confounding, has been extensively studied, but sometimes with puzzling and seemingly contradictory results. Here, we introduce a broad theoretical framework that brings mathematical clarity to spatial confounding. Our focus is the finite‐sample bias affecting simulation results and practical applications. Our intention is not to challenge existing results in the literature, but rather to build intuition and provide a unifying perspective that helps explain and connect them. We derive a compact analytical expression where, in the metric induced by the precision structure of the chosen analysis model, the bias is expressed in terms of the empirical correlation between a covariate and its confounder and their relative sizes. Thus, our formulation directly identifies these relationships as the key driver of bias; moreover, using the eigendecomposition of the precision structure, we obtain detailed quantitative characterizations of its behavior. The eigendecomposition has a natural interpretation as spatial frequencies and nonspatial information, with the eigenvalues encoding the precise effect of spatial smoothing in the model. This allows the insights from our bias expression to be translated into interpretable scenarios and explain subtle and counter‐intuitive behaviors. Finally, we propose a general strategy for addressing the bias in practice. When a covariate has nonspatial information, we show that a general form of the so‐called spatial+ method can reduce the bias. If not, explicit assumptions for identifiability are needed, and we develop a procedure in which multiple capped versions of spatial+ are applied. We illustrate our approach with an application to air temperature in Germany.
Bounded continuous data on the unit interval frequently arise in applied fields and often exhibit a non-negligible proportion of observations at the boundaries. Inflated regression models address this feature by combining a continuous distribution on the unit interval with a discrete component to account for zero- and/or one-inflation. In this paper, we propose a class of Bayesian structured additive quantile regression models for inflated bounded continuous data that accommodates zero- and/or one-inflation. The proposed approach enables direct modeling of both the conditional quantiles of the continuous component and the probabilities of observing zeros and/or ones, with structured additive predictors incorporated in both parts, including nonlinear effects, spatial effects, random effects, and varying-coefficient terms. Posterior inference is carried out using Markov chain Monte Carlo algorithms implemented through the software Liesel, a probabilistic programming framework for semiparametric regression. The practical performance of the proposed models is illustrated through simulation studies and two real-data applications: one analyzing the proportion of traffic-related fatalities across Brazilian municipal districts, and another evaluating speech intelligibility in cochlear implant recipients under different experimental conditions.
The under-representation of women as sources cited in news media is one prominent representation of gender bias. Understanding where gender bias concentrates and how it evolves is essential for targeted mitigation. Because gender representation varies across topics, time, and reported-on regions, creating complex dependencies that are difficult to capture parametrically, we employ a nonparametric model to uncover latent cluster structures and temporal dynamics. We combine time-dependent Bayesian mixture modeling techniques with a Beta mixture kernel tailored to female quote shares, bounded between 0 and 1. Fitted on Canadian news articles from 2019 to 2024, the model reveals structural under-representation of women across all clusters, with news topic driving differences in female quote shares more strongly than the reported-on region. More than 85
We read with interest the above article by Zavorsky (2025, Respiratory Medicine, doi:10.1016/j.rmed.2024.107836) concerning reference equations for pulmonary function testing. The author compares a Generalized Additive Model for Location, Scale, and Shape (GAMLSS), which is the standard adopted by the Global Lung Function Initiative (GLI), with a segmented linear regression (SLR) model, for pulmonary function variables. The author presents an interesting comparison; however there are some fundamental issues with the approach. We welcome this opportunity for discussion of the issues that it raises. The author's contention is that (1) SLR provides "prediction accuracies on par with GAMLSS"; and (2) the GAMLSS model equations are "complicated and require supplementary spline tables", whereas the SLR is "more straightforward, parsimonious, and accessible to a broader audience". We respectfully disagree with both of these points.
Gas fluxes between soil and atmosphere play a significant role as global greenhouse gas sinks or sources. Accurate estimation of these gas fluxes is challenging due to the heterogeneous nature of soil properties, dynamic environmental factors, and the complexity of measurement in soil systems. The flux-gradient method is a reliable approach for estimating CO2 gas fluxes. This method utilises Fick's first law of diffusion to calculate gas fluxes by analysing gas concentration profiles. The gas concentration gradient is multiplied with the apparent gas diffusion coefficient of the gas species in the soil. The estimation of the apparent gas diffusion coefficient is dependent on numerous parameters, including soil pore space and soil water content, which must be carefully measured or derived. These parameters typically cannot be measured at the exactly same location as not to interfere with the gas measurements. All these factors are subject to different uncertainties depending on the parameters and the measured soil location.In order to address and comprehend the extent of these uncertainties, Bayesian inference was employed, as this methodology enables uncertainty to be measured through probability distributions with credible intervals as opposed to point estimates. Furthermore, Bayesian inference functions effectively with small datasets and permits the incorporation of prior knowledge, a factor which also benefits soil gas modelling.We used a previously published dataset (Wordell-Dietrich et al. 2020) to estimate CO2 fluxes by using the flux-gradient method. The gas diffusion coefficient was derived through the use of a Bayesian inference model. The resulting data were then compared with chamber measurements and other modelling approaches.AcknowledgmentWordell-Dietrich, P.; Wotte, A.; Rethemeyer, J.; Bachmann, J.; Helfrich, M.; Kirfel, K.; Leuschner, C.; Don, A. (2020): Vertical partitioning of CO2 production in a forest soil. Biogeosciences, 17, 6341-6356. https://doi.org/10.5194/bg-17-6341-202
This article celebrates the life and scientific legacy of Carmen Cadarso, a leading figure in the development of Biostatistics in Spain. Combining personal reflections with an overview of her research, we begin by revisiting Carmen's academic path, including her early training, key influences, and her longstanding commitment to research grounded in collaboration and social relevance. The second part of the article focuses on her scientific contributions, guided by a thematic thread around flexible statistical models. Selected examples illustrate how she contributed to methods that are both rigorous and practical, consistently shaped by biomedical applications. Carmen's work, teaching, and leadership have left a lasting impact on the field and on the many colleagues and students who had the privilege to work with her. This article is dedicated, with our deepest affection, to Klaus Ebert, Carmen's husband.
Recent advancements in Artificial Intelligence (AI), notably the development of Large Language Models (LLMs) and text-to-image diffusion models, have facilitated the creation of realistic textual content and images. Specifically, platforms like ChatGPT and Midjourney have simplified the creation of high-quality text and visuals with minimal expertise and cost. The increasing sophistication of Generative AI presents challenges in ensuring the integrity of news, media, and information quality, making it increasingly difficult to distinguish between real and artificially generated textual and visual content. Our work addressed this problem in two ways. First, by means of ChatGPT and Midjourney, we created a comprehensive novel multimodal news corpus named SyN24News based on the N24News corpus, on which we evaluated our model. Second, we developed a novel explainable synthetic news detector for discriminating between real and synthetic news articles. We leveraged a Neural Additive Model (NAM)-like network structure that ensures effect separation by handling input data in separate subnetworks. Complex structures and patterns are extracted by deep features from unstructured data, i.e., images and texts, using fine-tuned VGG and DistilBERT subnetworks. We ensured further explainability by individually processing carefully chosen handcrafted text and image features in simple Multilayer Perceptrons (MLPs), allowing for graphical interpretation of corresponding structured effects. Our findings indicate that textual information are the main drivers in the decision-making finding process. Structured textual effects, particularly Flesch-Kincaid reading ease and sentiment, have a much higher influence on the classification outcome than visual features such as dissimilarity and homogeneity.
Density regression models allow a comprehensive understanding of data by modeling the complete conditional probability distribution. While flexible estimation approaches such as normalizing flows (NFs) work particularly well in multiple dimensions, interpreting the input-output relationship of such models is often difficult, due to the blackbox character of deep learning models. In contrast, existing statistical methods for multivariate outcomes such as multivariate conditional transformation models (MCTMs) are restricted in flexibility and are often not expressive enough to represent complex multivariate probability distributions. In this paper, we combine MCTMs with state-of-the-art and autoregressive NFs to leverage the transparency of MCTMs for modeling interpretable feature effects on the marginal distributions in the first step and the flexibility of neural-network-based NF techniques to account for complex and non-linear relationships in the joint data distribution. We demonstrate our method's versatility in various numerical experiments and compare it with MCTMs and other NF models on both simulated and real-world data.