
This article introduces a parametric mixed-tail quantile model which flexibly combines Gumbel, Fréchet and Weibull tail behaviours through weighted quantile functions. Unlike classical extreme value models, our approach simultaneously captures multiple tail types, enhancing finite-sample adaptability and modelling accuracy. The extension of the approach for the context of regression for extremes is here termed mixed-tail quantile regression. Parameters are estimated using a computationally efficient least squares approach based on the expected order statistics. We establish asymptotic properties and address inference challenges via bootstrap methods. Theoretical results, including tail dominance and adaptive bias–variance decomposition, guide practical quantile estimation. Simulations and an application on global ice extents demonstrate improved performance in modelling extreme quantiles.
Generalized additive smooth models, including those for location scale and shape, would sometimes benefit from the imposition of shape constraints on some smooth components. When these components are represented using spline basis expansions, many constraints, such as monotonicity or convexity, can be expressed as linear inequality constraints on the basis coefficients. Given smoothing parameters, model fitting by Newton's method then becomes a sequential quadratic programming problem, while the equally important task of finding initial coefficients satisfying the constraints is a simple linear programming problem. Smoothing parameters can then be estimated by an adaptation of the extended Fellner-Schall method. This article brings together and reviews the necessary methods required to provide a general purpose framework for shape constrained smooth additive modelling, covering the widely applicable case in which not all smooth terms are constrained. As well as providing examples of shape constrained generalized additive models, it is shown how the same approach facilitates non-parametric density estimation. The article is accompanied by a new function scasm in R package mgcv.
We introduce a novel extension of Benford’s Law that shifts attention from leading digits to the joint distribution of final digit pairs in numerical data. Traditional Benford-based tests rely on first-digit frequencies and require observations to span several orders of magnitude, a condition frequently violated in precinct-level election returns. To overcome this limitation, we develop a formal probabilistic framework for analyzing the behaviour of the last two digits under Benfordian assumptions, which we denote L2-BL 10. Within this framework, we derive closed-form expressions for digit-pair probabilities and establish convergence bounds toward their limiting uniform distribution, providing new theoretical results that extend the statistical foundations of Benford analysis beyond its classical first-digit form. We then show that precinct-level vote counts are well approximated by a discrete Weibull distribution, linking this theoretical development to realistic data-generating processes encountered in elections. Monte Carlo simulations and empirical illustrations using data from nine states in the 2016 and 2020 US presidential elections demonstrate how the L2-BL 10 framework can be operationalized as a practical diagnostic for identifying anomalous digit behaviour. The result is a method that is both theoretically grounded and directly applicable to election forensics in settings where standard Benford tests are not appropriate.
Longitudinal studies have known widespread use in the last years in several fields of research, as they allow to distinguish between different sources of variation. We may observe differences at the beginning of the study that stay persistent through time, and changes in the response that are due to temporal dynamics in the observed covariates. Individual-specific, time-constant, effects are often included in the linear predictor to allow for unobserved individual-specific, time constant, heterogeneity motivated by omitted individual features. The random effect approach to estimation is based on considering such effects as random variables, usually with a specific parametric distribution. This approach has been frequently criticized, as it is often employed not considering correlation between observed (i.e., covariates) and unobserved (i.e., random effects) terms. To solve this issue, we may explicitly account for correlation between observed and unobserved heterogeneity, using the so-called correlated effects approach. In this article, we show that a more general solution may be developed by estimating the random effect conditional distribution non-parametrically via a discrete probability distribution on a finite number of locations. The approach we propose is assessed via a large-scale simulation study and illustrated by the analysis of a benchmark dataset.
A technique is proposed for a more structured approach to modelling over- and under-dispersed counts: variance factorized loglinear (VFL) models are a parameterization of an (enlarged) distribution whereby the overdispersed variance is expressed as a product of the nominal variance and a remainder term involving a highly interpretable loglinear effect with a potentially different set of covariates. This facilitates mean-variance analysis by allowing the variance to be modelled more separately from the mean, especially for generalized linear models (GLMs). Consequently, VFL parameterized models offer substantial advantages over existing methods for investigating overdispersion, such as a unifying framework and greater interpretability. Examples given here of enlarged distributions that supply the equidispersion as a special case include the extended beta-binomial (with a newly proposed 'clog' link and special offset) and negative binomial (NB) and generalized Poisson (GP-1 and GP-2 variants). VFL parameterized models are constructed as a vector generalized linear model (VGLM) with a judicious combination of constraint matrices, offsets and link functions. Underdispersed data can be handled by expanding the response by a multiplier and offset adjustment. Useful for disentangling the mean-variance relationship in GLMs, VFL parameterized models are implemented within the very broad framework of the vector generalized additive model (VGAM) R package available from comprehensive r archive network (CRAN). The technique is generally applicable to mean-parameterized distributions.
We expand logistic regression models to accommodate spatial structure in a functional covariate process, direct spatial structure in responses and a geographically structured random effect. Our proposed model builds upon conditionally specified models, where the joint distribution corresponding to a specified full conditional can be identified up to an unknown constant of proportionality. To tackle the challenges of estimation and inference, we employ two separate double Metropolis-Hastings procedures in the Bayesian framework. Furthermore, we suggest methods for making modelling decisions to distinguish the sources of spatial structure through data-driven diagnostics and model assessments. Finally, we apply the model to a real-world problem involving unemployment rates over time in counties of the Midwestern United States as a functional covariate, poverty rates in those counties at a given time as response variables and state-level random effects.
Motivated by the analysis of chemical metal compounds and their properties for catalysis, we developed a gradient boosting model that explores graph structures to perform prediction tasks. Taking advantage of the iterative nature of boosting, our novel approach, called PathBoost, explores the graphs to identify the relevant paths and simultaneously fits a prediction model. Advantages of PathBoost include automatic variable selection, as only relevant paths are kept in the model and explainability, as a measure of variable importance is provided. The novel algorithm is applied to the tmQM dataset, where the molecules are represented as graphs with atoms as nodes and bonds as edges. The goal is to predict specific quantum properties of the molecules, such as the highest occupied molecular orbital (HOMO) and lowest unoccupied molecular orbital (LUMO) gap. These properties usually require heavy computational power to be computed, while our model aims to provide comparable results using much fewer resources.
Optimal designs describe the sampling distributions that achieve most efficient estimation. In this study, we apply optimal design theory to find optimal designs for the beta binomial (BB) regression model, where both location and dispersion are functions of a predictor. A major application of the BB regression model is in psychological test norming, where the location and the dispersion of test scores can be age dependent. Motivated by this context, we derive and characterize locally D-optimal designs for the BB regression model. In a practical example, we consider the designs' sensitivity to specification of model parameters, model form and optimality criterion. The designs were found to be relatively robust to these effects and thus show promise for practical implementation. However, more research is needed to investigate the suitability of the D-optimality criterion for psychological test norming and to extend the designs to be globally optimal.
We propose a distributional copula regression modelling approach for bivariate responses comprised of non-commensurate (i.e. mixed) variables. In our case, the margins are a right-censored time-to-event outcome and a non-time-to-event variable. The underlying hazard rate of the time-to-event margin is modelled using discrete-time-to-event (DT) or piecewise-exponential (PW) methods. A flexible statistical model is achieved by relying on the correspondence of the likelihood of the aforementioned time-to-event approaches with well-known univariate distributions. We construct joint bivariate distributions for these mixed responses by means of parametric bivariate copulas. This allows for separate specification of the dependence structure between the margins and their individual distribution functions. All coefficients of the distributional copula regression models considered here are estimated simultaneously via penalized maximum likelihood. We showcase the versatility of our proposed approach in an analysis of red-light running behaviour of E-cyclists by modelling the joint distribution of a mixed response comprised of a binary response and a time-to-event outcome that indicates the time of red traffic light running.
A class of new 1-parameter underdispersed distributions is introduced. Mixed with Poisson distributions; they generate 2- and 3-parameter discrete distributions that generalize the Poisson distribution and can be both under and over-dispersed. Probabilities are easy to compute and moments and random number generation are tractable. The distributions are described, and they are fitted to some underdispersed and overdispersed datasets. We show how inference for the effect of covariates sharpens on moving from the Poisson model. The fits compare favourably to two benchmarks, the COM Poisson distribution and the weighted Poisson distribution.
A realized covariance model specifies a dynamic process for a conditional covariance matrix of daily asset returns as a function of past realized variances and covariances. We propose parsimonious parameterizations enabling a spillover effect in the conditional variance equations, and a specific nonlinear, time-varying, effect of the lagged realized covariance between each asset pair on the corresponding conditional covariance. We introduce these parameterizations in four classes of realized covariance models. In an application to the components of the Dow Jones index, we find that the extended models improve the fit of their less flexible scalar versions and show a good out-of-sample forecast performance, in particular for short forecast horizons.
Discrete-time hazard models are widely used when event times are measured in intervals or are not precisely observed. While these models can be estimated using standard generalized linear model techniques, they rely on extensive data augmentation, making estimation computationally demanding in large-scale, high-dimensional settings. In this article, we demonstrate how the recently proposed batchwise backfitting algorithm, a general framework for scalable estimation and variable selection in distributional regression, can be effectively extended to discrete hazard models. Using both simulated data and a large-scale application on infant mortality in sub-Saharan Africa, we show that the algorithm delivers accurate estimates, automatically selects relevant predictors and scales efficiently to large datasets. The findings underscore the algorithm's practical utility for analyzing large-scale, complex survival data with high-dimensional covariates.
Human microbiome studies based on genetic sequencing techniques produce longitudinal data of the relative abundances of microbial taxa over time, allowing to analyze, through mixed-effects modelling, how microbial communities evolve in response to clinical interventions, environmental changes or disease progression. In particular, the Zero-Inflated Beta Regression (ZIBR) models jointly and over time the presence and abundance of each microbe taxon, considering the bounded nature of the data, its skewness and the over-abundance of zeros. However, as for other complex random effects models, maximum likelihood (ML) estimation suffers from the intractability of likelihood integrals. Available estimation methods rely on log-likelihood approximation, prone to limitations such as biased estimates or unstable convergence. In this work we develop an alternative maximum likelihood estimation approach for the ZIBR model, via the Stochastic Approximation Expectation-Maximization (SAEM) algorithm. The proposed methodology can accommodate unbalanced data, not always possible in existing approaches. We also provide estimations of standard errors and log-likelihood of the fitted model. The performance of the algorithm is established through simulation and its use is demonstrated on two microbiome studies, showing its ability to detect changes in presence and abundance of bacterial taxa over time and in response to treatment.
Advancements in computational power and methodologies have enabled research on massive datasets. However, tools for analyzing data with directional or periodic characteristics, such as wind directions and customers' arrival time in 24-hour clock, remain underdeveloped. While statisticians have proposed circular distributions for such analyses, significant challenges persist in constructing circular statistical models, particularly in the context of Bayesian methods. These challenges stem from limited theoretical development and a lack of historical studies on prior selection for circular distribution parameters.In this article, we propose a framework for selecting hyperpriors that contracts to a simpler model in circular scenarios, especially when there is insufficient information to guide prior selection. We introduce well-examined Penalized Complexity (PC) priors for the most widely used circular distributions. Comprehensive comparisons with existing hyperpriors in the literature are conducted through simulation studies and a practical case study. Final, we discuss the contributions and implications of our work, providing a foundation for further advancements in constructing Bayesian circular statistical models.
Ordinal data are quite common in applied statistics. Although some model selection and regularization techniques for categorical predictors and ordinal response models have been developed over the past few years, less work has been done concerning ordinal-on-ordinal regression. Motivated by a consumer test and a survey on the willingness to pay for luxury food products consisting of Likert-type items, we propose a strategy for smoothing and selecting ordinally scaled predictors in the cumulative logit model. First, the group lasso is modified by the use of difference penalties on neighboring dummy coefficients, thus taking into account the predictors' ordinal structure. Second, a fused lasso-type penalty is presented for the fusion of predictor categories and factor selection. The performance of both approaches is evaluated in simulation studies and on real-world data.
Uncertainty in machine learning models is a timely and vast field of research. In supervised learning, uncertainty can already occur in the first stage of the training process, the annotation phase. This scenario is particularly evident when some instances cannot be definitively classified. In other words, there is inevitable ambiguity in the annotation step and hence, not necessarily a single 'ground truth' associated with each instance. This work approaches the problem from a statistical modelling perspective. The main idea is to drop the assumption of a ground truth label and instead embed the annotations into a multidimensional space. This embedding is derived from the empirical distribution of annotations within a Bayesian setup, modelled using a Dirichlet-Multinomial framework. We estimate the model parameters and posteriors using a stochastic Expectation Maximizsation algorithm with Markov Chain Monte Carlo (MCMC) steps. The methods developed in this article readily extend to various situations in which multiple annotators independently label instances. To showcase the generality of the proposed approach, we apply our approach to three benchmark datasets for image classification and natural language inference (NLI), in which multiple annotations per instance are available. Besides the embeddings, we can investigate the resulting correlation matrices, which reflect the semantic similarities of the original classes for all three exemplary datasets.
In generalized regression models, the effect of continuous covariates is commonly assumed to be linear. This assumption, however, may be too restrictive in applications and may lead to biased effect estimates and decreased predictive ability. While a multitude of alternatives for the flexible modelling of continuous covariates have been proposed, methods that provide guidance for choosing a suitable functional form are still limited. To address this issue, we propose a detection algorithm that evaluates several approaches for modelling continuous covariates and guides practitioners to choose the most appropriate alternative. The algorithm utilizes a unified framework for tree-structured modelling which makes the results easily interpretable. We assessed the performance of the algorithm by conducting a simulation study. To illustrate the proposed algorithm, we analysed data of patients suffering from chronic kidney disease.
We present an approach for modelling heterogeneous gas flow patterns in gas transmission networks based on mixtures of generalized nonlinear models (GNMs). This enables the use of a gamma distribution for the dependent variable and to incorporate standard loading profile specifications based on nonlinear functions where the specific form to be used is explicitly prescribed by a regulatory framework agreement. We focus in particular on the modelling of the maximum daily gas flow to predict gas flow at low temperatures where we take into account if the weekday is a working day or a non-working day. The application to a gas flow dataset from Western Austria exemplifies the benefits of using mixture models to obtain maximum gas flow predictions while taking day characteristics into account.