
As large language models (LLMs) become increasingly integrated into analytical workflows, an urgent question arises: Can these models replace the trained statistician? This paper presents a controlled experiment to directly test the statistical reasoning ability of LLMs. We employ Monte Carlo simulations to generate datasets with known ground-truth parameters and pose four canonical statistical questions to five commercially prominent models across three linguistically distinct prompt formulations and five sampling temperature settings, yielding 3000 observations in a full-factorial design. In order to find an answer to our question, we design prompts that reflect how decision-makers with varying degrees of statistical knowledge would query AI in the absence of a trained statistician. We find that prompts that portray higher statistical competency can result in higher accuracy for some (but not all) LLMs; we also find that LLMs can fail catastrophically on tasks requiring quantitative precision. We connect these findings to architectural differences among models and to recent literature on epistemic mirroring in LLMs and argue that the observed patterns reveal models are performing linguistic pattern matching on statistically flavoured text rather than genuine statistical reasoning. We conclude that current LLMs cannot replace the statistician, though certain architectures approach useful performance on pattern-recognition subtasks.
In this paper, we deal with the problem of estimating the population quantiles using the judgement post stratification (JPS) sampling scheme. We introduce a general class of quantile function estimators, which includes quantile estimators based on both empirical and kernel distribution functions. We next study the asymptotic properties of the quantile estimators. Specifically, we prove that the estimators in the proposed class converge completely to the true quantile function under some mild conditions. We also establish the Bahadur representation for the JPS sample quantiles and address their multivariate normality. We then conduct an extensive Monte Carlo simulation study to compare the performance of the quantile function estimators in the JPS sampling design with their simple random sampling (SRS) competitors. Our study considers various factors such as sample size, set size, ranking quality, parent distribution and kernel function. Our findings show that the JPS estimators significantly enhance the efficiency of the quantile function estimators compared to their SRS counterparts across a broad range of scenarios. Finally, we illustrate the application of the proposed estimators to a real bone mineral density dataset from the Third National Health and Nutrition Examination Survey (NHANES III) to demonstrate their usefulness and practical benefits.
The importance of accurate measurement of economic or social inequality is generally accepted. Estimators of its most popular measure, the Gini coefficient, are commonly criticised for becoming increasingly imprecise for increasingly skewed distributions and small to moderate samples, i.e., two phenomena that we face more frequently today. More robust inequality measures, typically based on quantile ratios, are well studied in theory, but still attract little attention in practice. We compare through simulations bias, mean squared error and sensitivity to outliers of several inequality estimators, namely, the Gini and quantile-based indicators, including the recently proposed quantile ratio index (QRI). Our results, based on Italian SILC and synthetic data, demonstrate that the QRI estimator offers superior precision and robustness, making it a reliable tool for monitoring inequality.
Non-symmetric correspondence analysis (NSCA) is a powerful analytical technique designed to facilitate the exploration, interpretation and visualisation of the asymmetric association among nominal or ordinal categorical variables. We extend the theory of three-way ordinal NSCA to multiway data; in particular, we consider the case of four-way data using four-variate moment decomposition (FMD). The present study focuses on ordinal NSCA, specifically tailored for a completely ordered contingency table. Ordinal NSCA visualises the asymmetric association among the ordered categories of the variables through the decomposition of the centred predictor profile array using orthogonal polynomials. The traditional method of ordinal NSCA for three-way data utilises the Kronecker product for decomposing the centred column-tube profile array. We propose the use of tensor operations to perform FMD instead of the Kronecker product. We summarise some interesting tensor properties. Furthermore, we develop an algorithm to perform four-way ordinal NSCA using a tensorial approach. The advantage of the proposed algorithm is that it can be easily generalised to a multiway case. We explore the association between the variables through an interactively coded predictor isometric biplot. An R function is developed to perform ordinal NSCA on four-way data, and its implementation is demonstrated through a numerical example.
Coarsened data, where the true value of a variable is not fully observed, arise frequently. One common form is right-censoring-where the true value exceeds an observed lower bound-familiar from survival analysis as a feature of outcomes but less commonly recognized as a feature of covariates. Right-censored covariates resemble missing covariates, making the missing data literature a natural source of estimators. The two problems are not equivalent, however: A missing covariate is completely unknown, whereas a right-censored covariate is bounded below. How estimators must be adapted from one problem to the other has not received systematic treatment. Through the framework of -estimation, we establish how estimators developed for the missing covariate problem change when adapted to the right-censored covariate problem and show that adaptation has unexpected consequences: Augmented estimators can lose double robustness, weight specifications natural for missingness can introduce bias and efficiency gains can disappear. We develop two new augmented estimators for the right-censored covariate problem and provide improved versions of all augmented estimators with closed-form expressions and guaranteed efficiency gains. Simulation studies and an application to cognitive health in Huntington disease confirm these findings. The framework clarifies how estimators change when adapted and guides analysts in choosing among them.
We propose a generalised information criteria ( gic) that accounts for sparsity pattern in the model. We obtain both asymptotic and nonasymptotic results for model selection. Moreover, we show that the gic is useful for selecting the regularisation parameter in regularised estimation in high-dimensional scenarios. The results are illustrated in two examples: group LASSO in the context of generalised linear regressions and low-rank matrix regression.
In high-dimensional survival analysis, effective variable selection is crucial both for model interpretation and predictive performance. This paper investigates Cox regression with lasso and adaptive lasso penalties in genomic datasets where covariates far outnumber observations. We propose and evaluate four weight calculation strategies for adaptive lasso specifically designed for high-dimensional settings: ridge regression, principal component analysis (PCA), univariate Cox regression and random survival forest (RSF)-based weights. To address the inherent variability in high-dimensional model selection, we develop a robust procedure that evaluates performance across multiple data partitions and selects variables based on a novel importance index. Extensive simulation studies demonstrate that adaptive lasso with ridge and PCA weights significantly outperforms standard lasso in variable selection accuracy while maintaining similar or better predictive performance across various correlation structures, censoring proportions (0% to 80%) and dimensionality settings. These improvements are particularly pronounced in highly censored scenarios, making our approach valuable for real-world genetic studies with limited observed events. We apply our methodology to triple-negative breast cancer data with 234 patients, over 19 500 variables and 82% censoring, identifying key genetic and clinical prognostic factors. Our findings demonstrate that adaptive lasso with appropriate weight calculation provides more stable and interpretable models for high-dimensional survival analysis.
Bipartite record linkage has the goal of identifying observations referring to the same individual, called coreferent observations, across two distinct non-duplicated datasets. The two main approaches to solve this task are the Fellegi-Sunter model, which relies on pairwise comparisons of observations, and the graphical record linkage model, which directly models the data and groups together coreferent observations. In this paper, we aim to investigate the similarities between these two methods. We show that both models can be expressed in terms of a latent binary matrix indicating coreferent record pairs, that they can be framed as particular latent class analysis models and that they admit a direct relationship between their parameters under a common data model. Moreover, we propose a unified estimation framework based on a classification expectation-maximization algorithm. The proposed estimation method properly incorporates the problem constraints, while still allowing for a computationally efficient implementation. Moreover, it allows for an interchangeable use of the same distributional assumptions on the linkage distribution between the two models. Empirical results using the proposed estimation method demonstrate satisfactory and mostly equivalent performance for two models both on simulations and on a real dataset commonly used as a benchmark for record linkage.
When two random variables are positive quadrant dependent (PQD), they are more likely to assume small (or large) values simultaneously compared with when the random variables are independent. This dependence structure is of interest in many areas, including finance, actuarial science and engineering. We propose a new nonparametric goodness-of-fit testing procedure to assess whether PQD holds between two random variables. Our test uses empirical likelihood (EL) and is motivated by the seminal work of Owen and McKeague on this topic. We reviewed the statistics and econometrics literature and identified six nonparametric tests for the same problem, all of which estimate copula functions first. An advantage of our test is that it avoids copula estimation and can be implemented using asymptotic or finite-sample critical values. Our comparisons reveal the EL test performs as well as or better than copula-based approaches for a variety of dependence structures. We analyse three data sets and provide online R resources. A useful by-product of our work is that we synthesize a complex set of existing methods and offer data analysts the ability to implement all available nonparametric goodness-of-fit tests at once.
Surveys are based on questionnaires for generating data that are statistically analysed. The analysis can be performed descriptively and/or with log linear models, structural equations, Poisson regressions, decision trees, Bayesian networks or other such models. Artificial intelligence (AI) and machine learning (ML) methods are gaining presence in research and data analysis. These methods are based on splitting data into training and validation sets and not on stochastic assumptions. Predictive models are learned on the training data and assessed with validation data. This paper provides a perspective on applications of analytics, AI and ML to survey data analysis. By definition, ML is a subfield of AI focused on the algorithms that allow computer systems to perform specific tasks without explicit instructions. We use AI, statistical analysis and analytics as interchangeable terms. From a strategic perspective, the paper presents an information quality framework that is relevant to survey data analytics. The objective is to encourage a transition of survey data analysis, from the traditional enumerative context to future looking analytics.
In clinical trials planning, evaluation of the probability of success of an experiment is of central interest, for instance, in sample size determination. This assessment typically involves analyses of the power function of a test on a parameter of interest, such as a relevant treatment effect. In this article, we adopt a hybrid frequentist-Bayesian approach that is lately becoming more and more popular in the literature. Specifically, we focus on superiority trials, and we study the distribution of the power function induced by a design prior assigned to the parameter. Under mild assumptions, we derive general expressions for the cumulative and density functions of the random power in terms of its inverse. We then specialise this result to tests based on pivotal quantities, and we consider some classes of problems, both exact and asymptotic, conventionally employed in clinical trials. Ideas are exposed by resorting to four biomedical settings adapted from real data applications.
A changing survey landscape with increasing nonresponse rates and survey costs has caused organizations to explore new data sources for statistics production. There is great potential to use new types of data for statistics production, especially when blending them with existing data. We present a total error framework that covers all types of data sources, but our examples focus on digital data. We define digital data to include data from social networks, traditional business systems and Internet of Things. We review and build on existing frameworks for surveys, administrative-, found- and digital trace data. The framework describes steps, concepts and error sources when single-source or multiple-source statistics are produced based on digital or survey data. Blending data sources is a vital step in the framework. Furthermore, the unified framework offers terminology to describe and document errors in digital data, aligned with terminology used in the classical Total Survey Error framework. We also provide indicators for evaluating the quality of statistics produced based on single or multiple data sources.
Induction is a form of reasoning that starts with particular examples and generalizes them into a rule, namely, a hypothesis. However, establishing the truth of a hypothesis is problematic because conflicting events may occur, a difficulty known as the induction problem. The sunrise problem is a quintessential example of probability-based induction. It shows that zero probability must be assigned to the hypothesis that the sun rises forever, regardless of the number of observations. This reveals a fundamental deficiency of probability-based induction: A hypothesis can never be accepted through the Bayes-Laplace approach. Although alternative approaches have been proposed to address this issue, none has fully resolved the deficiency. In this paper, we investigate how this deficiency arises and demonstrate that confidence can overcome it. Confidence not only reconciles the epistemic and the aleatory interpretations of uncertainty but also aligns with the evidence by allowing the hypothesis to be accepted with complete confidence in rational decision-making, while remaining falsifiable if conflicting evidence arises.