
In South Africa, the country with the highest income inequality in the world and an unemployment rate of 32.9%, survey research forms a crucial enabler for evidence-based policymaking to drive inclusive growth (Francis and Webster, 2019; Statistics South Africa, 2025a). For decades, nationally representative household surveys have provided the country with key information on livelihoods and areas for targeted decision making by capturing measures such as household income, food expenditure and employment status. More recently, notable survey research has been undertaken in early childhood education (ECE)-seen through the Thrive by Five Index, a nationally representative survey of child outcomes using ECE tools developed and validated for the South African context. There are several computational methods developed in the literature offering solutions for optimum boundary determination and sample size allocation in the stratified sampling approach to survey research. The uptake of computational methods in the South African context, however, remains limited. Our study offers the first quantitative evaluation of more common stratification approaches used in South Africa in comparison to five prominent computational methods in the literature-random search, genetic algorithm, biased random key genetic algorithm, grouping genetic algorithm, and variable neighbourhood search. The findings indicate that substantial precision gains can be realised when adopting these novel methods. Additionally, this study is the first application of the methods to South African datasets, contributing to the literature using notable, recent research use cases in the country: the Thrive by Five Index 2021, ECD Census 2021, and General Household Surveys of 2023 and 2024. Through comprehensive evaluation, the work offers insights for performing stratified sampling in these applied contexts using existing methods available in the R programming language.
The analysis of multivariate serological data derived from blood serum samples and tested for the presence of antibodies against multiple pathogens gained attention in recent years. Despite the common use of a so-called threshold approach to classify individuals as seronegative or-positive, limitations of such an approach have been reported in the literature, with the subjective choice of the threshold being the most important. Here, we consider a Bayesian mixture approach to model continuous IgG antibody concentrations directly while accounting for the presence of individual heterogeneity and implied association between antibody titer levels for two infections. We fitted the proposed model to Belgian bivariate serological data on the varicella-zoster virus (VZV) and parvovirus B19 (PVB19). Given the existing body of evidence with respect to possible reinfections with PVB19, we investigated whether models explicitly accounting for waning of humoral immunity improved model fit. Our results showed that although after a steep rise with age, the observed seroprevalence for PVB19 decreases between the ages of 20 and 40, the mean IgG antibody concentration remains constant with age among individuals in the seropositive component. This could provide evidence of a direct impact of reinfections with PVB19 on the observed IgG antibody levels, while individuals with loss of humoral immunity after natural infection imply an increase in susceptibility. For VZV, the mean IgG antibody levels slightly decrease with increasing age among seropositive individuals, indicating only very limited waning of humoral immunity as age-dependent seroprevalence estimates are monotonically increasing with increasing age. In general, based on our analyses, we showed that mixture models provide additional insights concerning the waning of humoral immunity as compared to more traditional frailty approaches, which focus on estimating the seroprevalence solely while the model is sufficiently flexible to capture observed dynamics in IgG antibody decay.
Seasonality impacts various industries and sectors, influencing agricultural cycles, economic planning, and healthcare resource allocation. We propose a novel approach using an attribute based fuzzy lattice data structure to create overlapping catchment areas using the fundamentals of label propagation and graph clustering. This approach considers both the link structure and attribute similarities between nodes in a network, where the nodes are points of interest in a road network. Nodes may be close or far apart based on connectivity and shared attributes, such as common interests or in a geographical application considering topography features. In this study, we incorporate static and seasonal attributes for geographical nodes, allowing us to explore seasonal catchment areas and provide a more realistic view of accessibility throughout the year. This integrated approach offers a comprehensive framework for assessing spatial accessibility and understanding seasonal variations in regions to enhance planning for essential services.
Rhino poaching in South Africa continues to threaten the existence of African rhino species. Since poachers often attack wildlife parks frequently, predictive models are essential for exploiting the availability of data to gain information about the poachers. Although a number of statistical methods have been applied to poaching prediction, they either do not take the spatial variation of observations into account, require additional observational data, depend on known priors, or result in models that are overfitted and challenging to interpret. This paper proposes the use of point process models to predict the spatial distribution of poaching activity within a wildlife park. Descriptive statistics of poaching spatial point patterns have been considered, as well as univariate non-parametric kernel density estimation. However, the focus of this work is on fitting multivariate parametric point process models, using a number of environmental factors. Since real-world poaching data could not be obtained for this work, due to the sensitivity of the data, a simulation study is performed, where numerous point patterns are generated from the same underlying point process. The method can be used when no data is available, and is based on environmental preferences of poachers, which can be obtained through expert knowledge, literature reviews, or by making intelligent assumptions. The results indicate that the point process models are able to predict the initial probabilities well, for most data generating processes. Point process models thus appear to be a promising method for predicting the spatial distribution of poaching activity.
In this paper, Bayesian statistical process control limits are derived for Cronbach's coefficient alpha (a) in the case of the balanced one-way random effects model. Cronbach's alpha is one of the most commonly used measures for assessing a set of items' internal consistency or reliability, thereby assessing the assumption that they measure the same latent construct. By using the available data and the Jeffreys independence prior, the posterior distribution of a and the predictive density of a future (unknown) Cronbach's alpha ((alpha)over cap(f)) can be derived. Given a stable Phase I process, the predictive density function (f ((alpha)over cap(f) | (alpha)over cap)) and the conditional predictive density functions (f ((alpha)over cap(f)|alpha)) are used to calculate central values, variances, control limits, run-lengths and the average run-length. The predictive density of a future run-length is the average of a large number of geometrical distributions, each with its own parameter value. Three applications of interest are included in this paper. From the results, it can be seen that the average and median run-lengths are usually larger than the theoretical values. An advantage of the Bayesian procedure, however, is that the control limits, in other words, beta, can be adjusted in such a way that the average or median run-length has a specific value.
We propose new goodness-of-fit tests for the Poisson distribution. The testing procedure entails fitting a weighted Poisson distribution, which has the Poisson as a special case, to observed data. Based on sample data, we calculate an empirical weight function which is compared to its theoretical counterpart under the Poisson assumption. Weighted Lp distances between these empirical and theoretical functions are proposed as test statistics and closed form expressions are derived for L1, L2 and L1 distances. A Monte Carlo study is included in which the newly proposed tests are shown to be powerful when compared to existing tests, especially in the case of overdispersed alternatives. We demonstrate the use of the tests with two practical examples.
Considerable attention in reliability literature has been given to studying various repair models. These models are often described under the assumption of minimal repair. However, repairs of a failed system can be unsuccessful and, therefore, should be repeated until a successful attempt. At certain instances, the assumption of minimal repairs is too restrictive, as the repair can be worse than minimal due to adverse effects of the previous repair attempts, external and internal shocks to a system, insufficient quality of repair, etc. Therefore, in this paper, we use the generalised Polya process to define and derive useful properties for a more general multiple-repair-attempt process with the worse than minimal repair. This has the minimal repair model as a special case. As an application, the corresponding age replacement policy for the general multiple-repair-attempt model is defined and the optimal solutions are obtained. Detailed numerical illustrations support our findings.
In this paper, we propose a natural wavelet estimator of multivariate copula densities as a ratio of the linear wavelet estimator of the underlying joint population density function and the product of the linear wavelet estimators of the corresponding marginal density functions. It is proven that the new estimator attains the optimal almost sure convergence rate over Besov balls for the supremum norm whenever the resolution level is suitably chosen.
In the present work, we propose a method of constructing Archimedean copulas using the total time on test transform, extensively used in reliability modelling. It is observed that the copula can be specified in terms of a univariate life distribution with a finite mean and monotone hazard rate. We discuss some new properties of the Kendall distribution arising from the proposed new generator and the associated measures of dependence.
In this paper, we investigate the almost sure convergence, in supremum norm, of the rank-based linear wavelet estimator for a multivariate copula density. Based on empirical process tools, we prove a uniform limit law for the deviation, from its expectation, of an oracle estimator (obtained for known margins), from which we derive the exact convergence rate of the rank-based linear estimator. This rate reveals to be optimal in a minimax sense over Besov balls for the supremum norm loss, whenever the resolution level is suitably chosen.
We consider a functional linear regression model with a real -valued response and a functional random variable with its derivative as covariates. We are interested in testing the null hypothesis of no covariates effect using a spatially dependent sample. We propose two test statistics which take into account the proximity between sites and we establish the asymptotic normality of cross covariance operator between both interest variables. From this result, we derive asymptotic distributions of these both statistics. Then, we illustrate our test procedure by means of a simulation study.
Hotspot detection in spatial analysis identifies geographic areas with elevated event rates, facilitating more effective policy interventions aimed at reducing such incidents. In the current literature, several methods have been used to detect hotspots such as measures for local spatial association and spatial scan methods. However, the performance of these methods is limited for small-scale hotspots as well as spatial domains where the number of areas is small. In this work, we propose a new approach, making use of the Discrete Pulse Transform (DPT) to decompose spatial lattice data along with the multiscale Ht -index and the spatial scan statistic as a measure of saliency on the extracted pulses to detect significant hotspots. The proposed method outperforms the well -used local Getis-Ord statistic in a simulation study, especially on small-scale hotspots. The method is also illustrated on South African COVID-19 cases and South African crime data.
In some standard applications of spatial point pattern analysis, window selection for spatial point pattern data is complex. Often, the point pattern window is given a priori. Otherwise, the region is chosen using some objective means reflecting a view that the window may be representative of a larger region. The typical approaches used are the smallest rectangular bounding window and convex windows. The chosen window should however cover the true domain of the point process since it defines the domain for point pattern analysis and supports estimation and inference. Choosing too large a window results in spurious estimation and inference in regions of the window where points cannot occur. We propose a new algorithm for selecting the point pattern domain based on spatial covariate information and without the restriction of convexity, allowing for better estimation of the true domain. Amodified kernel smoothed intensity estimate that uses the Euclidean shortest path distance is proposed as validation of the algorithm. The proposed algorithm is applied in the setting of rural villages in Tanzania. As a spatial covariate, remotely sensed elevation data is used. The algorithm is able to detect and filter out high relief areas and steep slopes; observed characteristics that make the occurrence of a household in these regions improbable. Keywords: Covariate, Euclidean shortest path, Nonconvex, Spatial point pattern, Window selection
An automated logistic regression solution framework (ALRSF) is proposed to solve a mixed integer programming (MIP) formulation of the well known logistic regression best subset selection problem. The solution framework firstly determines the optimal number of independent variables that should be included in the model using an automated cardinality parameter selection procedure. The cardinality parameter dictates the size of the subset of variables and can be problem-specific. A novel regression parameter fixing heuristic that utilises a Benders decomposition algorithm is applied to prune the solution search space such that the optimal regression parameter values are found faster. An optimality gap is subsequently calculated to quantify the quality of the final regression model by considering the distance between the best possible log-likelihood value and a log-likelihood value that is calculated using the current parameter values. Attempts are then made to reduce the optimality gap by adjusting regression parameter values. The ALRSF serves as a holistic variable selection framework that enables the user to consider larger datasets when solving the best subset selection logistic regression problem by significantly reducing the memory requirements associated with its mixed integer programming formulation. Furthermore, the automated framework requires minimal user intervention during model training and hyperparameter tuning. Improvements in quality of the final model (when considering both the optimality gap and computing resources required to achieve a result) are observed when the ALRSF is applied to well-known real-world UCI machine learning datasets. Keywords: Best subset selection, Independent variable selection, Logistic regression, Mixed integer programming
In this paper, we present a numerical method based on the fast Fourier transform (FFT) to price call options on the minimum of two assets, otherwise known as two-asset rainbow options. We consider two stochastic processes for the underlying assets: two-factor geometric Brownian motion and three-factor stochastic volatility. We show that the FFT can achieve a certain level of convergence by carefully choos-ing the number of terms and truncation width in the FFT algorithm. Furthermore, the FFT converges at an exponential rate and the pricing results are closely aligned with the results obtained from a Monte Carlo simulation for complex models that incorporate stochastic volatility.
Machine learning and statistical models are increasingly used in a prediction context and in the process of building these models the question of which variables to include often arises. Over the last 50 years a number of procedures have been proposed, especially in the statistical literature. In this paper a newvariable selection procedure is introduced for linear models. A subset of variables is defined here to be “good at margin λ” if it has two properties, namely (i) its associated criterion of fit will be improved in relative terms by less than λ if any variable is added to it, and (ii) its criterion of fit will deteriorate in relative terms by at least λ if any variable inside it, is dropped from it. Thus, such a subset contains all variables that are individually important and none that are unimportant at a given margin λ ≥ 0. This paper discusses calculation of such λ-good subsets. The “good” approach extends readily to generalised linear and many other models by using an appropriate criterion of performance. The approach is illustrated on an artificial data set and a number of real data sets. Keywords: Good subsets, Linear regression, Logistic regression, Robust regression, Subset selection, Variable importance, Variable selection
The problem of multiple hypothesis testing with correlated test statistics is a very important problem in statistical literature. Specifically, we consider the case when the joint distribution of the test statistics is a multivariate normal distribution with an unknown mean vector and compound symmetric correlation structure. Our goal is to identify nonzero entries of the mean vector. Bogdan et al. (2011) solved this problem when test statistics are independent normals along with the study of asymptotic optimality in a Bayesian decision theoretic sense. The case under dependence was left as a challenging open problem. The solution is intuitive and permutation invariant, does not assume sparsity unlike Bogdan et al. (2011) and is validated through simulation studies.
We propose a method for variable selection in multivariate regression with random predictors. This method is based on a criterion that permits to reduce the variable selection problem to a problem of estimating a suitable set. Then, an estimator for this set is proposed and the resulting method for selecting variables is shown to be consistent. A simulation study that permits to study several properties of the proposed approach and to compare it with existing methods is given.
Since the introduction of factorisation machines in 2010, it became a popular prediction technique amongst machine learners who applied the method with success in several data science challenges such as Kaggle or KDD Cup. Despite these successes, factorisation machines are not often considered as a modelling technique in business, partly because large companies prefer tried and tested software for model implementation. Popular modelling techniques for prediction problems, such as generalised linear models, neural networks, and classification and regression trees, have been implemented in commercial software such as SAS which is widely used by banks, insurance, pharmaceutical and telecommunication companies. To popularise the use of factorisation machines in business, we implement algorithms for fitting factorisation machines in SAS. These algorithms minimise two loss functions, namely the weighted sum of squared errors and the weighted sum of absolute deviations using coordinate descent and nonlinear programming procedures. Using a simulation study, the above-mentioned routines are tested in terms of accuracy and efficiency. The prediction power of factorisation machines is then illustrated by analysing two data sets. Keywords: Factorisation machines, Fitting algorithms, Parameter estimation
Public health surveillance of overweight prevalence is essential to assess the extent of the problem, identify regions and groups most affected and inform policy-making. However, the needed reliable data at disaggregated levels is lacking in Kenya. The Kenya STEPwise Survey for Non-communicable Diseases and RiskFactors (KSSNDRF) was nationally representative. It was used to obtain various indicators of non-communicable diseases and risk factors including overweight. However, due to small sample sizes at lower levels like at the county, overweight prevalence estimates are statistically imprecise (i.e., high variance). Therefore, to increase the effective sample size we combine data from the KSSNDRF and the Kenya Population and Housing Census by model-based small area methods. In particular, we fit an arcsine square-root transformed Fay–Herriot model. To transform back to the original scale, we use a bias-corrected back transformation. For this model, we smooth the design variance using Generalised Variance Functions. We compute the mean squared error estimates using a bootstrap procedure. We found that counties within urban areas — including the major towns like Nairobi, Nakuru, Nyeri and Mombasa — have a higher prevalence of overweight compared to rural counties. Although the paper focuses on overweight prevalence in Kenya, the presented method can also be applied to other indicators in developing countries with similar data sources.