The cryptocurrency market offers significant investment opportunities, but with high levels of financial risk compared to traditional asset classes. This study analyzes the daily returns of Bitcoin and Ethereum, focusing on tail behavior and peakedness to assess risk. Using the flexible Generalized Tempered Stable (GTS) distribution, we capture significant deviations from normality. Results show Bitcoin returns are more concentrated around the mean, with 80 % of returns between $-1.27 \%$ and 2.84 %, while Ethereum is more dispersed-only 40 % of its returns fall in that range. Bitcoin's distribution is more sharply peaked; Ethereum has heavier tails and greater exposure to extreme fluctuations. These findings underscore the importance of using advanced models like the GTS for accurate risk management and portfolio optimization in cryptocurrencies.
This study develops a hybrid framework integrating ensemble learning with explainable artificial intelligence to address the methodological challenge of balancing predictive accuracy and interpretability in credit risk model comparison. Using the German Credit Dataset, we implemented a comprehensive preprocessing pipeline, including feature encoding, scaling, and SMOTE for class imbalance handling. Four base models, logistic regression, Random Forest, XGBoost, and Multilayer Perceptron, were combined through a Stacked Ensemble with a logistic regression meta learner. The ensemble demonstrated strong performance, achieving an AUC of 0.761, precision of 0.783, recall of 0.806, and an F1 score of 0.794, which represented the highest scores among all models tested. Notably, Random Forest (AUC = 0.749) surpassed XGBoost (AUC = 0.733), challenging conventional algorithmic hierarchies. SHAP analysis provided transparent global and local interpretability, identifying Current Account status (SHAP = 0.153), Loan Duration (0.064), and Savings Account (0.063) as dominant predictor variables. Class-imbalance handling and threshold optimisation enhanced practical utility by reducing false positives from 39 to 16, thereby aligning with financial risk priorities. The framework provides a reproducible methodological pipeline for systematically comparing credit scoring approaches, demonstrating how predictive performance can be evaluated alongside interpretability considerations within a benchmark dataset context.
This study investigates the effectiveness of Bayesian probabilistic methods for stock price forecasting on the Johannesburg Stock Exchange by implementing and comparing Gaussian process regression (GPR), Bayesian long short-term memory (Bayesian LSTM), and Bayesian neural networks (BNNs). Using daily open, high, low, close, and volume (OHLCV) data and engineered technical indicators for FirstRand and Discovery from January 2005 to June 2025 (5187 observations), models were trained and evaluated with the mean absolute error (MAE), root mean squared error (RMSE), and mean squared error (MSE). The GPR produced reliable, well-calibrated intervals in relatively stable regimes, but its performance degraded on the more volatile Discovery series. Bayesian LSTM delivered conservative uncertainty estimates with wide predictive intervals but showed the largest point forecast errors. The BNNs achieved the best balance between accuracy and uncertainty quantification, producing the lowest errors for FirstRand and competitive performance for Discovery. Comparative analysis indicates that BNNs are most suitable when point accuracy and calibrated uncertainty are both priorities, GPR is valuable for smaller or more stable data regimes, and Bayesian LSTM is preferable where conservative, risk-conscious intervals are required. This study highlights the practical value of embedding uncertainty into financial forecasts and recommends matching Bayesian model choice to market volatility, data availability, and decision maker risk appetite.
This paper identifies and characterizes the background driving Lévy process (BDLP) associated with the Generalized Tempered Stable (GTS) distribution, a flexible seven-parameter family of infinitely divisible distributions with applications in physics and quantitative finance. We show that the corresponding BDLP is a finite-variation, infinite-activity Type B Lévy process and derive its explicit background driving characteristic exponent function (BDCEF). The resulting BDLP provides a unified representation that encompasses several important special cases, including the bilateral stable, bilateral Gamma, and Variance Gamma distributions. Building on these results, we develop a simulation framework based on a stationary Ornstein–Uhlenbeck (OU)-type process driven by the GTS BDLP. The mean-reversion speed parameter of the OU process is calibrated using maximum likelihood estimation applied to daily return data from the SPY ETF and Ethereum over the period 2010–2024. The proposed simulation methodology produces realistic daily cumulative return trajectories, and comprehensive numerical error analyses demonstrate the accuracy and efficiency of the resulting discretization scheme. These findings provide both a theoretical extension of Lévy-driven OU models and a practical framework for simulating complex financial return dynamics.
This paper investigates extreme risk in cryptocurrency markets by comparing Bitcoin and Ethereum daily returns with those of S&P 500 and SPY ETF. Using the Generalized Tempered Stable (GTS) distribution to model heavy tails and Quantile-Quantile (Q-Q) plots to assess fitness, we find that all assets deviate sharply from normal distribution. Within this framework, Ethereum exhibits a higher frequency of extreme returns than Bitcoin, highlighting differences in risk profiles even among leading cryptocurrencies.
This study reviews advanced extreme value theory techniques and applies them to extreme rainfall events recorded at two meteorological stations, Port Edward and Virginia, in the KwaZulu-Natal province of South Africa. The study aims to provide a comparative analysis of the performance of three extreme value theory models—the generalised extreme value distribution (GEVD), the generalised extreme value distribution for r-largest order statistics (GEVDr), and the blended generalised extreme value distribution (bGEVD)—in modelling extreme rainfall events. The monthly maximum rainfall data used in the study was obtained from the South African Weather Service. The Shapiro–Wilk test demonstrated the non-normality of the rainfall datasets. Parameter estimation was performed using maximum likelihood estimation and Bayesian estimation methods, both yielding positive shape parameters consistent with the Fréchet class of distributions. The goodness-of-fit tests confirmed the suitability of the GEVD model for the data. The results of both the standard GEVD and GEVDr models provided consistent return level estimates, suggesting strong model performance. The bGEVD model produced lower return level estimates compared to the GEVD and GEVDr models. Overall, the findings of the study offer valuable insights into the behaviour of extreme rainfall in KwaZulu-Natal province, with significant implications for risk management, infrastructure planning, and disaster preparedness. This study will add value to the literature and knowledge of statistics.
The increasing global reliance on wind and solar energy underscores the critical vulnerability of renewable systems to extreme weather, which can severely disrupt power generation. Accurately modelling the complex, multivariate dependencies of weather extremes is essential for building grid resilience, yet conventional statistical models often fail to capture critical tail dependencies. This study aims to develop a robust framework using vine copulas to model the tail dependencies among key meteorological variables, extreme temperature, wind speed, and relative humidity, across the Eastern Cape province, South Africa, in order to identify optimal seasons for renewable energy production. We first clustered weather stations across the province into five distinct groups using Partitioning Around Medoids (PAM), based on geographical features (elevation, longitude, and latitude). This study explored an automatic selection of the optimal vine copula structure that adequately describes the dependence structure of the meteorological variables employed. The analysis demonstrated that R-vine copulas successfully captured the multivariate tail behaviour of temperature and relative humidity, while D-vine copulas were highly effective for wind speed. The models revealed significant tail dependencies, indicating a high potential for concurrent extreme weather events that impact energy generation. Our findings confirm that vine copulas offer a superior framework for assessing the risks associated with extreme weather to renewable energy systems. The results provide critical insights for regional energy policy and grid resilience planning, highlighting the importance of advanced risk assessment to safeguard renewable energy production against climate extremes.
The escalating frequency and intensity of extreme rainfall events driven by climate change threaten infrastructure resilience and societal safety, underscoring the urgent need for robust models to predict these events. Previous studies on the integration of Extreme Value Theory (EVT) and machine learning in modelling extreme rainfall events have not explored the use of a time-varying threshold. This study introduces a novel time-varying threshold Generalised Pareto (GP) regression tree for modelling extreme rainfall in Durban, South Africa. The proposed hybrid model combines EVT with covariate-driven regression tree partitioning, allowing the threshold to evolve dynamically with meteorological conditions. Using daily rainfall and meteorological covariate data from 1981 to 2025, the model was developed, pruned, and benchmarked against a static-threshold GP regression tree and a time-varying threshold Generalised Pareto Distribution (GPD). Evaluation based on the Bayesian Information Criterion (BIC) and log-likelihood demonstrated the superior performance of the proposed model in capturing covariate-driven heterogeneity and temporal variability of rainfall extremes. Four distinct climatic regimes with different tail behaviours and return levels were identified. This study provides the first meteorological application of a time-varying threshold GP regression tree and offers practical insights into flood risk assessment and climate resilience planning in the city of Durban.
Extreme events of climatic nature such as floods, heat waves and cold waves have become more frequent worldwide in recent years than in the past decades. Many studies have attempted to model these extreme events using various geohazards, environmental, mathematical, and statistical techniques without conclusive findings. Extreme value theory (EVT) is a well-established branch of statistics used to study extreme or rare events. Several studies that have applied EVT to climatic events modelled these events using at-site and regional approaches. These approaches have performed fairly well but lacked the ability to address the spatio-temporal nature associated with climatic events such as precipitation, floods, droughts, wind and heat waves. Other EVT studies have attempted to extend these studies to multivariate extremes, but still within either the at-site or regional realisations. The present study looks at the possibility of application of spatio-temporal extreme value theory to the complex extreme value events in the Southern African Development Community (SADC) region in a changing climate. This study is motivated by the existence of spatio-temporal variability of extreme weather events in the SADC region revealed from the previous studies, particularly with reference to the 1991/1992 severe droughts in the SADC region, the year 2000 disastrous floods in the region and the recent 2019 disastrous floods caused by Cyclone Idai which affected parts of Malawi, Mozambique, South Africa, and Zimbabwe. Although the present study is theoretical and does not present empirical findings in its current form, it proposes and discusses in detail EVT techniques such as hierarchical modelling for spatial extremes and extension to spatio-temporal processes with reference to precipitation in the SADC region.
The Generalized Tempered Stable (GTS) distribution extends classical stable laws through exponential tempering, preserving the power-law behavior while ensuring finite moments. This makes it especially suitable for modeling heavy-tailed financial data. However, the lack of closed-form densities poses significant challenges for simulation. This study provides a comprehensive and systematic comparison of GTS simulation methods, including rejection-based algorithms, series representations, and an enhanced Fast Fractional Fourier Transform (FRFT)-based inversion method. Through extensive numerical experiments on major financial assets (Bitcoin, Ethereum, the S&P 500, and the SPY ETF), this study demonstrates that the FRFT method outperforms others in terms of accuracy and ability to capture tail behavior, as validated by goodness-of-fit tests. Our results provide practitioners with robust and efficient simulation tools for applications in risk management, derivative pricing, and statistical modeling.
Accurate parameter estimation is essential in diverse fields such as finance, ecology, and physics, where stochastic processes form a fundamental component. This study aims to compare resampling estimation methods including, bootstrap, jackknife, and jackknife-after-bootstrap for parameter estimation in the Ornstein-Uhlenbeck (OU) process. The findings from the results indicated that the bootstrap method outperforms the other resampling methods with respect to bias estimation and standard error. Consequently, the bootstrap method provided more precise standard error estimates (average of 0.08) compared to jackknife (0.12) and jackknife-after-bootstrap (0.11). The Jarque-Bera test as well as Q-Q plot of normality confirmed that applying the bootstrap method to estimate the OU parameters of the model is adequate. It is, therefore, recommended to implement the bootstrap resampling method on real-life data (preferably lending rate) to demonstrate its behavior in estimating mean-reverting parameters and also serve as validation for the efficacy of our methodology.
Subjective Bayesian methods, which incorporate expert knowledge into disease modelling, remain underutilised in epidemiology. This is despite the growth of knowledge in statistical approaches to disease analysis. Objective priors are often favoured for their simplicity. However, subjective Bayesian approaches can produce more informative models by using expert insights, such as those related to malaria transmission. This study focuses on translating expert knowledge into prior probability distributions through a process known as prior elicitation. Prior elicitation presents several challenges. Converting expert judgments into probability distributions is complex and often requires specialised tools. Established methods like the Sheffield approach are resource-intensive, requiring considerable time and cognitive effort from both experts and researchers. To address these limitations, this study makes a major contribution by integrating the Analytic Hierarchy Process (AHP) with statistical validation techniques to quantify expert knowledge into prior probability distributions. Expert knowledge was collected through questionnaires and structured as pairwise comparisons. These were quantified into AHP weights, representing the relative importance of environmental factors influencing malaria transmission. The weights were then fitted to various probability distributions and evaluated using goodness-of-fit tests. Results showed that the beta, gamma and normal distributions best represented the elicited expert knowledge. These findings suggest that beta, gamma and normal distributions are suitable as prior distributions in Bayesian models of malaria transmission. By simplifying the elicitation process and reducing technical complexity, this approach offers a practical framework for applying subjective Bayesian methods in epidemiology. Future research will compare these elicited priors with objective priors to evaluate their impact on model performance across domains.
The paper presents an enhanced numerical framework for computing the one-dimensional fast Fractional Fourier Transform (FRFT) by integrating closed-form Composite Newton-Cotes quadrature rules. We show that a FRFT of a QN-length weighted sequence can be decomposed analytically into two mathematically commutative compositions: one involving the composition of a FRFT of an N-length sequence and a FRFT of a Q-length weighted sequence, and the other in reverse order. The composite FRFT approach is applied to the inversion of Fourier and Laplace transforms, with a focus on estimating probability densities for distributions with complex-valued characteristic functions. Numerical experiments on the Variance-Gamma (VG) and Generalized Tempered Stable (GTS) models show that the proposed scheme significantly improves accuracy over standard (non-weighted) fast FRFT and classical Newton-Cotes quadrature, while preserving computational efficiency. The findings suggest that the composite FRFT framework offers a robust and mathematically sound tool for transform-based numerical approximations, particularly in applications involving oscillatory integrals and complex-valued characteristic functions.
The study explores the relationship between economic growth, employment, and forest conservation in sub-Saharan Africa. Using a semi-parametric regression approach, we found distinct economic growth patterns across different regions, emphasizing the need for region-specific development strategies. Our findings revealed a significant link between economic growth and unemployment, underscoring the importance of job creation in economic expansion. We also identified complex relationships between economic growth and forest areas, highlighting the necessity for sustainable development. This research offers valuable insights for policymakers promoting job creation, sustainable development, and inclusive growth in sub-Saharan Africa.
Binary classification using Artificial Neural Networks (ANN) is a fundamental problem in machine learning, and the choice of activation functions plays a vital role in determining the performance of the model. This study investigates the performance of various activation functions for binary classification tasks using neural networks across multiple datasets. The results show that ReLU and Tanh, compared to Logistic, consistently excel in accuracy, precision, recall, F1-Score, and AUC-ROC, making them versatile choices for diverse tasks. However, Identity exhibits variable performance, highlighting the need for careful activation function selection based on task specifics. Additionally, the study emphasizes the impact of the dataset size on activation function performance, with ReLU and Tanh offering consistency across varying data volumes. Practitioners can leverage these insights to optimize neural network designs, improving model efficacy in binary classification tasks.
Classical binomial interval methods often exhibit poor performance when applied to extreme conditions, such as rare-event scenarios or small-sample estimations. Recent frequentist and Bayesian approaches have improved coverage in small samples and rare events. However, they typically rely on fixed error margins that do not scale with the magnitude of the proportion. This distorts uncertainty quantification at the extremes. As an alternative method to reduce these boundary distortions, we propose a novel hybrid approach. It blends Bayesian, frequentist, and approximation-based techniques to estimate robust and adaptive intervals. The variance incorporates sampling variability, Wilson score margin of error, a tuned credible level, and a gamma regularization term that is inversely proportional to sample size. Extensive simulation studies and real-data applications demonstrate that the proposed method consistently achieves better coverage proportions at all sample sizes and proportions. It provides more conservative interval widths below a sample size of 50 and competitively narrower widths from moderate to large sample sizes, especially beyond 50, compared to the Jeffreys’ and Wilson score intervals. Geometric analysis of the tuning curves demonstrates how the blended method adaptively tunes credible levels across binomial extremes. It starts at higher values for small samples and gradually flattens into near-linear, symmetric trajectories as sample size increases. This ensures robust coverage and balanced sensitivity. Our method offers a theoretically grounded, computationally efficient, and practically robust estimation of rare-event intervals. These intervals have applications in safety-critical reliability, epidemiology, and early-phase clinical trials.
In this paper we assess the impact of the 2010 Federation of International Football Association (FIFA) World Cup tournament on the monthly hotel accommodation income. Seasonal autoregressive integrated moving average (SARIMA) and the Box-Tao intervention analysis were applied to the data which is from Statistics South Africa is for the period 2005 to 2012. The SARIMA model is fitted to the pre-intervention period that is for the period before the world cup. The intervention model is applied using the pulse and step functions to assess if the impact of the FIFA World Cup was temporary or permanent, respectively. The pulse model has a lower AIC and hence adjudged to be a better fit for the data than the step. Empirical evidence suggest that the 2010 FIFA World Cup has an abrupt and temporary impact on the monthly hotel accommodation income in South Africa for the reviewed period.
This study proposes a more robust methodological approach to modeling the effect of weather and calendar variables on the number of bike rentals. We employ penalized splines quasi-Poisson regression (a semi-parametric model), which involves some form of regularization, like those used in lasso, ridge, and other types of parametric regularization models. We demonstrate that this modeling approach reveals hidden relationships that a pure parametric model fails to identify.The findings show that visibility, windspeed, season, working day, and year all significantly impact bike rentals. Increased rentals are associated with increased visibility and lower wind speed. Rentals are negatively affected by the spring and winter seasons, while working days and the year show positive trends except in a few cases. The analysis of rentals by registered and casual users reveals similar patterns, though the magnitudes of the effects differ. These findings highlight the importance of considering weather and calendar variables when managing and promoting bike-sharing services. The study has implications for bike-sharing system operators and policymakers, suggesting strategies such as improving visibility and wind protection, seasonally tailoring promotional campaigns, targeting non-working days for casual users, and adapting to changing user demands. The study adds to our understanding of the factors that influence bike rentals and provides suggestions for improving the utilization and accessibility of bike-sharing systems.
This paper proposes and implements a methodology to fit a seven-parameter Generalized Tempered Stable (GTS) distribution to financial data. The nonexistence of the mathematical expression of the GTS probability density function makes maximum-likelihood estimation (MLE) inadequate for providing parameter estimations. Based on the function characteristic and the fractional Fourier transform (FRFT), we provide a comprehensive approach to circumvent the problem and yield a good parameter estimation of the GTS probability. The methodology was applied to fit two heavy-tailed data (Bitcoin and Ethereum returns) and two peaked data (S&P 500 and SPY ETF returns). For each historical data, the estimation results show that six-parameter estimations are statistically significant except for the local parameter, μ. The goodness of fit was assessed through Kolmogorov–Smirnov, Anderson–Darling, and Pearson’s chi-squared statistics. While the two-parameter geometric Brownian motion (GBM) hypothesis is always rejected, the GTS distribution fits significantly with a very high p-value and outperforms the Kobol, Carr–Geman–Madan–Yor, and bilateral Gamma distributions.
This study investigates wind speed prediction using advanced machine learning techniques, comparing the performance of Vanilla long short-term memory (LSTM) and convolutional neural network (CNN) models, alongside the application of extreme value theory (EVT) using the r-largest order generalised extreme value distribution (GEVDr). Over the past couple of decades, the academic literature has transitioned from conventional statistical time series models to embracing EVT and machine learning algorithms for the modelling of environmental variables. This study adds value to the literature and knowledge of modelling wind speed using both EVT and machine learning. The primary aim of this study is to forecast wind speed in the Limpopo province of South Africa to showcase the dependability and potential of wind power generation. The application of CNN showcased considerable predictive accuracy compared to the Vanilla LSTM, achieving 88.66% accuracy with monthly time steps. The CNN predictions for the next five years, in m/s, were 9.91 (2024), 7.64 (2025), 7.81 (2026), 7.13 (2027), and 9.59 (2028), slightly outperforming the Vanilla LSTM, which predicted 9.43 (2024), 7.75 (2025), 7.85 (2026), 6.87 (2027), and 9.43 (2028). This highlights CNN’s superior ability to capture complex patterns in wind speed dynamics over time. Concurrently, the analysis of the GEVDr across various order statistics identified GEVDr=2 as the optimal model, supported by its favourable evaluation metrics in terms of Akaike information criteria (AIC) and Bayesian information criteria (BIC). The 300-year return level for GEVDr=2 was found to be 22.89 m/s, indicating a rare wind speed event. Seasonal wind speed analysis revealed distinct patterns, with winter emerging as the most efficient season for wind, featuring a median wind speed of 7.96 m/s. Future research could focus on enhancing prediction accuracy through hybrid algorithms and incorporating additional meteorological variables. To the best of our knowledge, this is the first study to successfully combine EVT and machine learning for short- and long-term wind speed forecasting, providing a novel framework for reliable wind energy planning.