
ABSTRACT In many fields of research, it is common to conduct experiments where several observations of the response variable are made on the same subject over non‐random conditions, generating longitudinal studies. A wide variety of techniques deal with this type of data when analysing population variations, many based on estimating functions. It is also common to find longitudinal data where the assumption of asymmetry and/or heavy tails is necessary for the distribution associated with the outcome(s) and also where we do not know the relationship between these outcome(s) and the covariates. In these situations we can use the scale mixture of skew‐normal distributions and the partially linear models. Said that, in this article, we developed estimating functions for the additive partially linear models based on the scale mixture of skew‐normal distributions using a centred parameterisation. This parameterisation does not present inferential problems, as the non‐quadratic form of the log‐likelihood, that the classical skew‐normal distribution can present. Additionally, it provides a better interpretation of the related parameters, since they represent the mean, variance and skewness coefficient. Our methodology focuses on the flexibility of the functional part when using semi‐parametric models and also on the distributional flexibility and robustness when using the aforementioned class of distributions. Additionally, we presented residuals and influence diagnostic tools. A Monte Carlo experiment is conducted to evaluate the performances of the estimators in finite samples. The methodology is illustrated with the analysis of the Framingham cholesterol study.
ABSTRACT In functional data analysis, methods based on basis expansion are already well developed. However, when faced with collinearity in multivariate functional covariates, these methods may not perform effectively. In this paper, by using a new signature‐based method, we introduce a novel method within the scalar‐on‐function regression framework, leveraging signature method and partial least squares (PLS), which we call SigFPLS . This method outperforms traditional functional regression models and the signature‐based functional linear model by offering several key benefits: simple process for parameter tuning, easy implementation and effective handling of multicollinearity in multivariate functional data. Through extensive simulation studies and rigorous real data analysis, we demonstrate the effectiveness of our proposed method with focus on the scalar‐on‐function regression data, especially when dealing with multi‐dimensional and rough functional covariates.
Current bushfire warning systems communicate the probability of a warning given danger, but residents require the probability of danger given a warning. This misalignment, combined with the extreme rarity of catastrophic fires, often leads to the dangerous 'wait and see' behaviour. A necessary preliminary result establishes that the Catastrophic fire danger rating system possesses genuine discriminative skill despite the accuracy paradox: a na & iuml;ve policy of never issuing warnings achieves higher raw accuracy than the actual system; yet, the Peirce Skill Score () demonstrates strongly positive discrimination that no skill-free forecasting strategy can replicate. The ratio constitutes a triple identity-Bayes factor, positive diagnostic likelihood ratio, and cost-loss multiplier-that unifies the Bayesian, medical diagnostic and meteorological forecast verification literature. We then develop a calibrated Bayesian decision-theoretic framework for property-level bushfire warnings using Australian historical data. By specifying priors for daily threat probability (), likelihoods for Catastrophic ratings and asymmetric loss functions, we derive a closed-form optimal evacuation threshold. Under baseline calibration, the posterior probability of threat given a Catastrophic rating rises to only . Crucially, however, under realistic asymmetric losses (where the cost of staying during a threat exceeds the cost of evacuating), the optimal decision threshold is even lower (), so 'leaving early' is rational even when the absolute probability of fire impact is extremely low. Sensitivity analysis demonstrates that this conclusion is robust across two orders of magnitude in loss ratios and wide ranges of prior probabilities, with the threat-to-destruction ratio identified as the key calibration parameter.
Motivated by an application to air quality data, we provide a comprehensive investigation of coherent forecasting techniques for ordinal time series. Here, the coherency requirement expresses that the computed forecast values have to be consistent with the ordered qualitative range of the data-generating process (DGP). Three classes of coherent forecasts are proposed, namely ordinal point forecasts, prediction intervals and probability mass function (pmf) forecasts. In addition, corresponding criteria for evaluating their predictive performance are established. An extensive simulation study analyses the performance of the coherent forecasting techniques across various ordinal DGPs, taking into account the impact of estimated model parameters and model misspecification. The proposed forecasting approaches are then applied to our central motivating application, the real-world time series on ordinal air quality levels and practical implications are discussed. In this context, it is also demonstrated that the coherent point and pmf forecasts can be combined visually in a meaningful and interpretable way. While centrally motivated by environmental monitoring, the proposed framework applies more broadly to ordinal time series arising in other settings where ordered categorical outcomes are used for assessment and decision making.
Joint models have been developed recently for multivariate semicontinuous data; however, joint modelling implies structured covariance and nonnegative correlation among responses, and thus is inflexible in accommodating complex covariance structures in practice. In addition, zero-inflated continuous data have been traditionally handled by two-part mixed models where zero and positive responses are analysed separately; therefore, the multivariate nature of the data would be destroyed. In this paper, we introduce a new model for multivariate semicontinuous data by incorporating distribution-free multivariate random effects of unstructured covariance into Tweedie compound Poisson regression model. An optimal estimation of our model has been developed using the orthodox best linear unbiased predictors of multivariate random effects. Our method is illustrated through the analysis of Mali family farmer data.
The choropleth map is a common tool for communicating spatial distributions across geographic areas. However, the size of geographic units can distort interpretation, influencing how users perceive the distribution. A common alternative is the cartogram, which resizes areas based on population. Yet, in Australia, the stark disparities in geography and population make cartograms less suitable. We explore the hexagon tile map as an alternative. We report results from a task-based experiment involving human participants, using the lineup protocol to assess how well hexagon tile maps and choropleths convey spatial patterns. Three spatial patterns were tested: one reflecting geography, with values increasing monotonically from the northwest to southeast of Australia, and two with clustered high concentrations. Results show that the hexagon tile map outperforms the choropleth map. These findings support the use of alternative map displays and suggest that hexagon tile maps are effective for visualising spatial distributions in heterogeneous regions.
Total variation (TV)-based methods, such as the fused lasso, are standard for change point detection but are impaired by issues like local monotonicity. To address these limitations, this study comparatively analyses the fused lasso with two alternative methodologies. These methodologies are the Penalised Regression Spline (PRS) and a novel Adaptive Penalised Regression Spline (APRS) that incorporates data-adaptive weights. Using an efficient coordinate descent algorithm, we evaluated the three methods on both simulated and real-world data. The comparative analysis, conducted under various Signal-to-Noise Ratio (SNR) conditions, revealed that APRS demonstrated superior performance, particularly in minimising false positives.
This paper develops a unified asymptotic framework for estimating conditional distortion risk measures and conditional expectiles within the APARCH-X model. The proposed framework extends that of Hoga (International Journal of Forecasting 37, 675-686, 2021), which focuses exclusively on extreme tail levels, by allowing for both extreme and intermediate tail regimes through a single asymptotic condition, , where is the risk level and corresponds to the extreme case and to the intermediate case. Building on extreme value theory, we derive the asymptotic distributions of the proposed estimators and show that the limiting variances differ structurally across the two regimes. The extreme case studied in Hoga (2021) is thereby recovered as a special instance. In addition, we establish a uniform limit theory over a continuum of tail levels, which facilitates the construction of confidence bands for conditional tail risk measures. Empirically, we apply the proposed methodology to daily returns of four major equity indices and incorporate Chicago Board Options Exchange (CBOE) volatility indices as exogenous variables. The results show that including these volatility measures leads to statistically significant improvements in forecasting accuracy for conditional tail risk measures, particularly during crisis periods, as evidenced by Diebold-Mariano tests. Moreover, the model confidence set procedure consistently favours the APARCH-X specification over conventional alternatives, including APARCH, AR-GARCH and GARCH models.
Estimation of large covariance matrices plays important roles in high-dimensional data analysis. The sub-Gaussian condition of a random vector is usually assumed, which requires the existence of infinite moments. Avella-Medina et al. (2018) provide an upper bound estimation in the sense of probability over a sparse covariance matrix space under the weak assumption of bounded moments (), see Biometrika, 105, 271-284. In particular, their estimation attains the minimax optimality when . The authors conjecture that their estimation is optimal as well for . In this paper, we first extend their upper bound estimation to a larger space and then prove the optimality of our estimation. This can be considered as a solution to their conjecture. Moreover, we give an optimal estimation in terms of expectation on the same space. Finally, numerical experiments support our theoretical analysis.
In fitting a non-parametric regression model, the spline-mixed model connection is well known, which associates fitting of a penalised spline model with that of a linear mixed model. We establish a theoretical framework for justifying the spline-mixed model connection, and study the asymptotic behaviour of the empirical best linear unbiased predictor of the regression mean and associated model parameters.
R and Python are the two dominant language tools for data science today. Which one is better, and in what senses? This paper explores such questions, in terms of aspects such as learning curve, clarity of expression, coding philosophy, high-performance computing capability and so on. Also, the article treats base-R and the tidyverse as two separate 'dialects' of R, so that the above comparisons are in many cases tripartite in nature.
Density function estimation is a cornerstone of statistical analysis. This paper focuses on the generalised edge frequency polygon estimator, establishing its asymptotic normality for identically distributed -mixing random variables. This finding complements the asymptotic theory outlined by Dong and Zheng (2001. Generalized edge frequency polygon for density estimation. Statistics and Probability Letters, 55, 137-145). Theoretical results are substantiated through simulations that assess finite-sample performance and an analysis of a real-world dataset.
Planning for a future workforce requires forecasts of age structure changes to inform policy decisions, particularly related to universities and immigration. We propose a new dynamic statistical model for forecasting the age structure of a workforce. Our approach is inspired by a stochastic model used in population forecasting, replacing births with graduate entry, modelling exits through death and retirement and including a remainder term that captures migration and career changes. Functional data models are used to model age-specific components, while ARIMA models are used for time series components. Simulation is employed to generate forecast distributions, capturing uncertainty from all components. The approach is illustrated using data on Australia's scientific workforce, allowing us to forecast the age distribution of various scientific disciplines for the next ten years. This analysis was central to an Australian Academy of Science initiative examining the capability of Australia's science system and identifying workforce gaps.
In this article, I trace the evolution of the tidyverse, a cohesive ecosystem of R packages for data science. Beginning with early packages I created during my PhD at Iowa State University, I explain how they coalesced into a unified collection with consistent principles. I detail key innovations including tidy data, tibbles, the pipe operator, tidy evaluation and the role of hex stickers in community building. I describe the transition from my individual efforts to a collaborative enterprise supported by a team at Posit and a vibrant global community. I highlight the importance of human-centred design, consistency, composability and inclusivity as our guiding principles. I discuss how our recent priorities have shifted from innovation to maintenance as the ecosystem has matured. Looking to the future, I discuss our current areas of focus including Positron (a new data science IDE), R in production environments and integrating large language models into data science workflows.
We discuss the development of the package repository of CRAN, its design principles and how these are put into practice. We provide insights into how the regular and submission checks are organised and the actions resulting from these checks. We describe some of the challenges of the current approach, including constraints on the available human resources and how these can possibly be mitigated in order to keep CRAN feasible and successful in the future.
Gain-Probability (G-P) analysis quantifies the probability that a randomly selected individual from one group scores higher or lower than an individual from another group, by varying magnitudes. While G-P methods have been developed under normality and various skewed distributions, symmetric heavy-tailed settings remain largely unexplored, despite their prevalence in finance, environmental science, and other applied domains. We extend the G-P framework to the broad family of scale mixtures of normal (SMN) distributions, including the Student's t, slash, variance gamma (VG), and Pearson Type VII distributions. Analytical expressions for G-P under SMN are derived for both independent and matched data, and parameter estimation is performed using the expectation maximisation (EM) algorithm. Simulation studies show that the proposed estimators are accurate, robust to heavy tails, and improve with sample size, with performance most sensitive to group separation and noise level. An application to daily returns of US and Chinese equity indices demonstrates how G-P analysis captures distributional tail effects that are overlooked by traditional tests. The results support G-P analysis under SMN as a practical, interpretable alternative to significance testing, enabling robust inference for symmetric heavy-tailed data in diverse applied settings.
Model averaging for categorical response variables has gained a lot of attention in recent years. To further improve the prediction accuracy, we present a partially linear multinomial logit model averaging (PLMLMA) technique. Our candidate models are built by selecting each continuous covariate in turn as the index variable of the non-parametric function, thus avoiding both the artificial selection of index variables and the curse of dimensionality in estimation. The model averaging weights are determined by minimising the Kullback-Leibler (KL) loss. We demonstrate asymptotic optimality by showing that the KL loss between the true model and the model where the log-odds ratio is estimated by the 'working' log-odds ratio is asymptotically equivalent to that of the best but impractical model averaging estimator. Furthermore, we establish the convergence rate of the weight estimator without assuming that the true model is included among the candidate models. The superior performance of the proposed method is evidenced by lower mean squared error (MSE) and higher hit rate (HR) in simulations, outperforming various competitors. We also apply our method to wheat variety classification to illustrate the merits of PLMLMA.