
In this study, we propose the use of the concentration function as a novel tool for assessing the adequacy of predictive distributions in a Bayesian framework. The methodology relies on a rolling-origin evaluation strategy, in which the available time series is partitioned into two segments: the first is used for model estimation, while the second is used for out-of-sample prediction. The concentration function is then constructed for each set of forecasts, providing a distributional assessment of predictive performance across competing models. In this framework, the model with the least inequality in its concentration function is interpreted as offering superior predictive fit. To complement this curve-based analysis, we compute two well-established inequality measures-the Gini concentration coefficient and the Pietra index-which provide further evidence and consistency regarding model performance. These measures allow for a comprehensive evaluation of how closely predictive distributions align with observed outcomes. The proposed methodology is validated through a simulation study, in which its ability to discriminate between models is assessed under controlled conditions. We illustrate its practical applicability by analyzing real-world data, thus demonstrating the usefulness of concentration functions and associated inequality measures as diagnostic tools for Bayesian predictive evaluation.
A class of exchangeable multivariate max-id distributions with 1-norm pound symmetric exponent measure was introduced by Genest, Neslehova and Rivest in a paper published in the journal Bernoulli in 2018. Three years later, an extended class of exchangeable multivariate max-id distributions with p-norm pound symmetric exponent measure was proposed by Mai and Wang in an article which appeared in the Journal of Multivariate Analysis. A new class is proposed here which encompasses them both and which allows for non-exchangeability, thereby providing extra flexibility for modeling multivariate block maxima data and extreme risks. Some properties of members of this class are studied, and conditions are given under which they are multivariate extreme-value distributions. The maximum attractor of each class member is also determined under broad conditions, and an algorithm due to Jan-Frederik Mai is adapted for simulation purposes.
We study early warning of short-horizon volatility spikes by combining three complementary blocks of information: functional data analysis summaries based on functional principal component analysis scores and derivative norms, multi-scale topological data analysis descriptors derived from persistent homology, and heterogeneous autoregressive lagged-volatility features. All models are trained under a strict chronological design, with operating thresholds pre-specified from out-of-fold scores to ensure that model selection and evaluation remain separate. Statistical significance is assessed using paired DeLong tests for the receiver operating characteristic area under the curve and is complemented by a calendar block bootstrap that more directly accounts for temporal dependence in the scores. Across four large-capitalization equities (INTC, COP, NVDA, DIS), the combined specification (BOTH, defined as functional data analysis features plus topological data analysis features plus heterogeneous autoregressive features) generally outperforms the functional data analysis-only and topological data analysis-only baselines in receiver operating characteristic area under the curve and precision-recall area under the curve, and improves their probabilistic accuracy and calibration while remaining competitive with a heterogeneous autoregressive-only gradient-boosted tree benchmark. In a regime-switching generalized autoregressive conditional heteroskedasticity stress experiment, a heterogeneous autoregressive-based gradient-boosted tree model attains the highest discrimination, illustrating that the empirical gains from combining functional and topological information are not guaranteed when the data-generating mechanism is close to a low-dimensional lag-volatility structure. Methodologically, we propose a time-respecting method that integrates functional and topological geometry with paired inference under temporal dependence. Practically, we provide a reproducible script for early warning and for monitoring short-run volatility stress.
We develop a new method based on the Lq-likelihood to perform robust covariance estimation on spatial data. We propose the maximum composite Lq-likelihood estimator (MCLqE), which is obtained from a combination of the maximum Lq-likelihood and the composite likelihood. In order to down-weight a single value at a spatial location, we make use of the composite likelihood of the spatial data, which is in the form of the product of the pairwise likelihood of all pairs of the observations. We apply the maximum Lq-likelihood on each pair to form the MCLqE, which effectively down-weights abnormal observations in a spatial dataset and gives more robust covariance parameter estimation results when the data are contaminated by outliers. We adapt the hyperparameter tuning procedure developed for the MLqE to the new MCLqE framework. The robustness of the MCLqE is further verified by both simulation studies and experiments on a precipitation dataset from the US.
Reliable quantification of treatment benefit in late-phase clinical trials increasingly requires modeling patient histories that include progression, adverse events, and treatment switches. Conventional multi-state analyses often invoke the Markov property and assume independent right censoring-conditions rarely satisfied in oncology, immunology, or cell-therapy programs, where intermediate events and informative dropout are common. This article presents a systematic review and bibliometric synthesis of 48 peer-reviewed studies published through 11 June 2025 that (i) relax the Markov assumption and (ii) address complex observation schemes such as left truncation, interval censoring, or informative censoring, identified through Web of Science and Scopus searches following preferred reporting items for systematic reviews and meta-analyses 2020 guidelines.A recurring set of methodological strategies emerges across the literature, including semi-Markov transition-intensity models, illness-death and semi-competing risks frameworks, landmarking for dynamic prediction, and inverse-probability-of-censoring weighting. Estimation approaches range from nonparametric product integrals to semiparametric weighted likelihoods and Bayesian Markov chain Monte Carlo, with recent contributions exploring saddle-point approximations and subsampling for large-scale electronic health records. To complement this synthesis, we include a compact simulation contrasting baseline and landmark Aalen-Johansen estimators under semi-Markov dynamics with history-dependent censoring, and a bibliometric network analysis mapping collaboration patterns, thematic clusters, and structural gaps. The findings highlight the need for scalable, auditable software, robust diagnostics aligned with the International Council for Harmonization E9(R1) estimand framework (which links clinical trial objectives to precise statistical targets), and better integration of high-dimensional biomarkers; limitations include the English-language restriction and reliance on bibliometric meta-data. Addressing these priorities may enhance both the methodological robustness and regulatory applicability of non-Markov survival models.
The Ziggurat algorithm is a well-established rejection-sampling method designed for the efficient generation of pseudo-random numbers from unimo dal distributions, particularly the standard normal. In this work, we extend and adapt the Ziggurat algorithm to enable the tail-adaptive generation of random numbers from the gamma-order generalized normal distribution -a flexible family characterized by a tail-shaping parameter that governs transitions between light, Gaussian, and heavy-tailed regimes. The resulting algorithm retains the computational speed of the original Ziggurat algorithm while supporting both univariate and multivariate implementations. This extension is especially relevant in simulation-intensive contexts, such as Bayesian modeling, quantitative finance, and machine learning. We provide the mathematical foundation, reproducible implementation details, and extensive benchmarking results that validate the method's efficiency and accuracy. A multivariate extension based on radial decomposition is also introduced, demonstrating the feasibility of generating random variables from symmetric multivariate distributions in practice. To illustrate the practical utility of the properformance across various shape and scale configurations. Additionally, we apply the method to real-world data from biomedical signal processing, highlighting its robustness and adaptability to empirical settings where tail behavior plays a crucial role.
Bootstrap-based inference is investigated for non-nested hypothesis testing in beta regression models with small samples. Specifically, the J and MJ tests are considered, which assess whether one model outperforms its non-nested alternatives by augmenting it with predictors from competing specifications. Their standard bootstrap and fast double bootstrap versions are examined. Monte Carlo simulations reveal that conventional asymptotic tests suffer from size distortions in finite samples. The fast double bootstrap versions of the J and MJ tests yield more accurate inference while requiring minimal additional computational cost. The theoretical results are supported by an empirical application that illustrates the practical value of the proposed methods for model selection in beta regression.
Financial literacy and education have become pivotal on the global policy agenda, as reflected in initiatives such as the Organisation for Economic Co-operation and Development / International Network on Financial Education High-Level Principles on National Strategies for Financial Education. To improve financial well-being, Caja de Compensacion Los Andes and Tapp implemented the Financial Inclusion I90 Program (I90 Program) in Los Muermos, Chile. However, challenges related to participant consent and compliance within the treatment-control design introduced biases that compromised the study's internal validity. This article presents an alternative methodology for impact evaluation based on partial-identification techniques to address these uncertainties and biases. By modeling a range of plausible counterfactual behaviors, the proposed approach delineates partial-identification regions for key probabilities, thereby quantifying the extent of uncertainty in the findings. Furthermore, it emphasizes that impact evaluation should inform policy decisions -consistent with Neyman's concept of inductive behavior-rather than merely predict outcomes in similar contexts. Ultimately, the method offers a framework on evaluation by incorporating policymakers' beliefs and explicitly acknowledging inherent uncertainties. By quantifying potential risks through partial-identification regions, it enables more informed and flexible policy decisions based on a realistic appraisal of implementation challenges.
Zero-inflated regression models are widely used to relate covariates to count outcomes that exhibit an excess of zeros. In their marginal formulations, these models yield direct inference on the population mean instead of on an unobserved susceptible subpopulation. The recently proposed marginalized zero-inflated Bell (MZIBell) model provides a parsimonious alternative for highly overdispersed counts, yet the available theory assumes that all covariates are fully observed. This article extends the MZIBell framework to settings in which covariates are missing at random. Parameter estimation is carried out through parametric inverse probability weighting, and the resulting estimators are shown to be consistent and asymptotically normal under standard regularity conditions. one third of the covariate information is absent. An application to physician-visit counts demonstrates that the proposed inverse-probability-weighted MZIBell model achieves a superior fit, according to the Akaike and Bayesian information criteria, when compared with competing marginalized Poisson and negative binomial specifications.
In this paper, we derive closed-form estimators for the parameters of certain exponential family distributions through the maximum a posteriori (MAP) equations. A Monte Carlo simulation is conducted to assess the performance of the proposed estimators. The results show that, as expected, their accuracy improves with increasing sample size, with both bias and mean squared error approaching zero. Moreover, the proposed estimators exhibit performance comparable to that of traditional MAP and maximum likelihood (ML) estimators. A notable advantage of the proposed method lies in its computational simplicity, as it eliminates the need for numerical optimization required by MAP and ML estimation.
This article investigates the statistical and asymptotic properties of first-order periodic autoregressive conditional heteroscedasticity (P-ARCH(1)) models, emphasizing both theoretical developments and empirical insights. The study introduces the class of PARCH models, highlighting how periodicity allows for distinct volatility dynamics across different seasons or regimes. By employing a vector autoregressive representation of the squared process, key probabilistic properties are derived under periodic stationarity conditions. A moment-based estimation approach using periodic Yule-Walker equations is proposed, and its consistency and asymptotic normality are established. Monte Carlo simulations assess the finite-sample performance of these estimators, particularly in comparison with least squares methods, and provide practical guidelines for implementation. The methodology is further illustrated through an empirical application to monthly log-stock returns of Intel Corporation, demonstrating that periodic ARCH models effectively capture time-varying heteroscedasticity driven by recurring seasonal or cyclical factors, outperforming approaches that assume constant volatility parameters. This study underscores the flexibility of P-ARCH(1) models in financial and economic applications where volatility patterns repeat over known intervals and highlights the advantages of a moment-based estimation strategy as a computationally efficient alternative to quasi-maximum likelihood methods. The findings provide clear directions for future research in extending the periodic ARCH framework to accommodate more complex dynamics and additional stylized features of real-world time series.
Goodness-of-fit tests based on the likelihood ratio are widely used to assess whether a given probability distribution adequately describes observed data. However, certain likelihood ratio-based tests do not have known asymptotic distributions, making it necessary to rely on pre-tabulated critical values obtained through Monte Carlo simulations. A major limitation of this approach is that practitioners must generate additional critical values via simulations for sample sizes not explicitly tabulated, which restricts the applicability of these tests in practice. This study addresses this limitation by developing asymptotic critical value functions for likelihood ratio-based goodness-of-fit tests under the exponential distribution. The proposed methodology employs response surface analysis to express simulated critical values as functions of sample size, enabling rapid computation of finite-sample critical values without requiring extensive simulations. The response surface regressions are estimated using median regression, ensuring robustness to outliers and heteroskedasticity. Extensive Monte Carlo experiments demonstrate that the estimated asymptotic critical value functions provide highly accurate test sizes across a wide range of sample sizes.
This work introduces a novel version of the Stahel-Donoho multivariate outlier detection procedure, which considers 5p + 1 specific random directions, where p is the dimensionality of the data, that is, the number of variables in the dataset. These directions are derived by maximizing the squared third sample moment of the projected observations, which then serves as a seed to obtain 5p additional directions via a stratified sampling. Compared with the standard Stahel-Donoho estimator and other outlier detection methods, this new version exhibits competitive performance across various high-dimensional datasets and contamination scenarios. By leveraging maximum skewness projection within the Stahel-Donoho framework, the proposed estimator maintains stable results in high dimensions, showing its advantage in efficiently handling complex data structures.
This article introduces a measure-theoretic definition of the likelihood function via Radon-Nikodym (RN) derivatives and establishes a new likelihood proportionality theorem showing that likelihoods obtained from any pair of dominating measures differ only by a parameter-free factor. This result validates the RN-based definition in light of the likelihood principle, particularly in settings where a single canonical dominating measure does not exist or multiple choices are natural-such as certain infinite-dimensional or missing-data problems. The role of continuous versions of RN derivatives is highlighted, demonstrating how continuity both ensures well-behaved, often unique likelihoods and aligns with Fisher's original intuition. The prior predictive measure is also examined as an alternative dominating measure in Bayesian contexts, while exponential families are shown to retain their defining exponential structure under any dominating measure. Collectively, these findings unify and refine fundamental measure-theoretic questions about likelihood, offering a rigorous framework for likelihood-based inference across a variety of statistical models.
In high-dimensional semiparametric regression, balancing accuracy and interpretability often requires combining dimension reduction with variable selection. This study introduces two novel methods for dimension reduction in additive partial linear models: (i) minimum average variance estimation (MAVE) combined with the adaptive least absolute shrinkage and selection operator (MAVE-ALASSO) and (ii) MAVE with smoothly clipped absolute deviation (MAVE-SCAD). These methods leverage the flexibility of MAVE for sufficient dimension reduction while incorporating adaptive penalties to ensure sparse and interpretable models. The performance of both methods is evaluated through simulations using the mean squared error and variable selection criteria, assessing the correct detection of zero coefficients and the false omission of nonzero coefficients. A practical application involving financial data from the Baghdad Soft Drinks Company demonstrates their utility in identifying key predictors of stock market value. The results indicate that MAVE-SCAD performs well in high-dimensional and complex scenarios, whereas MAVE-ALASSO is better suited to small samples, producing more parsimonious models. These results highlight the effectiveness of these two methods in addressing key challenges in semiparametric modeling.
In the last two decades, several classes of distributions have been introduced to provide great flexibility in modeling real data. We define a new Weibull flexible-G family and construct a regression model based on this distribution. Maximum likelihood methods are adopted. Three real applications reveal that the new models can perform better fits than other competing ones.
This article introduces a generalized method for estimating parameters of an alphastable distribution based on the characteristic function. We propose a novel approach that extends existing techniques. As an application, we compare our proposed method with the maximum likelihood method. Our results demonstrate the efficacy of the proposed method in accurately estimating the parameters of alpha-stable distributions. Additionally, we highlight the advantages of the proposed method ver maximum likelihood method in capturing the tails and skewness of the distribution. This research contributes to the advancement of statistical methods for modeling heavy-tailed and skewed distributed data, with applications in finance, risk management, and other fields.
This study comprehensively explores the research landscape within statistical and reliability studies, focusing on the Birnbaum-Saunders distribution, Gaussian inverse distribution, cumulative damage models, and fatigue life prediction. Using a combination of bibliometric analysis, network visualization, thematic mapping, and latent Dirichlet allocation, we analyze 465 articles from the ISI Web of Science database. These articles were selected for their relevance based on a targeted search strategy. Our analysis identifies key trends, collaboration networks, and emerging research themes. Notable growth in scholarly activity was observed from 2015 to 2021, with a peak around 2021, followed by a decline in the number of publications. Relevant contributions were noted from countries such as Brazil, Canada, Chile, China, Iran, Japan, and the United States. The thematic analysis of keywords reveals influential motor themes like the Birnbaum-Saunders distribution and expectation-maximization algorithm; specialized niche areas such as producer risk; emerging or declining themes like the generalized Birnbaum-Saunders distribution; and foundational themes including cumulative damage and fatigue life distributions. A cluster analysis states key focus areas, such as material durability and advanced statistical methods. Integrating latent Dirichlet allocation, six main topics are derived, capturing broad thematic structures. However, some niche areas do not align directly due to their specialized nature and limited cross-field impact. These findings map the current research on this thematic and suggest future research directions, including deeper exploration of niche themes, integration of advanced statistical methods in practical applications, and increased collaboration across diverse research areas to enhance the robustness and applicability of reliability models.