
In this paper, we introduce the Beta-pG family of distributions as a flexible framework for modeling long-term survival data. This family extends the traditional Beta-G distribution by incorporating a new parameter p, from which a nonzero long-term survivor fraction is induced through the limiting survival probability. This structure allows the model to capture survival patterns in which a portion of the population remains unaffected by the event of interest over extended follow-up. In particular, we focus on the case where G, the baseline distribution, is specified as the exponential distribution, leading to the Beta-pExponential model. Parameter estimation is performed using maximum likelihood method, allowing likelihood-based inference for the model parameters and the long-term survivor proportions. Monte Carlo simulation studies are conducted to evaluate the finite-sample performance of the estimators across various scenarios. The proposed model is then applied to two melanoma survival datasets, including data from a clinical trial and a population-based cancer registry. In both applications, the fitted model captures the main features of the observed survival curves and provides direct estimates of the long-term survivor proportions. These results illustrate the usefulness of the Beta-pG family as an alternative framework for analyzing survival data with long-term survivors.
Granger causality models are widely used in time series analysis to manage causal relationships between variables based on temporal precedence. Traditional Granger causality models always assume linear relationships and normality assumptions, which often fail to capture the complexities of real-world data involving nonlinear and non-normal behavior. In this work, we propose a copula-based approach that offers a novel alternative method by capturing nonlinearity and non-normality in multivariate time series modeling. We develop a copula-based joint regression model tailored for Granger causality analysis and carry out some predictions in these models. We assess and validate the performance of our method using simulation studies in different scenarios. Two applications of our proposal in real-world data analysis are also presented. Our findings highlight the superiority of copula-based approaches compared to other Granger causality models in complex time series data.
The manuscript introduces a new first-order integer-valued autoregressive process in a random environment, based on the binomial thinning operator. The proposed model is driven by two control processes: the first determines the marginal distribution, while the second governs the correlation structure and the dynamic behavior of the series. The inclusion of two control processes enhances the model’s flexibility and allows for a more accurate fit to the data. Two variants of the model are considered—one with a Poisson marginal distribution and the other with a geometric marginal distribution—and both specifications are theoretically analyzed. The fundamental properties of the model are investigated in detail, providing comprehensive theoretical insight into its structure, moment characteristics, and the dependence structure it generates. For the estimation of unknown parameters, two methods are employed: the Yule–Walker estimator and the conditional maximum likelihood estimator. Their performance is evaluated through extensive simulation studies. Finally, the practical applicability of the proposed model is illustrated using real-world data.
The Bayes factor reversal (BFR) paradox shows that for any significant result under a point-null hypothesis, there exists a prior scale τ∗—the “flip point”—where BF01=1. Below τ∗, evidence favors H1; above, it favors H0. We prove this phenomenon is universal across thirteen standard test families: z-tests, t-tests (one-sample, two-sample, paired), one-way ANOVA, chi-squared, proportion tests, correlation, regression, and nonparametric tests (Mann–Whitney and Wilcoxon signed-rank). The mathematical mechanism is the same in every case: continuity, a unique minimum, and divergence of BF01(τ) guarantee a unique flip point. We propose the flip point as a robustness diagnostic: just as power analysis quantifies frequentist fragility to sample size, the flip point quantifies Bayesian fragility to prior specification. When τ∗ lies within the range of common default priors, conclusions are sensitive to arbitrary software choices. We provide closed-form solutions where available, R code, and an interactive web-based Flip Point Calculator covering all test families, enabling researchers to assess and report the robustness of their Bayesian inferences.
This study proposes a two-dimensional index for measuring the degree of deviance from the quasi-symmetry (QS) model corresponding to the conditional QS (CQS) model for ordinal square contingency tables. The CQS model indicates the structure of asymmetry for the local odds ratios. Previous study proposed a two-dimensional index for measuring the degree of deviance from the QS model corresponding to the extended QS (EQS) which has a different probability structure from the CQS model. The asymmetry parameters of the CQS and EQS models indicate both the degree and direction of deviance from the QS model. When the CQS (or EQS) model does not fit for the presented data well, we cannot evaluate them using the asymmetry parameter. The two-dimensional index can address this issue. The proposed two-dimensional index is assembled by combining two new sub-indexes. One represents the degree of deviance from the QS model corresponding to the CQS model, and the other represents its direction. Additionally, this study derives a plug-in estimator and an approximate confidence region for the proposed two-dimensional index. For comparing several datasets, we reveal that the magnitude relationship between the degrees of deviance from the QS model changes by whether the proposed or existing two-dimensional index is used.
Let (Zn) be a supercritical branching process in an independent and identically distributed (i.i.d.) random environment . For the mean estimator Mn = n-1 & sum;n-1 k=0(Zk+1/Zk) introduced by (The Annals of Statistics 7 (1979) 680-685), we establish a Berry-Esseen bound and an algebraic nonuniform Berry-Esseen bound, and some applications of the nonuniform Berry-Esseen bound to confidence interval estimation are discussed.
In the usual statistical inference problem, we estimate an unknown parameter of a statistical model using the information in the random sample. A priori information about the parameter is also known in several real-life situations. One such information is the order restriction between the parameters. This prior information improves the estimation quality. In this paper, we deal with the componentwise estimation of location parameters of two exponential distributions with ordered scale parameters under a bowl-shaped affine invariant loss function and generalized Pitman closeness criterion. We have shown that several benchmark estimators, such as maximum likelihood estimators (MLE), uniformly minimum variance unbiased estimators (UMVUE) and the best affine equivariant estimators (BAEE) are inadmissible. We have given sufficient conditions under which the dominating estimators are derived. Under the generalized Pitman closeness criterion, a Stein-type improved estimator is proposed. As an application, we have considered special sampling schemes such as type-II censoring, progressive type-II censoring and record values. We conducted a simulation study to compare the risk performance of the improved estimators. Finally, we performed a real-life data analysis to demonstrate the practical applications of our findings.
The study of animal telemetry is crucial in ecology, providing valuable information on movement patterns, behavior, and habitat use of various species, which is essential for conservation and management efforts. Numerous models in the literature address animal telemetry by modeling velocity, telemetry data itself, or both processes jointly through a Markovian approach. In this work, we propose a novel approach by modeling the velocity of each coordinate axis for animal telemetry data using a fractional Ornstein-Uhlenbeck (fOU) process. The integral of the fOU process models the position data in animal telemetry. This proposed model is particularly flexible in capturing long-range memory effects. The Hurst parameter H is an element of (0, 1) plays a crucial role in the integral fOU process, determining the long-range memory. The integral fOU process is nonstationary; a higher Hurst parameter (H > 0.5) indicates stronger memory, leading to trajectories with transient tendencies, while a lower Hurst parameter (H < 0.5) implies weaker memory, resulting in trajectories with recurring trends. When H = 0.5, the process reduces to a standard integral Ornstein-Uhlenbeck process. We develop a simulation algorithm for telemetry trajectories using finite-dimensional distributions and employ the maximum likelihood method for parameter estimation, with its performance evaluated through simulation studies. Finally, we present a telemetry application involving Fin Whales dispersing throughout the Gulf of California.
In simple linear regression with response variable Y and covariate X, the classical Pearson correlation measures the strength of the linear association between Y and X, and corresponds to the standardized slope of the regression line. This paper explores the concept of local linear correlation to capture the locally linear association between Y and X as a function of X, while preserving key properties of the Pearson correlation. Without assuming a parametric form for the joint distribution of (X, Y), we show that the kernel-weighted local linear correlation measures the strength of locally linear association, and is connected to local linear regression through its interpretation as a locally standardized slope. We derive the finite-sample and asymptotic properties of the population and sample versions of local linear correlations and the optimal order of the bandwidth is provided. Numerical results confirm the asymptotic theory and a baseball data example is given for illustration.
This article is concerned with the optimal design problem of efficient statistical inference for comparing multivariate linear models estimated from samples of independent measurements. The objective is to find the & micro;D-optimal designs that minimize the volume of the confidence tube for the multivariate linear models. The definitions of the volume for the confidence tube are obtained. General equivalence theorems are established to verify the & micro;D-optimality in the set of all approximate designs. Two examples are presented to illustrate the applications of the obtained results.
In this paper, we propose a new bivariate integer-valued autoregressive process with interaction effects, referred to as the IEBINAR(1) model. We consider three estimation methods for the unknown parameters of interest: conditional least squares, conditional maximum likelihood, and a two-step estimation approach. The asymptotic properties of the estimators are established. The performance of these estimation methods is compared through simulation experiments. Furthermore, we conduct hypothesis testing to examine the existence of interaction effects in the IEBINAR(1) model. Finally, a real data application is presented to evaluate the performance of the proposed model.
In this work, we developed a new Bayesian method for variable selection in function-on-scalar regression (FOSR). Our method uses a hierarchical Bayesian structure and latent variables to enable an adaptive covariate selection process for FOSR. Extensive simulation studies show the proposed method's main properties, such as its accuracy in estimating the coefficients and high capacity to select variables correctly. Furthermore, we conducted a substantial comparative analysis with the main competing methods, the BGLSS (Bayesian Group Lasso with Spike and Slab prior) method, the group LASSO (Least Absolute Shrinkage and Selection Operator), the group MCP (Minimax Concave Penalty), and the group SCAD (Smoothly Clipped Absolute Deviation). Our results demonstrate that the proposed methodology is superior in correctly selecting covariates compared with the existing competing methods while maintaining a satisfactory level of goodness of fit. In contrast, the competing methods could not balance selection accuracy with goodness of fit. We also considered a COVID-19 dataset and some socioeconomic data from Brazil as an application and obtained satisfactory results. In short, the proposed Bayesian variable selection model is highly competitive, showing significant predictive and selective quality.
This paper introduces a ridge penalization scheme to enhance the numerical stability of conditional maximum likelihood estimation of the parameters indexing the beta ARMA model. The proposed approach involves adding a simple penalty term to the conditional log-likelihood function to enhance its curvature. This modification reduces the chance of convergence failures and implausible estimates. We also present a bootstrap-based parameter estimation strategy. It is particularly useful when penalization alone is insufficient to address numerical issues, providing a complementary solution for obtaining more reliable estimates. Our numerical results show the effectiveness of the proposed approaches in addressing numerical instability issues in beta ARMA parameter estimation. Two empirical applications are presented and discussed.
In this article, we propose the zero-adjusted defective Gompertz model incorporating gamma frailty, aimed at jointly modeling survival data where excess zeros and cure fractions coexist, a frequent challenge in biomedical and public health studies. Traditional survival models often fail to capture these complexities simultaneously, limiting their applicability in real-world medical data. The proposed approach integrates the flexibility of the Gompertz distribution with structural adjustments for zero inflation and defective survival functions, while the inclusion of a gamma frailty term accounts for unobserved heterogeneity at the individual level. Through extensive Monte Carlo simulations and bootstrap analyses, we demonstrate the model's consistency and reliability of parameter estimates. Applications to real medical data, including insulin use among pregnant women with gestational diabetes, illustrate the model's practical utility in accurately identifying both cured and zero-adjusted subpopulations, offering a robust and versatile framework for analyzing complex survival patterns in health research.
The Unit Gompertz (UGo) distribution has two parameters and is adequate for modeling data with support in the interval (0, 1). It was introduced as an alternative to the beta and Kumaraswamy distribution for modeling double-bounded variables. The UGo distribution can accommodate asymmetric data and has been applied in various situations, such as environmental studies, industrial applications, and survival analysis. An attractive characteristic of the UGo is its closed-form expression for the quantile function. It allows for the formulation of a quantile-based parameterization and accommodates different dependence structures for modeling the conditional quantiles. Therefore, in this study, we introduce a simple alternative for modeling double-bounded variables under serial correlation in a conditional quantile of the UGo distribution. The so-called UGo-ARMA is constructed considering an autoregressive moving average structure using the UGo distribution as the random component. The maximum likelihood method was used for parameter estimation. Subsequently, Monte Carlo simulations are conducted to investigate the performance of the maximum likelihood estimators and the asymptotic confidence intervals of the parameters. The proposed model is an alternative for modeling double-bounded variables with serial correlation, especially in contexts where UGo has already proven competitive with classic unitary distributions. To illustrate the practical relevance of the proposed model, we apply it to a financial time series: the average monthly interest rate for credit card installment operations in Brazil. The results highlight the model's ability to capture serial dependence and distributional features typically found in real-world bounded data.
Modeling count time series is essential in many fields, especially when data exhibit complex features such as overdispersion and zero modification. While the first-order integer-valued autoregressive (INAR(1)) model is a fundamental tool, it often fails under these conditions. This study explores and systematically compares two extensions: zero-inflated INAR(1) and hurdle INAR(1), utilizing both Poisson and negative binomial innovations. We propose a unified Bayesian framework using Hamiltonian Monte Carlo in Stan to assess performance under zero-inflated and zero-deflated scenarios. Simulation results reveal a critical trade-off: in zero-inflated settings, zero-inflated models generally offer superior predictive fit, whereas hurdle models provide more accurate recovery of structural zero parameters. In zero-deflated settings, however, zero-inflated models fail structurally, making hurdle models the only viable alternative. These theoretical findings are corroborated by applications to two contrasting datasets from the same urban census tract: drug-related offenses (zero-inflated) and sex offenses (zero-deflated). To support reproducibility and broader adoption, we provide an open-source R package, ZIHINAR1, available on CRAN and GitHub, for model fitting and comparison. These findings offer practical guidance for selecting models that accommodate complex zero structures in discrete-valued time series.
In this paper, we propose a family of multivariate asymmetric distributions over an arbitrary subset of set of real numbers which is defined in terms of the well-known elliptically symmetric distributions. We explore essential properties, including the characterization of the density function for various distribution types, as well as other key aspects such as quantiles, stochastic representation, conditional and marginal distributions, moments and parameter estimation. A Monte Carlo simulation study is performed for examining the performance of the developed parameter estimation method. Finally, the proposed models are used to analyze socioeconomic data.
Our study addresses the estimation of the locations of discontinuities (or jumps) within multivariate signals from noisy observations in the nonparametric regression setting. Departing from standard analytical approaches, we propose a new framework, based on geometric control over the jump locations. This allows us to consider larger classes of signals, of any dimension, with potentially wild discontinuity set (exhibiting, typically, self-intersections and corners). We study a simple estimation procedure relying on histogram differences and show its consistency and near-optimality for the Hausdorff distance over these new classes. Furthermore, exploiting this new geometric framework, we design procedures to estimate consistently several topological descriptors of the jump locations.
Missing data is a common issue across fields such as engineering, finance, healthcare and social sciences. If not addressed appropriately, it can lead to biased analyses and reduced statistical power. This paper introduces a family of novel Exponential-Type Imputation Models (NEtIMs) designed under three distinct strategies to handle missing data effectively. These models exploit the properties of newly formulated mean estimators by incorporating measures such as absolute relative bias and mean squared error to better capture uncertainty. The performance of NEtIMs is evaluated through extensive simulations on both symmetric and asymmetric datasets, as well as real-world data. Results show that NEtIMs consistently outperform conventional imputation methods in terms of accuracy, efficiency and robustness across different missing data mechanisms. Additionally, optimality constraints are established to demonstrate the broader applicability and reliability of the proposed mean estimators formulated through NEtIMs.