
In recent years, considerable attention has been devoted to statistical models designed for continuous doubly bounded dependent variables, defined on a finite interval (a, b) with known limits satisfying −∞ < a < b < ∞. A predominant focus, however, has been on the special case in which the response is restricted to the standard unit interval, namely (0, 1). In this context, we introduce a new class of quantile regression models based on the unit-gamma distribution, along with a mean-based model. For inference, we adopt a Bayesian framework, employing Markov Chain Monte Carlo (MCMC) stochastic simulation algorithms. We conducted extensive simulations to verify the computational implementation and assess the properties of the Bayesian estimators. All analyses were carried out using the nimble package in R, which provides a flexible environment for building and fitting Bayesian hierarchical models. Our results suggest that the proposed Bayesian approach provides unbiased and consistent estimators for all model parameters. Additionally, we perform a detailed model fit assessment, comparing the proposed models with several established alternatives, and conduct an influence analysis to identify potential outliers or influential data points. Finally, we apply the proposed models to a real-world dataset, demonstrating their practical utility and systematically comparing their performance with that of existing models commonly used in the literature.
The nonlinearity and seasonality of the business cycle account for the majority of short-run movements in quarterly or monthly macroeconomic time series. The multiplicative seasonal threshold autoregressive model with exogenous inputs (TSARX) combines threshold dynamics, multiplicative seasonality, and external regressor, the model is a special case of a general nonmultiplicative TAR model, but there is limited evidence on how it performs in forecasting relative to simpler models and machine-learning algorithms. We evaluate the performance out-of-sample of TSARX models using seasonally unadjusted macroeconomic time series for three economies (Colombia, the United States of America, and the United Kingdom) and three key variables (gross domestic product (GDP), unemployment rate, and inflation). For each country-variable pair, we compare the TSARX models with four competitor models: a TAR model, a linear seasonal autoregressive model (SAR), Holt-Winters exponential smoothing (ES), and long short-term memory (LSTM) neural networks. Multi-step forecasts at horizons from one to four periods ahead are produced under a rolling-window design, and accuracy is assessed using Mean Squared Error (MSE) and Diebold-Mariano tests (DM) for equal predictive ability. Across 36 series-country-horizon combinations, the TSARX achieves the lowest MSE in three cases and is often statistically indistinguishable from the best benchmark; in many situations simpler models remain di‑cult to beat. These findings show that additional nonlinear and seasonal structure does not guaranty superior forecasts and that the benefits of the TSARX models are context- and horizon-dependent.
In the financial context, survival analysis has been used in a variety of ways, like in the study (Borelli & Lucena, 2022), which aims to estimate the time until the recovery of overdue credit portfolios, with a focus on pricing non-performing credit portfolios. As in the study by Ramirez (2016), which aimed to estimate the time until default, evaluating the conditioning variables for credit delay. One of the objectives of this paper is to analyze a real dataset about credit, to model the time until the customer of a financial company, located in Rio Grande do Sul - Brazil, becomes defaulter. Therefore, in this paper it is proposed a model for survival data with cure rate, in which the distribution of the times is adjusted by the power piecewise exponential (PPE) distribution. The other objective is to propose a model that has not yet been explored jointly in the literature, through the construction of a long-term survival model, considering the approach of the promotion time models (Yakovlev & Tsodikov, 1996) and the power piecewise exponential distribution (Gómez et al., 2017). To evaluate the performance of the proposed model, a simulation study was carried out comparing the proposed model with models available in the literature (Weibull and piecewise exponential), using the R software. Finally, an application of the proposed model was conducted on a real data set related to personal loans, to evaluate the applicability of the model in a credit context.
This paper introduces a new time series model for non-negative continuous data based on the Maxwell distribution. The proposed model employs a reparameterization of the Maxwell distribution in which its parameter directly represents the mean. In this formulation, the conditional mean is modeled through a dynamic structure that combines autoregressive and moving average components, linked by an appropriate link function. Parameter estimation is carried out using the conditional maximum likelihood method, for which closed-form matrix expressions of the conditional score vector and the conditional Fisher information matrix are derived. Based on the asymptotic properties of the estimators, procedures for interval estimation and hypothesis testing are presented. Monte Carlo simulations assess the finite-sample performance, and provide evidence of the estimators' convergence toward the true parameter values as the sample size increases. An empirical application involving wind speed data from Brasília, the capital of Brazil, shows the practical relevance and effectiveness of the proposed model for real-world time series modeling and forecasting.
There are many linear and nonlinear measures of association between two continuous pairwise variables. They are used to indicate the strength of the relationship between the two variables. The question thus arises as to which of these measures should be used to explore relationships between two variables in general. The identification of linear and/or nonlinear relationship between two variables can help to avoid problems within a regression framework. The objective of this paper is to examine alternative measures of association that could be employed as a replacement or in conjunction with, standard linear correlation coefficients. The results lead us to conclude that the maximum correlation measure is particularly useful, and capable of detecting linear and nonlinear associations between two continuous variables, while also being relatively computationally efficient. It can be utilized in exploratory analysis and in a modern regression framework.
In this article, we introduce a new continuous probability distribution called the Inverse Exponential Logistic Lehmann Type II distribution, derived from the Lehmann Type II alternative. The main objective is to apply this new distribution to survival analysis, specifically with right-censored data. We discuss various properties of the proposed distribution, including quantiles, skewness, kurtosis, moments, order statistics, and Rényi entropy. The distribution exhibits a hazard rate function with different shapes depending on the parameter values. Simulation studies were conducted to evaluate the performance of maximum likelihood estimates under a right censoring scheme. Finally, we illustrate the usefulness and flexibility of the proposed distribution by applying it to two real datasets and comparing its performance with that of other distributions.
Quantile regression provides a parsimonious model for the conditional quantile function of the response variable Y given the vector of covariates X, and describes the whole conditional distribution of the response, yielding estimators that are more robust to the presence of outliers. Quantile regression models specify, for each quantile level τ , the functional form for the conditional τ -th quantile of the response, which brings complexity to perform variable selection using regularization techniques, such as LASSO or adaptive LASSO (adaLASSO), as one might obtain a different set of selected variables for each quantile level. In this work, we propose a method for global variable selection and coefficient estimation in the linear quantile regression framework, imposing few restrictions on the functional form of β(·), and applying group adaLASSO penalization for variable selection. We set up a Monte Carlo study comparing six different proposed estimators based on LASSO, adaLASSO and group LASSO in six scenarios that diversify sample and quantile levels grid sizes. The findings demonstrate that the selection of the tuning parameter λ for penalization is critical for model selection and coefficient estimation. It was observed that the methods using traditional LASSO are more prone to include the true model as compared to adaLASSO, but renouncing model shrinkage and not removing irrelevant covariates, while the grouped approaches are more effective in zeroing coefficients that are less relevant.
The exponentiated transmuted-G family was introduced along with some of its statistical properties derived from a power series expansion. This work demonstrates that the previously proposed expansions, which depend on a double sum, present convergence problems for some combinations of parameters. Therefore, a much simpler linear representation that depends only on a single sum is presented, allowing for the precise and general calculation of these properties.
Human capital theory posits a central hypothesis: there is a direct relationship between individuals' levels of education and their productivity, which in turn leads to higher earnings. To evaluate this hypothesis, the economic literature commonly estimates the Mincer equation using the classical linear regression model. In this paper, both the traditional approach and a GAMLSS model with the Dagum distribution are estimated for wage earners in private firms and the public sector in Colombia in 2024. The hypothesis that an individual's rate of return increases with years of education, and that it rises with years of experience up to a certain point in the life cycle before declining, is confirmed by both the linear regression and the GAMLSS model. Model selection is based on the Generalized Akaike Information Criterion (GAIC), and the results show that the GAMLSS provides a better fit than the linear regression model.
This work presents the development of a multilevel Bayesian nonparametric model that allows for the estimation of linear relationships in heterogeneous data sets, while simultaneously identifying clusters without the need to specify the number of groups in advance. The study includes the mathematical development of the model using the Chinese Restaurant Process and the implementation of algorithms for its fitting. The results obtained from real data show that the model performs well in both clustering data and characterizing linear relationships, achieving results comparable and even better to those obtained by traditional parametric methods.
Recently, it has been proposed to estimate the conditional mode of a response, given a vector of covariates, using a computationally scalable estimator derived from the linear quantile regression model. Alternatively, we propose to estimate the conditional mode by maximizing a smoothed conditional density estimator. This approach offers at least two benefits: computational efficiency and good asymptotic behavior which, in particular, bypasses the curse of dimensionality.
In sample survey, missing data is a common issue. Various imputation techniques have been developed to handle the missing data issue. But, a miniscule work has been done to handle missing data issue in the presence of measurement errors (ME) and correlated measurement errors (CME). This manuscript proposes a few logarithmic imputation techniques and the accompanying point estimators to address the missing data issue when the data are affected by CME. The mean square error (MSE) of the proposed imputation methods is reported to the first order approximation. The dominance conditions of the proposed imputation methods over the coeval imputation methods are obtained. Afterward, a simulation study using an artificially drawn population and a real data application are carried out to support the theoretical findings.
TheWelch-Satterthwaite (WS) methodology is typically used in medicine, biology and economic courses to make inferences about the difference between two population means. Despite his wide-spreading applications, it has been pointing out in many references the multiple limitations of the inferences based on it. In this work, we propose three simple ways to improve the classical WS approach. Under balanced samples scenarios, we give exact inference results of two of the proposed estimators. Additionally, under unbalanced samples scenarios, we offer first-order approximation results and through several Monte Carlo simulations, we assess the mean and variance of the proposed estimators under (very) small and moderate sample sizes. Nonetheless, the simplicity of the proposed approach we obtain a much better performance than the WS proposal. Lastly, one application is presented in which the proposed estimators potentially improve the performance of t-student interval estimation and hypothesis testing procedures.
This article introduces a novel bivariate probability distribution derived through a transformation-based approach, along with the closed-form expression of its l-th order joint moment. Although the distribution may be employed as a prior for the shape parameters of the Beta distribution, the main focus of this work lies in evaluating the convergence behavior of Markov Chain Monte Carlo (MCMC) algorithms designed to generate samples from this new distribution. A simulation strategy is analyzed, consisting of a Gibbs sampling scheme in which an adaptive random walk Metropolis-Hastings (ARWMH) algorithm is used to sample from one of the full conditional distributions, employing a four-parameter Beta distribution as the proposal. Convergence is assessed using diagnostics such as the effective sample size (ESS) and the potential scale reduction factor (R-hat). The results show that when the elements of the parameter vector ϕ differ, the empirical moments obtained from the chains approximate the theoretical values accurately. However, when all components of ϕ are equal, the estimates of variance and covariance deviate considerably, revealing a sensitivity to symmetry in the geometry of the new distribution. A brief application in a Bayesian context is also presented, in which the new distribution is used as a prior and the Beta distribution as the likelihood. This application confirms that the proposed sampling methods yield empirical moments that are consistent with the theoretical ones, thus supporting the validity of the strategy. The contributions of this study are relevant to the design, evaluation, and implementation of MCMC techniques for sampling from complex distributions in Bayesian inference.
The discriminatory capacity of a test is commonly expressed in terms of sensitivity and specificity, and there is generally a compromise relationship between these two measures, as an increasing threshold for defining the positivity of the test results in a decrease in sensitivity and an increase in specificity. Recommended methods for the meta-analysis of diagnostic tests, such as the bivariate model, focus on estimating a summary sensitivity and specificity at a common threshold, while the Hierarchical Summary Receiver Operating Characteristic (HSROC) model focuses on estimating a summary curve from studies that have used different thresholds. Therefore, we will explain the hierarchical modeling for meta-analysis of the study of precision in diagnostic tests, and we will design a decision scheme that helps to understand the models to choose the most appropriate in situations of heterogeneity in the studies, and illustrate its application in a systematic review, studying the properties and assumptions of meta-analytic procedures to synthesize the quantitative evidence of the parameters. For which, we used a systematic review that obtained summary estimates for the diagnosis of invasive aspergillosis, being our modeling framework the NLMIXED procedure of SAS, obtaining summary estimates for sensitivity and specificity for the bivariate model of 0.7708 and 0.8521, from a total of 27 studies involving 3,943 patients, and for the HSROC case the values were 0.7304 and 0.8867 respectively. Finally, we hope that this article will provide clinicians with a sufficient understanding of the terminology and statistical methods, obtaining plausible interpretations of the results in systematic reviews.
During human survey, participants are asked highly personal questions about a sensitive variable. This paper focus on estimation of population mean of sensitive study variable in the existence of non-response and measurement error under Optional Randomized Response Technique model using two-phase sampling. The bias and mean squared error of the proposed and considered family of estimators have been derived up to first order of approximation. Further, the properties of the proposed estimator have been discussed, and efficiency conditions have been derived. To demonstrate the theoretical findings, simulation study is conducted under various conditions using data set from hypothetical population and it is clear that our proposed family of estimator is always better than the other considered family of estimators.
In this paper, Expected Bayesian and Hierarchical Bayesian techniques have been discussed to estimate the shape parameter of the inverse power Lomax distribution. The proposed estimates for the shape parameter are obtained by using an informative gamma prior based on squared error, entropy and weighted balance loss functions. The definitions of the proposed estimators as well as their characteristics are provided. A Monte Carlo simulation is executed to compare the performance of the proposed estimators in terms of mean squared error. Finally a real life data set has been analyzed for further illustrations.
The study investigated the dynamics of _commencement-to-event-time-behaviour in life insurance portfolios, employing Maximum Likelihood Estimation (MLE) and Maximum A Posteriori (MAP) with the Markov Chain Monte Carlo (MCMC) simulation technique. Focusing on the Lognormal and Exponential distributions for their efficacy in modelling time-to-occurrence data, the research simulated 120 observations from both distributions and estimated parameters using the first 80 ordered samples. Remarkably, estimates for lognormal parameters obtained through MLE and MAP_MCMC were highly similar, with errors well within 10% of the actual values, highlighting the accuracy of both methods. The study also explored the robustness of the MAP_MCMC technique to various prior distributions, demonstrating its effectiveness across different priors, including Exponential, Normal, Gamma, Pareto, and Weibull prior distributions. In the case of the exponential distribution, both MLE and MAP_MCMC techniques performed exceptionally well, providing estimates within 5% of the true value, with MAP _MCMC exhibiting remarkable precision, just 1% off the true value. Real-life data fitted to the Gamma distribution showed that MLE and MAP _MCMC methods, using censored data, closely approximated benchmark estimates from the method of moments. The MAP_MCMC approach slightly outperformed the MLE.
In this paper, some estimators of the unknown parameter, reliability function and hazard rate function of Inverse Pareto distribution under Unified Hybrid Censoring were derived. The maximum likelihood method, Bayes and E-Bayes method were used for estimating the parameter, reliability function and hazard rate function of the Inverse Pareto Distribution. Approximate confidence intervals (confidence interval and credible interval) were also derived. Comparisons were made in sense of mean squared error and asymptotic relative efficiency through Monte Carlo simulation. Finally, the proposed methods can be understood through illustrating the results of the real data analysis.
In this paper, we examine the effectiveness of high-dimensional methods for real-time forecasting of inflation in Colombia. We utilize statistical dimension reduction techniques, such as sparse principal components and dynamic factor analysis, alongside machine learning algorithms that incorporate shrinkage methods. Our evaluation of out-of-sample forecasts, using a dataset of 102 macroeconomic and financial indicators, indicates that ensembles of multiple underlying models can enhance forecast accuracy for horizons of 11 and 12 months ahead. Additionally, stepwise models are suitable for horizons between 4 and 10 months ahead, and spectral component models are effective for short horizons.