
In an era of increasingly complex data, three-way arrays capturing information across units, variables and occasions are ubiquitous in fields from chemometrics to finance. However, extracting meaningful and interpretable patterns from such data remain a significant challenge. To address this, we introduce the Explainable Tucker3 Clustering (XT3Clus) methodology. XT3Clus performs clustering on units while simultaneously identifying explainable components for variables and/or occasions, significantly enhancing model interpretability. This approach functions as a constrained Tucker3 model, where each dimension is forced to contribute to a single component. The framework supports fully confirmatory, exploratory or hybrid analytical strategies. The optimization of the objective function is carried out by an efficient Alternating Least Squares algorithm. Finally, we propose a novel quantitative metric to evaluate the interpretability of a solution and confirm the practical utility of XT3Clus in three real-world scenarios.
The redundancy allocation problem (RAP) is considered as one of the important problems in reliability theory. In this paper, we consider a series system with several subsystems wherein each subsystem is a weighted k-out-of-n system formed by different types of multi-state components. The degradation of the performance level (i.e., the probability of changing from a given state i to the next state (i - 1)) of a component of the system is modeled by the Markov process. Then, we study the multi-objective RAP problem for this system, that is, we determine the optimum number of components of each type in each subsystem so that the maximum system reliability is achieved at minimum cost. Note that the given RAP problem is of NP-hard type, and consequently, we use the controlled elitism non-dominated ranked genetic algorithm (CE-NRGA) to solve this problem. At the end, we illustrate the proposed methodology through a numerical example. Moreover, we discuss a case study to validate the proposed model.
In this paper, we study the properties and performance of optimal transport autoregression in modeling and forecasting high-frequency financial data distributions. We build on a class of univariate autoregressive transport models recently proposed in the literature (Zhu and M & uuml;ller) where the distributional time series dynamics is modeled either through a single scalar, similarly with traditional Euclidean autoregressive models, or via a functional distribution-contraction coefficient. Properties and performance of the models are investigated through an empirical application to forecast distributions of high-frequency financial price returns and volatility of Bitcoin. Our results show that forecast errors are highly time- and quantile-dependent: while autoregressive transport models are generally able to predict return and volatility densities during "normal business" periods, forecast errors tend to rise in the proximity of extreme quantiles, though such increase is non-monotonic. We highlight the strengths and weaknesses of the method in modeling the distributional time series of high-frequency, noisy financial data, suggesting some potential directions for future research.
This article develops a valuation model for mortgages with partial prepayment risk under the equal principal payment method. To model partial prepayment, we assume prepayment events occur at the jump times of a point process. The proportion of the prepayment amount to the outstanding principal balance is characterized by a stochastic process. By defining specific forms for the point process and stochastic process, we derive explicit valuation formulas for the mortgage loan. Finally, after resolving the computational challenge of multiple integrals via matrix exponentiation techniques, numerical results are presented to investigate the impact of some parameters on the valuation.
The evaluation of healthcare services is a crucial aspect of public health management, as it provides insights into service effectiveness, efficiency, and user satisfaction. This paper proposes a multidimensional approach to measuring healthcare service experience by constructing a synthetic index that incorporates three key perceptual dimensions: Cost, accessibility, and quality. These latent constructs are measured using a set of elementary indicators from the European Quality of Life Survey. To develop the dimensional synthetic indices—one for each experience dimension—as well as the overall experience index and the customer satisfaction index, we employ Partial Least Squares Path Modeling (PLS-PM). This approach not only enables the synthesis of latent variables but also allows for the analysis of structural relationships between them. The results provide a comprehensive framework for assessing healthcare service experiences and offer valuable insights for policymakers and service providers aiming to enhance healthcare quality and accessibility while improving user satisfaction.
The advent of ensemble methods, such as Random Forest (RF), has led to a paradigm shift in supervised learning. These methods have achieved remarkable levels of prediction accuracy by aggregating multiple weak learners. However, a drawback of these methods is their lack of transparency, which often prevents users from understanding their prediction processes. In light of these challenges, Explainable Ensemble Trees (E2Tree) has recently been proposed, providing a graphical representation of the relationships between response variables and predictors in RFs for classification. E2Tree constructs a single decision tree based on (dis)similarities between observations. By summarizing both distances in terms of predictors and a forest as a single decision tree, E2Tree merges the strengths of both decision trees and decision tree ensembles. In this paper, we propose to extend the E2Tree methodology to regression contexts. We investigate the performance of E2Tree for regression using real-world datasets. We use the Mantel test to test the correlation between similarities of the RF and E2Tree.
The gamma process is a natural model for monotonic degradation processes. In practice, it is desirable to extend the single gamma process to incorporate measurement error and to construct models for the degradation of several nominally identical units. In this paper, we show how these extensions are easily facilitated through the Bayesian hierarchical modeling framework. Following the precepts of the Bayesian statistical workflow, we show the principled construction of a noisy gamma process model. We also reparameterise the gamma process to simplify the specification of priors and make it obvious how the single gamma process model can be extended to include unit-to-unit variability or covariates. We first fit the noisy gamma process model to a single simulated degradation trace. In doing so, we find an identifiability problem between the volatility of the gamma process and the measurement error when there are only a few noisy degradation observations. However, this lack of identifiability can be resolved by including extra information in the analysis through a stronger prior or extra data that informs one of the non-identifiable parameters, or by borrowing information from multiple units. We then explore extensions of the model to account for unit-to-unit variability and demonstrate them using a crack-propagation data set with added measurement error. Lastly, we perform model selection in a fully Bayesian framework by using cross-validation to approximate the expected log probability density of a new observation. We also show how failure time distributions with uncertainty intervals can be calculated for new units or units that are currently under test but have yet to fail.
Theoretical developments in sequential Bayesian analysis of multivariate dynamic models underlie new methodology for counterfactual prediction. This extends the utility of existing models with computationally efficient methodology, enabling routine exploration of post-intervention analyses with multiple time series in putatively casual studies. Methodological contributions also define the concept of outcome adaptive modelling to monitor and respond to changes in experimental time series following interventions. The benefits of sequential analyses with time-varying parameter models for such investigations are inherited in this broader setting. A case study in forecasting retail revenue following marketing interventions highlights the methodological advances.
In health economics, decision-makers rely on models to assess the cost-effectiveness of healthcare interventions and guide resource allocation. Health Technology Assessment (HTA) agencies employ cost-effectiveness models to determine the approval and market access of new therapies within their respective jurisdictions. Health economists use quantitative techniques to synthesize clinical, epidemiological, and economic data to model the costs and effectiveness of a new drug compared to the current standard of care over the lifetime of the patients. These models frequently integrate a wide range of assumptions and data inputs from various sources, which renders them vulnerable to a significant level of uncertainty. Economic models commonly confront multiple forms of uncertainty, such as stochastic uncertainty (first-order), which differs from parameter uncertainty (second-order), as well as the presence of heterogeneity within patient populations. Additionally, structural uncertainty related to the model itself adds another layer of complexity. Uncertainty assessment is essential in model-based health economic evaluations that inform regulatory and reimbursement decisions. Understanding these sources of uncertainty, taking steps to minimize their impact, and analyzing, quantifying, and reporting these inherent uncertainties are crucial for ensuring that health economic models provide robust and reliable insights for effective decision-making. This article examines different types of uncertainty in health economic models and methods to analyze and quantify them, offering practical guidelines with examples from recent literature.
This article presents the benefits of using Bayesian algorithms to fit regime-switching models to daily financial returns data in order to design trading strategies. Our study focuses on a Gaussian hidden Markov model (HMM). We show how the application of a simple smoothing technique preserves the hidden Markov structure and facilitates regime detection even in instances of highly volatile data. The effectiveness of a trading strategy, based on regime detection, may be hindered by a high rate of false signals, leading to numerous trades and, consequently, an escalation in transaction costs. By reducing variance through data smoothing, we enhance the persistence of regimes over time. We validate our statistical learning procedures using synthetic data prior to their application to real-world financial data.
This paper focuses on the weighted mean residual life (WMRL) in mixtures of time to failure distributions. WMRL is an aging index that accounts for transformations of nonnegative random variables. The time to failure of systems operating under changing environments is described by mixtures of distributions that capture the corresponding random effects. This study analyzes the preservation by mixtures of aging properties based on the WMRL and bending properties. The latter compare the WMRL of the mixture and the expected value of the WMRL of the distributions therein. We also analyze the combined effect of a frailty and lifetime functions in the case of mixtures following the proportional WMRL model. The results reveal the improved behavior in the WMRL of mixtures with respect to that in the sub-populations in the mixture. This pattern is relevant for the maintenance of systems.
In this paper, we introduce a new regression method tailored for data presented as distributions. Building on the latest advancements in Distributional Data Analysis (DDA), we propose a new regression model based on a transformation of quantile functions using Logarithmic Derivative Quantile (LDQ) functions. For each distributional variable (where ), we model the LDQ functions as functional data by applying smoothing B-splines at the points corresponding to the distributions' quantiles. The main contribution is the development of a regression model that considers functional regression coefficients. This allows for the consideration of distribution characteristics such as position, variability, and shape. Another contribution is the development of a robust procedure based on trimming distributions to reduce the instability of the tails and make more effective predictions. The proposed approach is corroborated by real environmental data. Cross-validation and bootstrap techniques have been employed to assess the effectiveness of both the new regression model and its robust variant.
In this article, an informative Bayesian approach is proposed for the bounded transformed gamma process, a novel stochastic process recently proposed in the literature to describe bounded above, monotonic increasing, degradation phenomena. The proposed approach is used to analyze a set of real wear data of the cylinder liners of a Diesel engine. Several scenarios, which differ in terms of the quality of the available prior knowledge, are considered and suitable prior distributions are suggested for each of them. In addition, detailed instructions are provided to help potential users incorporate into the suggested prior distributions all and solely the pieces of prior information that are available and sound. In particular, weak prior distributions are also suggested for situations in which available information is poor and/or there is no prior information to exploit. The proposed approach is used to estimate the process parameters and some functions thereof, such as the mean degradation level, the residual reliability of a unit, and to predict the future degradation growth and the useful lifetime. Point estimation and prediction under the (asymmetric) general entropy loss function are also performed to properly deal with situations where overestimation is costlier than underestimation, or vice versa. Estimates and predictions are computed by using proper Markov Chain Monte Carlo algorithms. Results obtained by analyzing wear data of the liners are compared both with those provided by classical methods and with those obtained by using Bayesian approaches based on vague priors. Finally, a sensitivity analysis is developed to study the impact of different prior distributions on the estimates of the parameters.
A new sparse recursive filtering is suggested for the efficient inference of the joint posterior of the number of change points and their locations. The computational and storage costs of the sparse recursive filtering are quadratic to the number of uniformisation times generated by the uniformisation scheme, which can be scaled down to the number of change points. This new version of sparse recursive filtering is generally applicable for either conjugate or nonconjugate priors. It is also applicable when either cross-segment dependence or cross-segment independence occurs. Its good performance in some complicated circumstances is demonstrated through examples from robust Bayesian change point detection using -models, Bayesian change point detection with dependence cross-segment, objective Bayesian change point detection and simulation studies, in which the marginal likelihood of the model is often difficult to obtain or intractable.
In this work, a copula-based pairs-trading strategy is proposed based on a new way of calculating the mispricing indicator. Its performance is evaluated on selected equities in the Brazilian Stock Exchange Index (Ibovespa) and compared to a naive buy-and-hold strategy, and the original indicator is applied to copula-based pairs trading. Additionally, the relationship between the strategies' profitability and key macroeconomic variables is assessed. The results indicate that the copula-based pairs-trading strategy is statistically superior to the buy-and-hold strategy in terms of profitability and risk. Last, although macroeconomic variables are associated with Ibovespa performance, they are neutral in relation to performance with the pairs-trading strategy.
We propose here a new multivariate Birnbaum-Saunders (BS-type) distribution characterized by its leptokurtic property, making it particularly useful in the field of finance. Unlike the approach of Romeiro et al., our proposal is also based on scale mixtures of normal distributions (SMN), but with the mixing variable following a BS distribution, resulting in an asymmetric distribution. This new distribution captures leptokurtic character in the distribution, which implies heavier tails and a more pronounced peak compared to BS or StBS distributions (BS based on the Student-t distribution), enabling more realistic modeling of financial data. The resulting multivariate BS-type distribution is an absolutely continuous distribution whose marginal and conditional distributions have leptokurtic properties as compared to the usual univariate BS distribution. These results are a potentially necessary supplement to the recent work of Romeiro et al. This new distribution has not been discussed yet in the literature, and it enriches the family of multivariate BS distributions as it adds new features that take advantage of the presence of observations quite concentrated around the mode. By using the nice hierarchical representation, we have developed a fast and accurate EM (Expectation-Maximization) algorithm for computing the maximum likelihood estimates, and simulation studies show its good performance, and the corresponding asymptotic properties of the estimates. Finally, we illustrate the results with a real dataset, showcasing the effectiveness and practical utility of the proposed distribution.
This paper studies a class of shock models for a system that is equipped with a protection block that has its own failure rate. Under the considered class, the system exposed to shocks at random times is protected by the protection block, and the probability of the shock damaging the system varies depending on whether the protection block operates or not. The system failure criteria is defined based on the pattern of the critical/damaging shocks. Exact expressions for the reliability and mean time to failure of the system are obtained, and detailed computations are presented for the run shock model, which is included in the class. The application of the extreme shock model, which is included in the relevant class, to wind turbine reliability is also discussed.
Bootstrap procedures represent a straightforward approach to assessing the uncertainty around estimates of interest in statistical models. However, with the rising prevalence of massive datasets in statistical problems, the computational cost of bootstrap methods can quickly become prohibitive in many settings. To this end, this paper proposes the Averaged Robbins-Monro Bootstrap (ARM-B), a scalable tool for estimating parameter variability via multiple chains of Robbins-Monro updates. The method is illustrated in large-scale Poisson regression and logistic regression settings and compared with the alternative scalable method given by the bag of little bootstraps (BLB). Some simulation experiments and an illustrative analysis on a large-scale dataset show that ARM-B has comparable accuracy with ordinary bootstrap, but, at the same time, it is significantly less computationally demanding and quite competitive with BLB.
Fairness is a key requirement for artificial intelligence applications. The assessment of fairness is typically based on group-based measures, such as statistical parity, which compares the machine learning output for the different population groups of a protected variable. Although intuitive and simple, statistical parity may be affected by the presence of control variables, correlated with the protected variable. To remove this effect, we propose to employ Shapley values, which measure the additional difference in output specifically due to the protected variable. To remove the possible impact of correlations on Shapley values, we compare them across different subgroups of the most correlated control variables, checking for the presence of Simpson's paradox, for which a fair model may become unfair when conditioning on a control variable. We also show how to mitigate unfairness by means of a propensity score matching that can improve statistical parity, building a training sample that matches similar individuals in different protected groups. We apply our proposal to a real-world database containing 157,269 personal lending decisions and show that both logistic regression and random forest models are fair when all loan applications are considered, but become unfair for high loan amounts requested. We show how propensity score matching can mitigate this bias.