The density power divergence (DPD) is a well-studied member of the Bregman divergence family and forms the basis of widely used minimum divergence estimators that balance efficiency and robustness. In this paper, we introduce and study a new sub-class of Bregman divergences, termed the exponentially weighted divergence (EWD), designed to generate competitive and practically interpretable inference procedures. The EWD is constructed so that its associated weight function remains bounded within the interval [0, 1], which facilitates a transparent interpretation of robustness through controlled downweighting of low-density observations and avoids excessive influence from high-density points. We develop minimum EWD estimators (MEWDEs) within a general framework accommodating independent but non-homogeneous data, thereby extending classical minimum divergence theory beyond the i.i.d. setting. Under standard regularity conditions, we establish Fisher consistency and asymptotic normality, and we analyze robustness properties through influence function calculations. The EWD framework is further extended to parametric hypothesis testing, for which we derive the asymptotic null distribution of a Bregman divergence-based test statistic. Extensive simulation studies and real-data applications demonstrate that the proposed estimators perform comparably to, and often more robustly than, existing DPD-based procedures, particularly under moderate to heavy contamination, while retaining high efficiency under clean data. Overall, the EWD provides a tractable and interpretable alternative within the Bregman divergence class for robust parametric estimation and testing.
Balancing the efficiency of an estimator under ideal conditions against its robustness under contamination remains a central challenge in robust statistics. While minimum divergence methods offer a flexible alternative to traditional M-estimation, choosing the appropriate discrepancy measure has historically relied on heuristic or empirical justifications. This manuscript introduces a rigorous optimality criterion for this selection process. By investigating the comprehensive Generalized Alpha-Beta Divergence (GABD) family, we explicitly characterize the Pareto frontier dictating the lowest possible asymptotic variance for any strictly enforced asymptotic breakdown point. Our main theoretical results establish that the estimator achieving this mathematical optimum invariably falls within the extended (ϕ, γ)-divergence class. Crucially, the derived optimal tuning parameter, ϕ^*, given other parameters, depends solely on the desired breakdown threshold and is entirely invariant to both the assumed parametric model and the exact nature of the data contamination. Supported by comprehensive derivations of asymptotic normality, influence functions, and breakdown thresholds for both continuous and discrete settings, this work offers a unified, theoretical resolution to the long-standing problem of optimal divergence selection in robust inference.
Covariance matrix estimation is an important problem in multivariate data analysis, both from theoretical as well as applied points of view. Many simple and popular covariance matrix estimators are known to be severely affected by model misspecification and the presence of outliers in the data; on the other hand robust estimators with reasonably high efficiency are often computationally challenging for modern large and complex datasets. In this work, we propose a new, simple, robust and highly efficient method for estimation of the location vector and the scatter matrix for elliptically symmetric distributions. The proposed estimation procedure is designed in the spirit of the minimum density power divergence (DPD) estimation approach with appropriate modifications which makes our proposal (componentwise minimum DPD estimation) computationally very economical and scalable to large as well as higher dimensional datasets. Consistency and asymptotic normality of the proposed componentwise estimators of the multivariate location and scatter are established along with asymptotic positive definiteness of the estimated scatter matrix. Robustness of our estimators are studied by means of influence functions. All theoretical results are illustrated further under multivariate normality. A large-scale simulation study is presented to assess finite sample performances and scalability of our method in comparison to the usual maximum likelihood estimator (MLE), the ordinary minimum DPD estimator (MDPDE) and other popular non-parametric methods. The applicability of our method is further illustrated with a real dataset on credit card transactions.
Robust inference based on the minimization of statistical divergences has proved to be a useful alternative to classical techniques based on maximum likelihood and related methods. Basu et al. (1998) introduced the density power divergence (DPD) family as a measure of discrepancy between two probability density functions and used this family for robust estimation of the parameter for independent and identically distributed data. Ghosh et al. (2017) proposed a more general class of divergence measures, namely the S-divergence family and discussed its usefulness in robust parametric estimation through several asymptotic properties and some numerical illustrations. In this paper, we develop the results concerning the asymptotic breakdown point for the minimum S-divergence estimators (in particular the minimum DPD estimator) under general model setups. The primary result of this paper provides lower bounds to the asymptotic breakdown point of these estimators which are independent of the dimension of the data, in turn corroborating their usefulness in robust inference under high dimensional data.
Composite likelihood (CL) methods provide a computationally efficient alternative to full likelihood inference for complex multivariate models by replacing the joint likelihood with a product of lower-dimensional marginal or conditional components. Like the MLE, however, the maximum CL estimator (MCLE) is highly sensitive to data contamination. On the other hand, robust divergence-based procedures such as the minimum density power divergence (DPD) estimator require the full joint density and so scale poorly to complex multivariate models. We introduce the composite DPD (CDPD), a genuine statistical divergence built entirely from the low-dimensional component densities defining a CL, combining the computational scalability of CL with the robustness of the DPD. The resulting minimum CDPD estimator (MCDPDE) robustifies the MCLE without requiring integration over the full multivariate sample space. We establish consistency, asymptotic normality, and the influence function of the MCDPDE under regularity conditions on the component models alone, without requiring correct specification of the full joint distribution. We show that it is qualitatively robust for every positive value of its tuning parameter, unlike the MCLE recovered as the limit. Because its components can be chosen at the pairwise or cell level, the framework guards simultaneously against casewise and cellwise contamination. Operating directly on component densities rather than elliptical distance structures, it extends robust inference beyond the elliptical models to which most existing cellwise-robust procedures are confined. We develop computational algorithms implemented in the accompanying R package mvdpd. Simulation studies and real-data applications show that the MCDPDE achieves substantial robustness gains over the MCLE while retaining competitive efficiency under the assumed model.
Traditional methods for linear regression generally assume that the underlying error distribution, equivalently the distribution of the responses, is normal. Yet, sometimes real life response data may exhibit a skewed pattern, and assuming normality would not give reliable results in such cases. This is often observed in cases of some biomedical, behavioral, socio-economic and other variables. In this paper, we propose to use the class of skew normal (SN) distributions, which also includes the ordinary normal distribution as its special case, as the model for the errors in a linear regression setup and perform subsequent statistical inference using the popular and robust minimum density power divergence approach to get stable insights in the presence of possible data contamination (e.g., outliers). We provide the asymptotic distribution of the proposed estimator of the regression parameters and also propose robust Wald-type tests of significance for these parameters. We provide an influence function analysis of these estimators and test statistics, and also provide level and power influence functions. Numerical verification including simulation studies and real data analysis is provided to substantiate the theory developed.
The minimum density power divergence estimator (MDPDE) has gained significant attention in the literature of robust inference due to its strong robustness properties and high asymptotic efficiency; it is relatively easy to compute and can be interpreted as a generalization of the classical maximum likelihood estimator. It has been successfully applied in various setups, including the case of independent and non-homogeneous (INH) observations that cover both classification and regression-type problems with a fixed design. While the local robustness of this estimator has been theoretically validated through the bounded influence function, no general result is known about the global reliability or the breakdown behavior of this estimator under the INH setup, except for the specific case of location-type models. In this paper, we extend the notion of asymptotic breakdown point from the case of independent and identically distributed data to the INH setup and derive a theoretical lower bound for the asymptotic breakdown point of the MDPDE, under some easily verifiable assumptions. These results are further illustrated with applications to some fixed design regression models and corroborated through extensive simulation studies.
This paper discusses a new superfamily of divergences that is similar in spirit to the S-divergence family introduced by Ghosh et al. (Bernoulli, 23, 2746–2783. 2017). This new family serves as an umbrella that contains the logarithmic power divergence family (Renyi 1961; Maji et al. RASHI, 2, 39–51. 2017) and the logarithmic density power divergence family (Jones et al. Biometrika, 88, 865–873. 2001) as special cases. Various properties of this new family and the corresponding minimum distance procedures are discussed with particular emphasis on the robustness issue; these properties are demonstrated both theoretically as well as through simulation studies. In particular the method demonstrates the limitation of the first order influence function in assessing the robustness of the corresponding minimum distance procedures. In this respect, we examine the necessity and usefulness of the third order influence functions for the divergence based test statistics, which is studied in the literature for the first time.