Bias in data may lead to prediction procedures which discriminate individuals from sensitive groups. In this paper we propose a Bayesian method for parameter estimation in the linear regression model taking fairness into account. For a given prior structure, namely, the Normal-Gamma prior, for a specific choice of unfairness measure (which does not require the use of private-sensitiveinformation), our method computes a constrained empirical Bayes estimator that induces fairness in posterior predictions. Specifically, the model hyper-parameters are calculated as the optimal solution of a constrained optimization problem: that in which the marginal likelihood is maximized in a region where the unfairness measure is upper bounded and thus, controlled. This guarantees that average posterior predictions are, at most, as unfair as the upper bound establishes. Experiments with synthetic and real data sets show the competitiveness of our approach in terms of fairness in the obtained solution and prediction error.
This article introduces a novel, Bayesian, semi-parametric approach to inference for both elliptical and normal/independent distributions. The location and scale parameters are modelled parametrically and a suitable transformation of the modular variable is modelled using Dirichlet process mixtures. A feature of our approach is that the partial lack of identifiability inherent in both elliptical and normal/independent distributions can be accounted for by incorporating a restriction on the diagonal elements of the scale matrix. Posterior computation is carried out using a Markov chain Monte Carlo algorithm.A novel technique for model selection, based on an approximation of the deviation information criterion, is introduced. As shown by a numerical study based on simulation, the approach can be used to discriminate between elliptical, and normal/independent distributions. Finally, our methodology is illustrated with both simulated and real data.
The wavelet spectra is a common starting point for estimating the Hurst exponent of a self-similar signal using wavelet-based techniques. The decay of the log_2 average energy of the detail wavelet coefficients as a function of the level of signal decomposition can be used to construct estimators for this parameter. In this paper, we expand on previous work which introduced the “dual" wavelet spectra, where decomposition levels are instead treated as a function of energy values, and propose a relationship between its slope and the Hurst exponent by inverting the standard wavelet spectra, thereby creating a new estimator. The effectiveness of this estimator and its sensitivity to several settings are demonstrated through a simulation study. Finally, we show how the technique performs as a feature extraction method by applying it to the task of detecting the presence of breast cancer in mammogram images. Dual spectra wavelet features had a statistically significant effect on the log-odds of Cancer.
In this paper we analyze the well-known 'Anonymous bank' call center dataset from a queueing science viewpoint. For this purpose, fitted distributions for both the inter-arrival and service times as well as for customers patiences are integrated in a simulator to infer quantities of interest related to call centers managerial decisions as waiting times, abandonment rates and queue lengths. In particular, it is shown how a type of Markov renewal process, the Markovian arrival process (MAP), is able to capture some of the characterizing properties of arrivals in a modern call center as overdispersion and positive correlation between arrival counts. The work provides a new inference approach for the MAP based on the count process descriptors and presents new properties concerning the dependence structure of the cumulated number of arrivals in a MAP.
A number of approaches have dealt with statistical assessment of self-similarity, and many of those are based on multiscale concepts. Most rely on certain distributional assumptions which are usually violated by real data traces, often characterized by large temporal or spatial mean level shifts, missing values or extreme observations. A novel, robust approach based on Theil-type weighted regression is proposed for estimating self-similarity in two-dimensional data (images). The method is compared to two traditional estimation techniques that use wavelet decompositions; ordinary least squares (OLS) and Abry-Veitch bias correcting estimator (AV). As an application, the suitability of the self-similarity estimate resulting from the the robust approach is illustrated as a predictive feature in the classification of digitized mammogram images as cancerous or non-cancerous. The diagnostic employed here is based on the properties of image backgrounds, which is typically an unused modality in breast cancer screening. Classification results show nearly 68% accuracy, varying slightly with the choice of wavelet basis, and the range of multiresolution levels used.
Support vector machines (SVMs) are widely used and constitute one of the best examined and used machine learning models for two-class classification. Classification in SVM is based on a score procedure, yielding a deterministic classification rule, which can be transformed into a probabilistic rule (as implemented in off-the-shelf SVM libraries), but is not probabilistic in nature. On the other hand, the tuning of the regularization parameters in SVM is known to imply a high computational effort and generates pieces of information that are not fully exploited, not being used to build a probabilistic classification rule. In this paper we propose a novel approach to generate probabilistic outputs for the SVM. The new method has the following three properties. First, it is designed to be cost-sensitive, and thus the different importance of sensitivity (or true positive rate, TPR) and specificity (true negative rate, TNR) is readily accommodated in the model. As a result, the model can deal with imbalanced datasets which are common in operational business problems as churn prediction or credit scoring. Second, the SVM is embedded in an ensemble method to improve its performance, making use of the valuable information generated in the parameters tuning process. Finally, the probabilities estimation is done via bootstrap estimates, avoiding the use of parametric models as competing approaches. Numerical tests on a wide range of datasets show the advantages of our approach over benchmark procedures.
The Naive Bayes has proven to be a tractable and efficient method for classification in multivariate analysis. However, features are usually correlated, a fact that violates the Naive Bayes' assumption of conditional independence, and may deteriorate the method's performance. Moreover, datasets are often characterized by a large number of features, which may complicate the interpretation of the results as well as slow down the method's execution. In this paper we propose a sparse version of the Naive Bayes classifier that is characterized by three properties. First, the sparsity is achieved taking into account the correlation structure of the covariates. Second, different performance measures can be used to guide the selection of features. Third, performance constraints on groups of higher interest can be included. Our proposal leads to a smart search, which yields competitive running times, whereas the flexibility in terms of performance measure for classification is integrated. Our findings show that, when compared against well-referenced feature selection approaches, the proposed sparse Naive Bayes obtains competitive results regarding accuracy, sparsity and running times for balanced datasets. In the case of datasets with unbalanced (or with different importance) classes, a better compromise between classification rates for the different classes is achieved.
In this article we consider an aggregate loss model with dependent losses. The losses occurrence process is governed by a two-state Markovian arrival process (MAP2), a Markov renewal process process that allows for (1) correlated inter-losses times, (2) non-exponentially distributed inter-losses times and, (3) overdisperse losses counts. Some quantities of interest to measure persistence in the loss occurrence process are obtained. Given a real operational risk database, the aggregate loss model is estimated by fitting separately the inter-losses times and severities. The MAP2 is estimated via direct maximization of the likelihood function, and severities are modeled by the heavy-tailed, double-Pareto Lognormal distribution. In comparison with the fit provided by the Poisson process, the results point out that taking into account the dependence and overdispersion in the inter-losses times distribution leads to higher capital charges.
The Naïve Bayes is a tractable and efficient approach for statistical classification. In general classification problems, the consequences of misclassifications may be rather different in different classes, making it crucial to control misclassification rates in the most critical and, in many realworld problems, minority cases, possibly at the expense of higher misclassification rates in less problematic classes. One traditional approach to address this problem consists of assigning misclassification costs to the different classes and applying the Bayes rule, by optimizing a loss function. However, fixing precise values for such misclassification costs may be problematic in realworld applications. In this paper we address the issue of misclassification for the Naïve Bayes classifier. Instead of requesting precise values of misclassification costs, threshold values are used for different performance measures. This is done by adding constraints to the optimization problem underlying the estimation process. Our findings show that, under a reasonable computational cost, indeed, the performance measures under consideration achieve the desired levels yielding a user-friendly constrained classification procedure.
Motivated by a real failure dataset in a two-dimensional context, this paper presents an extension of the Markov modulated Poisson process (MMPP) to two dimensions. The one-dimensional MMPP has been proposed for the modeling of dependent and non-exponential inter-failure times (in contexts as queuing, risk or reliability, among others). The novel two-dimensional MMPP allows for dependence between the two sequences of inter-failure times, while at the same time preserves the MMPP properties, marginally. The generalization is based on the Marshall-Olkin exponential distribution. Inference is undertaken for the new model through a method combining a matching moments approach with an Approximate Bayesian Computation (ABC) algorithm. The performance of the method is shown on simulated and real datasets representing times and distances covered between consecutive failures in a public transport company. For the real dataset, some quantities of importance associated with the reliability of the system are estimated as the probabilities and expected number of failures at different times and distances covered by trains until the occurrence of a failure.
Since the seminal paper by Bates and Granger in 1969, a vast number of ensemble methods that combine different base regressors to generate a unique one have been proposed in the literature. The so-obtained regressor method may have better accuracy than its components, but at the same time it may overfit, it may be distorted by base regressors with low accuracy, and it may be too complex to understand and explain. This paper proposes and studies a novel Mathematical Optimization model to build a sparse ensemble, which trades off the accuracy of the ensemble and the number of base regressors used. The latter is controlled by means of a regularization term that penalizes regressors with a poor individual performance. Our approach is flexible to incorporate desirable properties one may have on the ensemble, such as controlling the performance of the ensemble in critical groups of records, or the costs associated with the base regressors involved in the ensemble. We illustrate our approach with real data sets arising in the COVID-19 context.
One of the main challenges researchers face is to identify the most relevant features in a prediction model. As a consequence, many regularized methods seeking sparsity have flourished. Although sparse, their solutions may not be interpretable in the presence of spurious coefficients and correlated features. In this paper we aim to enhance interpretability in linear regression in presence of multicollinearity by: (i) forcing the sign of the estimated coefficients to be consistent with the sign of the correlations between predictors, and (ii) avoiding spurious coefficients so that only significant features are represented in the model. This will be addressed by modelling constraints and adding them to an optimization problem expressing some estimation procedure such as ordinary least squares or the lasso. The so-obtained constrained regression models will become Mixed Integer Quadratic Problems. The numerical experiments carried out on real and simulated datasets show that tightening the search space of some standard linear regression models by adding the constraints modelling (i) and/or (ii) help to improve the sparsity and interpretability of the solutions with competitive predictive quality.
The Batch Markov Modulated Poisson Process (BMMPP) is a subclass of the versatile Batch Markovian Arrival Process (BMAP) which has been proposed for the modeling of dependent events occurring in batches (such as group arrivals, failures or risk events). This paper focuses on exploring the possibilities of the BMMPP for the modeling of real phenomena involving point processes with group arrivals. The first result in this sense is the characterization of the two-state BMMPP with maximum batch size equal to K, the BMMPP2(K), by a set of moments related to the inter-event time and batch size distributions. This characterization leads to a sequential fitting approach via a moments matching method. The performance of the novel fitting approach is illustrated on both simulated and a real teletraffic data set, and compared to that of the EM algorithm. In addition, as an extension of the inference approach, the queue length distributions at departures in the queueing system BMMPP/M/1 is also estimated. (C) 2019 Elsevier B.V. All rights reserved.
This paper investigates how the production policy, as well as other factors, affect the facility location-allocation decisions. We focus on a p -median location problem in which one single perishable product is to be produced and shipped to a set of users. The time-correlated demands of the clients are generated by autoregressive processes, and they are forecasted from historical data. Empirically, we show that: (i) embedding the production policy in the location-allocation decision problem may lead to a facilities-clients assignment which does not necessarily correspond to the minimum cost allocation, but produces better profits, (ii) taking into account the autocorrelation of the demand can significantly improve the performance of the supply chain, and (iii) the variability of the demand strongly affects the performance of the supply chain, so a careful choice of production strategy is especially recommended in this case.
Support Vector Machine (SVM) is a powerful tool in binary classification, known to attain excellent misclassification rates. On the other hand, many realworld classification problems, such as those found in medical diagnosis, churn or fraud prediction, involve misclassification costs which may be different in the different classes. However, it may be hard for the user to provide precise values for such misclassification costs, whereas it may be much easier to identify acceptable misclassification rates values. In this paper we propose a novel SVM model in which misclassification costs are considered by incorporating performance constraints in the problem formulation. Specifically, our aim is to seek the hyperplane with maximal margin yielding misclassification rates below given threshold values. Such maximal margin hyperplane is obtained by solving a quadratic convex problem with linear constraints and integer variables. The reported numerical experience shows that our model gives the user control on the misclassification rates in one class (possibly at the expense of an increase in misclassification rates for the other class) and is feasible in terms of running times.
The batch Markov‐modulated Poisson process ( B M M P P ) is a subclass of the versatile batch Markovian arrival process ( B M A P ), which has been widely used for the modeling of dependent and correlated simultaneous events (as arrivals, failures, or risk events). Both theoretical and applied aspects are examined in this paper. On one hand, the identifiability of the stationary B M M P P m ( K ) is proven, where K is the maximum batch size and m is the number of states of the underlying Markov chain. This is a powerful result for inferential issues. On the other hand, some novelties related to the correlation and autocorrelation structures are provided.
Feature Selection is a crucial procedure in Data Science tasks such as Classification, since it identifies the relevant variables, making thus the classification procedures more interpretable, cheaper in terms of measurement and more effective by reducing noise and data overfit. The relevance of features in a classification procedure is linked to the fact that misclassifications costs are frequently asymmetric, since false positive and false negative cases may have very different consequences. However, off-the-shelf Feature Selection procedures seldom take into account such cost-sensitivity of errors. In this paper we propose a mathematical-optimization-based Feature Selection procedure embedded in one of the most popular classification procedures, namely, Support Vector Machines, accommodating asymmetric misclassification costs. The key idea is to replace the traditional margin maximization by minimizing the number of features selected, but imposing upper bounds on the false positive and negative rates. The problem is written as an integer linear problem plus a quadratic convex problem for Support Vector Machines with both linear and radial kernels. The reported numerical experience demonstrates the usefulness of the proposed Feature Selection procedure. Indeed, our results on benchmark data sets show that a substantial decrease of the number of features is obtained, whilst the desired trade-off between false positive and false negative rates is achieved.
In this article we describe a method for carrying out Bayesian estimation for the two-state stationary Markov arrival process (MAP(2)), which has been proposed as a versatile model in a number of contexts. The approach is illustrated on both simulated and real data sets, where the performance of the MAP(2) is compared against that of the well-known MMPP2. As an extension of the method, we estimate the queue length and virtual waiting time distributions of a stationary MAP(2)/G/1 queueing system, a matrix generalization of the M/G/1 queue that allows for dependent inter-arrival times. Our procedure is illustrated with applications in Internet traffic analysis.
The capability of modeling non-exponentially distributed and dependent inter-arrival times as well as correlated batches makes the Batch Markovian Arrival Processes (BMAP) suitable in different real-life settings as teletraffic, queueing theory or actuarial contexts. An issue to be taken into account for estimation purposes is the identifiability of the process. This paper explores the identifiability of the stationary two-state BMAP noted as BMAP 2 (k), where k is the maximum batch arrival size, under the assumptions that both the interarrival times and batches sizes are observed. It is proven that for k ≥ 2 the process cannot be identified. The proof is based on the construction of an equivalent BMAP 2(k) to a given one, and on the decomposition of a BMAP 2 (k) into k BMAP 2 (2)s.
This paper studies in detail different problems concerning the identifiability of the non-stationary version of the MAP2. First, a matrix-based methodology to build equivalent processes is given. Second, a unique, canonical representation of the process, so that the infinite, equivalent versions of a process can be reduced to its canonical counterpart is provided. Finally, a characterization of the process in terms of five descriptors representing moments of the three first inter-event time distributions is given.