Recognizing the critical importance of explainable clustering results for decision-making and the influence of sample importance on the clustering result, this study proposes a clustering method based on the centroid data envelopment analysis (DEA) cross-efficiency approach. Specifically, this study first introduces the centroid DEA cross-efficiency approach. The approach is constructed based on the unique set of centroid weights of the convex polytope formed by all optimal weight vectors for each DMU. Then, a gravity model is constructed based on the centroid DEA cross-efficiency approach. The gravity model simultaneously accounts for the sample importance and the distance between samples. Based on the gravity between samples, this study develops the gravity clustering method. This clustering method enhances interpretability and provides decision support by identifying the importance degree of the features for samples across different clusters through centroid weights. To validate the effectiveness, an empirical example is conducted, and the result shows that the proposed clustering method outperforms existing DEA-based clustering approaches. Furthermore, a clustering study is conducted on the healthcare levels of various provinces in China, and policy recommendations are provided for the medical development of provinces within different clusters.
The widespread adoption of online question and answer (Q&A) systems by e-commerce platforms has sparked interest in understanding their potential impacts. Nevertheless, little is known about how various types of consumer questions influence product popularity. Using a dataset from JD.com, this study categorizes consumer questions along the content and temporal dimensions and investigates how different types of questions influence product popularity. The findings reveal that questions concerning core, auxiliary, and peripheral attributes are generally associated with higher product popularity. By comparison, post-purchase questions, particularly those focused on core attributes, are associated with lower product popularity. Further analysis suggests that factors such as question length, product price, and product type moderate these relationships. Specifically, for high-priced search products, longer post-purchase core attribute questions are linked to stronger negative associations with product popularity, whereas longer post-purchase peripheral attribute questions show more positive associations. Questions regarding auxiliary attributes do not show similar patterns. These findings provide crucial theoretical and practical insights for e-commerce platforms and merchants in optimizing Q&A strategies.
Clustering, as an unsupervised technique, does not explain why certain objects are grouped together, which makes it challenging for decision-makers to interpret the characteristics that define each cluster. In this study, we propose a weight space clustering method based on Data Envelopment Analysis (DEA). This method offers a semantic explanation for each cluster, and it uses Monte Carlo simulation to explore feasible weight vectors that satisfy DEA constraints, thereby effectively characterising the weight space of the objects. The similarity between objects is then computed to construct a similarity matrix, which serves as the foundation for clustering. To validate the weight space clustering method, we use a Monte Carlo simulation with a piecewise-linear production function. Our results show that the average Rand index and F-score of the clustering exceed 0.9, demonstrating the effectiveness of the proposed method. Furthermore, comparative experiments show that the internal validity of our clustering results surpasses that of the existing DEA-based clustering methods, as well as conventional k-means and hierarchical clustering approaches. Finally, we apply the proposed method to industrial parks in Hunan Province, China, identifying three distinct clusters: the first is characterised by high water resource consumption, the second also emphasises water consumption but places greater importance on land output and the third is primarily associated with energy consumption. These explainable clustering results provide valuable insights into enhancing production efficiency and optimising resource utilisation in the studied industrial parks.
Data envelopment analysis (DEA) methods for fixed-sum outputs have received increasing attention in recent years. Among them, the generalized equilibrium efficient frontier DEA method (GEEFDEA) has gained widespread application due to its computational convenience. However, its uniform weight assumption is inconsistent with the differentiated weighting principle of traditional DEA models, which simplifies the structure of the generated equilibrium efficient frontier (EEF), limiting its ability to accurately reflect real-world production processes. Although recent studies have addressed the uniform weighting issue in single-stage systems, their approaches are not readily applicable to more complex systems. To address this limitation, this study develops a novel evaluation approach for network decision-making units (DMUs) with shared undesirable fixed-sum outputs. The proposed approach extends the relaxation of the uniform weight assumption to the two-stage network setting, allowing DMUs to select their preference weights when constructing the EEF. Depending on the presence of a central authority, EEF construction models are developed under centralized and decentralized scenarios. After obtaining the adjusted DMUs, the adjusted extreme efficient DMUs are further identified to form the corresponding locally supported extreme efficient DMU set (LSEEDS) for evaluating the original DMUs. Notably, the model for constructing the EEF is nonlinear. We demonstrate that it can be linearized when there is a single fixed-sum output and propose a parametric iterative algorithm for the case of multiple fixed-sum outputs. Finally, we verify the proposed algorithm through a numerical example and apply the approach to evaluate energy production-utilization efficiency in 30 Chinese provincial administrative regions.
Choice behavior reflects consumer preferences. Consumers often purchase products online nowadays, which can be viewed as a choice process. If a consumer makes multiple transactions over a period of time, then we can say the consumer make multiple intertemporal choices. This study focuses on the problem of learning consumer preferences from intertemporal choice data. The main challenges in this research include the ratio relationship between some attributes, the variability of choice set and the uncertainty of attribute values. To address these challenges, we propose a consumer preference model based on the chance-constrained data envelopment analysis (DEA) framework. In the model, we assume consumer choice has maximum utility value, and define a performance cost utility function to capture the ratio relationship between some attributes. We then develop two scenarios for the consumer preference model, depending on whether the uncertain variables are correlated. The estimated consumer preferences can be used to predict each consumer's choice and item ranking. To validate our model, we conduct two numerical experiments, and analyze the impact of some parameters on the preference and evaluation results. The results show that the estimated preference values are accurate when the values of risk indicator and correlation coefficients are small, and our model performs well on the predictions of choice and item ranking.
Network data envelopment analysis (DEA) approach can evaluate decision-making units (DMUs) with network structure. However, in the real world, the total amount of some inputs/outputs is fixed, which is called fixed-sum inputs/outputs. Few studies focus on fixed-sum DEA (FSDEA) evaluation issues with considering the internal network structure of DMUs. Existing studies on fixed-sum research mainly discuss series structures, parallel structures are often neglected although they are common in reality. In this study, we propose a novel fixed-sum DEA efficiency evaluation model for parallel structures, based on generalized equilibrium efficient frontier DEA. Furthermore, depending on whether there's a centralized decision maker, we propose models in decentralized scenario and centralized scenario to construct the efficient frontier. Then, efficiency evaluation models in two scenarios are proposed based on the established efficient frontier respectively. Finally, we demonstrate the practical application of the proposed models by evaluating the industrial performance of three major industries in each province in China for the year 2020. Through this application, we aim to compare the efficiency differences between provinces and provide insights for improving their industrial performances.
Fixed cost allocation is a significant issue for organisations that establish common platforms. Traditional studies on fixed cost allocation based on Data Envelopment Analysis (DEA) usually assume that the inputs and outputs of two-stage systems are deterministic. However, in practice, the input and output data of decision-making units (DMUs) may fluctuate around mean values due to measurement errors or changes in the market environment. Moreover, in the process of fixed cost allocation, a DMU obtains relative gain with reference to a DMU with little allocated cost and obtains relative loss with reference to a DMU with much allocated cost. Decision makers often exhibit risk-seeking behaviour towards losses, reflecting their bounded rationality. To address these issues, this study proposes a novel fixed cost allocation model for two-stage systems based on robust DEA and prospect theory. We first explore robust DEA model for two-stage systems with allocated cost and derive a novel fixed cost allocation possibility set. Building upon this, we incorporate prospect theory to construct a prospect value function for DMUs. Subsequently, a prospect allocation model is formulated by considering the non-cooperative relationship among two-stage DMUs. Finally, an empirical case study and comparative analysis are conducted to validate the proposed approach.
As a cluster of industries where pollutants are generated during production activities, industrial parks bear responsibilities for both economic development and environmental protection. Scientific and objective eco-efficiency evaluation of industrial parks not only clarifies development goals but also promotes the achievement of those goals. However, most existing data envelopment analysis approaches assume the convex production technology, which does not accurately reflect the production systems in industrial parks. To address this, this paper proposes a model to evaluate eco-efficiency and identify benchmarks for industrial parks under natural and managerial disposability within the nonconvex production possibility set. The rationality and superiority of the model proposed in this paper are verified by comparing its results with those of the classical model under weak disposability. The proposed method can clarify the current situation of green and sustainable development of industrial parks and provide scientific improvement targets for inefficient parks.
The opinions of experts often exhibit initial variance in group decision-making process due to the differences in background, knowledge, stance, and other influential factors. Thus, an incentive mechanism is critical to motivate experts to adjust individuals' opinions and achieve a consensus solution. The incentive mechanism is usually costly and included in the consensus reaching process (CRP), and its effect relies on the behavior interaction between the moderator and experts. Within the Stackelberg game framework, we present the maximum-experts and minimum-cost consensus models with multiple incentives and cost budget. First, from a two-side perspective of the moderator and experts, a consensus model with maximum-return modification and maximum-experts feedback (MRMECM) is built to pursue the maximum number of consensus experts under an established cost budget, where the incentive mechanism is realized with the allocation of unit return to each return-driven expert. Then, a double incentive mechanism (DIM) is designed with the modification and shared return incentives. Subsequently, the MRMECM with DIM is constructed to improve the utilization efficiency of cost budget. Finally, the DIM module is integrated into the maximum-return modification and minimum-cost feedback consensus model (MRMCCM), which aims to minimize the consensus cost in the feedback mechanism. All the proposed consensus models are built as bi-level programming models in a unified Stackelberg game framework. We propose a hybrid approach that combines the best-response update strategy with the differential evolution (DE) algorithm to solve the equilibrium solutions. Consequently, several experimental studies are performed to validate the effectiveness of the proposed models.
Clustering algorithms are commonly used to group units based on their intrinsic characteristics, represented by various features, typically divided into cost features and benefit features. Most existing clustering methods analyze these features based on distance functions, and the cost-benefit relationship is not a focus. In this study, we approach clustering from the perspective of constructing a composite index reflecting the cost-benefit ratio. We utilize data envelopment analysis (DEA) to develop a "k-DEA-weights" clustering method to identify clusters and reveal the relative degree of importance of different features within each cluster to enhance the explainability of clustering outcomes. The proposed method can be viewed as an optimization-based explainable machine learning technique. We first employ a DEA cross-efficiency model to calculate cross-evaluated weights for units and a group DEA cross-efficiency model to obtain common weights for clusters. Through an iterative heuristic algorithm, we optimize the clustering process by minimizing the distance between the cross-evaluated weights and the common weights within a cluster. To validate its effectiveness, we conduct numerical experiments that demonstrate strong clustering performance. We then apply it to Chunyu Doctor platform, a Chinese online healthcare consultation platform. The results provide actionable insights for optimizing services and improving user experience.
The literature on incentives suggests that the cost of effort is a key determinant of production, as an agent’s utility is dependent on both consumption and the cost of effort. However, the benchmarking literature has neglected to consider the cost of effort in performance improvement, as it primarily focuses on selecting best-practice benchmarks for underperforming agents. This paper aims to bridge these two literatures by examining the cost of effort in benchmarking and its applications. Our approach is a direct extension of the rational inefficiency hypothesis. We use the information of slacks regarding technology to make inference about the cost of effort in benchmarking. Importantly, we show that inference about the cost of effort gives new sights into activity planning, incentive provision, and employee layoffs. Our analysis provides a new explanation for benchmarking failure in business practices. It also contributes to the rational inefficiency hypothesis by revealing that inefficiency can be beneficial in benchmarking since it can be regarded as a form of fringe reimbursement provided to stakeholders to offset the cost of effort.
Previous studies on profit improvement mainly consider improving traditional profit efficiency including the technical, allocative, and price profit efficiencies. Few studies discuss profit improvement by optimizing the structural arrangement of the network system, which is about adjusting the combinations of upstream and downstream entities in the system. When achieving traditional efficiency, the system's profit can no longer be increased only by reducing inputs, increasing outputs, and adjusting prices. In this situation, structural rearrangement may provide a silver lining to profit improvement. In this study, we take the two-level supply chains as an example and address the research question of how to improve profit via structural rearrangement of the system containing several two-stage subsystems from the efficiency perspective. To answer this question, we view two-stage subsystems in each system as players and propose a data envelopment analysis permutation game method. Through this method, we find the optimal structural rearrangement plan, which is a new way to further improve the profit in addition to improving traditional profit efficiency. We propose a new efficiency measurement, the structural profit efficiency, and find overall profit efficiency is the product of structural profit efficiency and traditional profit efficiency. Finally, we conduct an experiment using the data of 27 supply chains to validate our proposed method and provide the optimal structural rearrangement plans for them.
In classification tasks with large sample sets, the use of a single classifier carries the risk of overfitting. To overcome this issue, an ensemble of classifier models has often been shown to outperform the use of a single “best” model. Given the rich variety of classifier models available, the selection of the high-efficiency classifiers for a given task dataset remains an urgent challenge. However, most of the previous classifier selection methods only focus on the measurement of classification output performance without considering the computational cost. This paper proposes a new ensemble learning method to improve the classification quality for big datasets by using data envelopment analysis. It contains the following two stages: classifier selection and classifier combination. In the first stage, the commonly used classifiers are evaluated on the basis of their performance on resource consumption and classification output performance using the range directional model (RDM); then, the most efficient classifiers are selected. In the second stage, the classifier confusion matrix is evaluated using the data envelopment analysis (DEA) cross-efficiency model. Then, the weight for the classifier combination is determined to ensure that classifiers with higher performance have greater weights based on the cross-efficiency values. Experimental results demonstrate the superiority of the cross-efficiency model over the BCC model and the benchmark voting method in model ensemble. Furthermore, our method has been shown to save more computational resources and yields better results than existing methods.
Effective resource allocation can assist organizations in utilizing infinite resources, thereby improving efficiency and productivity. For an organization, investment is an essential production resource, and how to appropriately allocate investment is one of the most important decisions they need to make. This study aims to rationally allocate investment across each stage of two-stage network structure under two scenarios. We first propose general investment allocation model with adjustment costs to efficiently allocate investments across each stage. Then, given a fixed total investment, we use the DEA-based fixed cost allocation method to achieve the optimal allocation across stages, and we also incorporate two cooperative and noncooperative relationships between two stages. Additionally, capital stock efficiency is defined to analyze the relationship between optimal and actual capital stocks, clarifying the degree to which the optimal capital stock is reached. Finally, the proposed approach is applied to 14 A-share listed companies in China, and suggestions on the optimal amount of investment allocation and capital stock required are provided to facilitate organizational production.
In the era of big data, the rapid solution of large-scale data envelopment analysis model has attracted extensive attention of scholars. Finding efficient DMUs has become a key consideration in large-scale evaluation. The parallel building hull method and the framework method offer good performance in terms of finding efficient DMUs. However, with the surge in time series data and high-frequency data, the evaluation problem of large-scale samples has put forward the requirement for faster and larger scale solutions for traditional methods. In this study, we propose a prescoring method for DMUs by synthesizing the angle and the index, hereafter called the angle-index synthesis method. Meanwhile, we combine this proposed method with the framework method to obtain all efficient DMUs. Numerical experiments and a real application to the bankruptcy data of Polish companies show the significant advantages of our algorithm in terms of computational time in the large-scale sampling context, and optimal parameters can be ignored. Finally, in the Monte Carlo simulation of very large-scale samples, we show that our algorithm has linear increasing trend for the computational time and good validity even in 1 billion DMUs.
Internal validity indices are crucial in evaluating the quality of clustering results, serving as valuable tools for comparing various clustering algorithms and determining the optimal number of clusters for datasets. Most existing internal validity indices use the worst-case scenario to represent the overall validity. Moreover, some indices assign equal weights to distances among different clusters, even when these distances might have varying degrees of influence on the overall validity. Data envelopment analysis (DEA) is an effective technique for evaluating the performance of decision-making units through the computation of the ratio of the weighted sum of outputs to the weighted sum of inputs. The weight assigned to each indicator signifies its degree of influence on efficiency. Furthermore, DEA can be viewed as a multiple-criteria evaluation methodology, wherein inputs and outputs are two sets of performance criteria. We propose a DEA-based internal validity index (DEAI) to evaluate the validity of the clustering results. In this approach, intra-cluster compactness and inter-cluster separation are employed for determining the input(s) and output(s). The DEAI is then applied to the artificial datasets and empirical examples. Experimental results illustrate that DEAI outperforms six classic internal validity indices in accurately identifying the optimal cluster across all 10 datasets.
The sharing economy affords new opportunities to the global tourism industry, accelerating the development of tourism at a striking rate. This paper examines productivity of the bed and breakfast (B & B) industry and how to improve the performance of the B & B industry in a sharing economy. A modified global Malmquist index using slack-based measures is developed to measure productivity, and benchmark selection models using technology forecasting are proposed to improve future performance. The future technology is forecasted by averaging the frontier shifts (technology changes) in previous periods. This study challenges the implicit assumption that the technology will be stationary in traditional benchmarking, and suggests that how the information of frontier shifts derived from Malmquist index can be used to forecast future technology in a rational manner. The empirical results inform that there is a productivity growth in the B&B industry in a sharing economy, and the main driver of the growth is technology progress rather than performance improvement. In particular, a counterintuitive but interesting result is that the B&B industry experienced a slight productivity growth after the outbreak of COVID-19. The benchmarks in 2021 and 2022 are forecasted to improve the performance of the B&B industry.
Convex and nonconvex nonparametric technologies have been applied to estimate plant capacity utilization. However, they face challenges in capturing the production characteristic that involves the simultaneous consideration of increasing, constant, and decreasing marginal production rates along the production surfaces, which may be inconsistent with the standard microeconomic production theory. This inconsistency raises concerns regarding the potential bias of plant capacity utilization estimates. Thus, this paper introduces a new plant capacity utilization measure based on the piecewise Cobb–Douglas technology, aligned with the standard microeconomic production theory. This measure is defined using two multiplicative directional distance functions with optimal endogenous directions. It can be interpreted by the Euclidean distance between two points associated with optimal capacities and the Euclidean norm of one optimal endogenous direction vector. The new measure captures potential production slacks to measure plant capacity in the sense of Pareto–Koopmans efficiency. It also suggests a remedy of non-existence of a maximal plant capacity in the Cobb–Douglas function by utilizing the piecewise Cobb–Douglas technology. The proposed plant capacity utilization measure is validated using a secondary dataset of 19 Chilean hydroelectric power plants.
Witold Pedrycz合作论文数School of Intelligent Systems Science and Engineering, Jinan University;Department of Electrical & Computer Engineering, Faculty of Engineering, University of Alberta2