An effective financial distress prediction method can help investors and financial institutions identify risk early, preventing investment loss. We consider the raw financial data that contain more relevant and valuable information as input features. To mine more knowledge and reduce the impact of irrelevant and redundant features, this study investigates a soft probability based random forest method for financial distress prediction. The random forest is employed to build various base decision trees efficiently and the soft probability is employed to select the trees dynamically and combine them. The proposed method was tested on real financial data from listed companies. The results proved that raw financial data can help improve the model’s predictive ability using suitable methods. In addition, the experiments results proved that the proposed method outperforms some other well-known individual classifiers and ensemble methods, whether using raw financial indicators or financial ratios as input. Furthermore, the proposed method has lower misclassification cost than the benchmark methods, indicating its effectiveness in predicting financial distress.
The Stock trend charts encode fine-grained market dynamics and investor reactions, yet their potential remains largely underexplored in corporate financial risk prediction (FRP). Existing studies predominantly rely on numerical financial indicators and textual disclosures, such as annual reports, which may fail to fully capture complex and rapidly evolving risk signals. In this study, we propose a tri-modal FRP framework that jointly leverages numerical financial ratios, textual information, and stock trend chart images. To effectively integrate these heterogeneous data sources, we develop a novel multimodal cross-attention framework comprising three distinct fusion strategies that explicitly model inter-modal interactions and enable complementary information to be fully exploited. Extensive experiments demonstrate that the proposed approach consistently outperforms state-of-the-art methods, achieving accuracy improvements of 0.8%–3.9% and gains of 0.6%–3.8% in the area under the receiver operating characteristic curve (AUC). Moreover, interpretability analysis reveals that image-based market features exert a stronger influence on prediction outcomes than textual disclosures, highlighting the critical role of fine-grained market dynamics in financial risk assessment. These findings provide both theoretical and practical insights into FRP, demonstrating the value of image-based market data and advanced multimodal fusion for more accurate, transparent, and robust risk modeling within the examined market setting.
The limitations of traditional multi-attribute decision-making (MADM) have become increasingly evident, highlighting the need for intelligent MADM. The Elimination and Choice Translating Reality (ELECTRE) III MADM method suffers from tedious calculation caused by pairwise comparison, and cannot effectively determine the indifference, preference, and veto thresholds. Recently, neural networks (NNs) have been widely used to address challenging MADM problems and develop intelligent methods. This study, therefore, focuses on building an intelligent ELECTRE III MADM model based on NNs for intelligent computation, assignment, and decisionmaking. In this case, a multiprocessing calculation algorithm is designed to enable ELECTRE III to perform efficiently in high-volume data environments. Further, an NN-based detection algorithm is designed to detect the threshold parameters of ELECTRE III. Finally, the application of the established intelligent ELECTRE III model is demonstrated using a real-world dataset from QS World University Rankings. The model is shown to have good stability, strong computing power, flexible applicability, and the ability to support practical decision-making.
Following the implementation of China’s 2020 delisting reforms, delisting has evolved into an acute and multidimensional risk that necessitates both accurate prediction and informed decision-making to facilitate mitigation. To address this need, this study introduces a novel three-stage interpretable machine learning approach. By combining 145 integrated variables with the LightGBM algorithm, this study successfully predicts delisting risk and identifies several non-mainstream dominant variables, such as performance-based mispricing, industry tech density, management-based mispricing, capacity utilization and Kaplan-Zingales index. More importantly, this study provides diverse risk-based decision-making options aimed at delisting risk mitigation. These counterfactual options reveal that the effort needed to alter risk-based predictions is similar across most delisting categories, except for material-violation-triggered delisting, which demands significantly greater variable changes. Beyond the universal dimensions of fundamental-based valuation, corporate governance, and market performance, each delisting category exhibits significant heterogeneity in its prioritization of specific dimensions.
The role of a top management team (TMT) with a digital technology background in advancing corporate digital governance, which represents a critical and distinct governance paradigm in the digital era, remains underexplored, particularly from the perspective of dynamic managerial capabilities. Consequently, this study draws on the dynamic managerial capabilities theory to analyze how such a TMT leverages its managerial cognition, human capital, and social capital to promote corporate digital governance. Meanwhile, the extent to which the TMT deploys its dynamic managerial capabilities to influence corporate digital governance is contingent upon organizational (the board with a digital technology background and managerial myopia) and environmental (industry complexity) contexts. Therefore, this study further investigates and finds the relationship strengthened by a board with a digital technology background and by industry complexity and weakened by managerial myopia.
News is reaching investors at an unprecedented rate since pop-up notifications and news recommendations are so common on websites and applications. For investors, these news sources have established themselves as an essential resource for stock market information. To study the impact of news on stock prediction and further enhance its predictive ability, this study innovatively merges the capacity of attention mechanisms to focus crucial information with the ability of temporal convolutional network (TCN) to discern temporal patterns, proposing a new algorithm named “Attention-TCN”. The results indicate that the Attention-TCN model has the smallest prediction error compared to well-known stock prediction algorithms such as Long Short-Term Memory and Gated Recurrent Units. The study demonstrates that incorporating a news perspective and utilizing a TCN model combined with an attention mechanism can hold the promise of helping investors to make more informed decisions and achieving higher returns.
This study seeks to explore the potential of data-driven methods for developing a financial statement fraud prediction model. We emphasize that building a fraud prediction model that can be used to detect fraud in real-world applications should receive attention from researchers. However, the severe class imbalance issue and the complex nature of fraudulent activities make it a rather challenging task. To address these problems, we apply the combinations of different sampling techniques and tree-based ensemble classifiers to an extensive set of raw financial statement data. The results show that the models using an extensive set of raw financial data, undersampling techniques and boosting tree classifiers are superior in fraud detection. Moreover, several features without a priori knowledge are identified to be important for fraud prediction models by feature importance evaluation. Accordingly, this study provides a methodological guide for designing fraud prediction models for real-world applications and serves as a preliminary step of the knowledge discovery process to complement fraud detection knowledge systems.
Missing data are frequently encountered in reality, which inevitably poses great challenges to data mining techniques devoted to data structure identification. In view of the fact that missing data generally exhibits high uncertainty, this paper first introduces the concept of information granule, and performs granular imputation on missing data in a more abstract and inclusive way. With the tolerant nature of the information granule to uncertainty, the error of data imputation and the adverse effects on the subsequent research caused by the low-quality data can be largely reduced. Second, the initial data structure (including granular cluster centers and numeric partition matrix) of the data set with missing values is identified by performing fuzzy clustering on the mixed data set (including both numeric values and information granules) formed by imputation. Third, the bounds of granular cluster centers are further optimized by using the principle of justifiable granularity, and a more robust and reliable granular partition matrix is formed subsequently. Finally, by constructing a reconstruction criterion for mixed data, clustering performance and the optimization of some critical parameters (e.g., the cluster number) used in the proposed method could be investigated. This paper conducts comprehensive experimental studies on both synthetic and publicly available data sets to show the feasibility and effectiveness of the proposed data structure exploration method.
Missing data is frequently encountered in reality, which inevitably poses great challenges to data mining techniques devoted to data structure identification. In view of the fact that missing data (values) generally exhibits high uncertainty, this paper first introduces the concept of information granules, and performs granular imputation on missing data in a more abstract and inclusive way. With the tolerant nature of the information granule to uncertainty, the error of data imputation can be alleviated and the adverse effects on the subsequent research caused by the low-quality data can be reduced. Second, the initial data structure (including granular cluster centers and numerical partition matrix) of the data set with missing values is identified by performing fuzzy clustering on the mixed data set (including numerical and information granules) formed by imputation. Third, the boundaries of granular cluster centers are further optimized by refining the principle of justifiable granularity for mixed data, and a more robust and reliable granular partition matrix is formed subsequently. Finally, by constructing a reconstruction error criterion for mixed data, the critical parameters used in the data clustering process (e.g., the clustering number, the degree of emphasis on the specificity of information granules) are optimized to achieve better data mining results. This paper conducts comprehensive experimental studies on both synthetic and publicly available data sets to show the feasibility and effectiveness of the proposed data structure exploration method.
The outranking multi-criteria decision-making (MCDM) method focuses on the systematic comparison of alternatives and the non-compensation among criteria to conduct decision analysis. Z-numbers are powerful tools for characterizing decision-making information and identifying information reliability. This paper aims to explore Z-number outranking theories and the corresponding MCDM method on the basis of the idea of Elimination and Choice Translating Reality (ELECTRE) III, which is a very popular and capable outranking model. The concordance and discordance indices of Z-numbers are defined by processing their bimodal uncertainty fully. And an objective computation method is investigated to determine the values of thresholds in these indices. Further, three types of novel outranking relations, including dominance, support and opposition relations, for Z-numbers are presented by comparing these indices systematically. Then, an extended MCDM method with the distillation and flow ranking rules is proposed by detecting the outranking relations among alternatives under multiple criteria and carrying out the outranking aggregation and exploitation. To verify the applicability and effectivity of the proposed method, an illustrative example is provided, and a simulation test and a comparison discussion are conducted
As there is a constant trade-off between carbon dioxide emissions against economic growth for every government, carbon efficiency is a key indicator to guide sustainable development. However, the energy crisis and COVID-19 recovery (declined cases of COVID-19 infection, flight recovery, manufacturing restart, and increasing import and export trading) could affect carbon efficiency. Therefore, this paper combines the fuzzy regression discontinuity and random forest algorithm (RF-FRD method) to estimate the discontinuity of the energy crisis and COVID-19 recovery on carbon efficiency. The findings show that discontinuity points in carbon efficiency were induced by the energy crisis and COVID-19 recovery. The positive treatment effect at the first discontinuity point proves that the “zero-tolerance” policies effectively promote carbon efficiency. Besides, the negative treatment effect at the second discontinuity point proves that electricity rationing has not always improved carbon efficiency during the energy crisis.
New energy vehicles (NEVs) are the future of the automotive industry, with their development being highly significant for the energy system and environment. An important challenge faced by the automotive industry is appropriately assessing NEVs, which acts an important role in the promotion of NEV products and development of the NEV industry. This paper constructs an integrated decision support framework to solve NEV evaluation problems on the basis of information reliability, decision perception, and criterion non-compensation. By this means, the Z-number is introduced to reliably characterise the information involved in the problems, a Z-number regret theory model is established to identify the perceived utility of information, and QUALIFLEX outranking exploitation techniques are presented to comprehensively process multiple criterion assessments. To demonstrate the applicability of the constructed decision framework, a case study of NEV evaluation in Zhuzhou City is carried out and the result analysis and management implication are conducted in detail. The study reveals that the pure electric vehicle is evaluated as the most suitable NEV for development and production in Zhuzhou. Crucially, the analysis results confirm that the constructed framework can effectively facilitate NEV evaluation.
Ensemble pruning becomes an important stage in multiple classifier systems, and it has been widely applied to solve binary classification problems. Diversity and performance measures are two widely used evaluation methods to build the selection criterion for ensemble pruning. However, few works consider both of them simultaneously, and they usually use one algorithm to measure the diversity or performance, which may not be enough to capture all the relevant diversities and performance of the base classifiers. To solve this problem, we propose a multiple criteria ensemble pruning method by employing multiple diversity and performance measures to capture the base classifiers’ diversity and evaluate their classification ability respectively. Moreover, a multi-criteria decision making method, based on fuzzy soft set and Dempster-Shafer theory of evidence, is used to build the final selection criterion, which can make a good trade-off between the diversity and performance measures. With sixteen binary data sets, the experimental studies show its effectivity and superiority for ensemble pruning over six state-of-the-art benchmark methods.
This paper aims to produce user-centered explanations for financial fraud detection models based on Explainable artificial intelligence (XAI) methods. By combining an ensemble predictive model with an explainable framework based on Shapley values, we develop a financial fraud detection approach that is accurate and explainable at the same time. Our results show that the explainable framework can meet the requirements of different external stakeholders by producing local and global explanations. Local explanations can help understand why a specific prediction is identi-fied as fraud, and global explanations reveal the overall logic of the whole ensemble model.
Coronavirus disease 2019 (COVID-19) has placed tremendous pressure on supply chain risk management (SCRM) worldwide. Recent technological advances, especially machine learning (ML) technology, have shown the possibility to prevent supply chain risk (SCR) by decreasing the need for human labor, increasing response speed, and predicting risk. However, the literature lacks a comprehensive analysis of the relationship between ML and SCRM. This work conducts a comprehensive review of the relatively limited literature in this field. An analysis of 67 shortlisted articles from 9 databases shows that this area is still in the rapid development stage and that researchers have shown extraordinary interest in it. The main purpose of this study is to review the current research status so that researchers have a clear understanding of the research gaps in this area. Moreover, this study provides an opportunity for researchers and practitioners to pay attention to ML algorithms for SCRM during the COVID-19 pandemic.
With the development of different kinds of techniques, especially the Internet of Things (IoT), a large amount of quantitative (either numeric or categorical) data have been generated, transmitted, and stored in the modern society. People hope to understand the interested phenomenon from the collected quantitative data by utilizing different data analysis methods. Exploring the structure of data (e.g., the cluster centers or prototypes) has always been a hot spot in the domain of data mining and knowledge discovery, yet it seems that the modeling and analyzing process still focus on a low-level abstraction of the data because normally, the structure found is only represented by some numeric data points. In this study, we highlight that a low-level abstraction may not be a user-friendly way for people to grasp the knowledge contained in the data. Instead, we explore the structure of the data from a perspective of symbolic analysis. Specifically, two modes of abstraction are proposed. In the vertical mode (i.e., values of each feature are abstracted), the numeric prototypes are characterized by the symbolic prototypes such that people could get rid of being stuck in minor details of each feature. In the horizontal mode (i.e., values of each prototype are abstracted), the linguistic summarization is used to describe all the features of each symbolic prototype such that people could immediately grasp the essential information conveyed in the symbolic prototype. We conduct comprehensive experimental studies on the publicly available data to illustrate the feasibility and validity of the proposed symbolic analysis process.
Vigorously developing new energy vehicles (NEVs) is a new strategic goal of the automotive industry nowadays. Stocks are important channels for NEV enterprises to raise funds, and stock price prediction is of great significance to equity financing, risk identification and policy formulation of NEV enterprises. Given this, this paper devotes to developing a novel prediction method by the integration of time series and cloud models to manage the stock price prediction of NEV enterprises. In this way, the concept and generator algorithm of time series clouds (TSCs) are presented, the time series techniques are introduced to model stock price data comprehensively, and the TSC model is established to detect the uncertainty of data changes. Finally, a case study for the stock price prediction of NEV enterprises is carried out to elucidate and testify the developed prediction method. The result analysis and discussion demonstrate that this method outperforms other models and can effectively support the stock price prediction of NEV enterprises.
Stock selection for effective investment decisions is a valuable and attractive research interest for many years. Owing to the uncertainty and complexity of the stock market, many fuzzy multicriteria decision-making (MCDM) methods were proposed to solve stock selection problems. However, these methods have difficulty in characterizing unreliable information, which is widespread in the stock market, and handling the non-compensation among multiple criteria. In this paper, an innovative method is developed from the perspectives of information reliability and criterion non-compensation to manage stock selection problems. First, the Z-number, which is a powerful tool for describing real-life information and identifying information reliability, is introduced to depict stock evaluation information. Second, the outranking degree of Z-numbers is defined based on the fuzzy and probability information. Subsequently, some outranking aggregation and exploitation procedures are presented based on the idea of Elimination and Choice Translating Reality (ELECTRE) I to handle the non-compensation among stock evaluation criteria. By integrating the above studies, a Z-number ELECTRE I MCDM method is developed. Finally, a stock investment object selection problem is solved, and some discussions and analyses are conducted to testify the applicability and validity of this method.
Witold Pedrycz合作论文数School of Intelligent Systems Science and Engineering, Jinan University;Department of Electrical & Computer Engineering, Faculty of Engineering, University of Alberta3