Background:Sclerosing adenosis (SA) and breast cancer (BC) often exhibit overlapping clinical, imaging, and pathological characteristics, making them difficult to differentiate. SA may also coexist with BC (SA + BC), including ductal carcinoma in situ (SA-DCIS) and invasive breast cancer (SA-IBC), which complicates diagnosis even when core-needle biopsy (CNB) suggests SA. This study aimed to develop interpretable AI-based binary and ternary classification models that leverage clinical and imaging features to distinguish SA-only from SA + BC and to further differentiate among SA-only, SA-DCIS, and SA-IBC. Methods:We retrospectively analyzed a cohort of 726 patients with SA (January 2006 to December 2021), comprising 537 SA-only and 189 SA + BC cases (90 SA-DCIS, 99 SA-IBC). Multiple machine learning algorithms-logistic regression, support vector machine, decision tree, XGBoost, and random forest-were compared using AUC, accuracy, F1-score, and C-index. Model interpretability was assessed with SHAP to elucidate feature contributions and identify key predictors. Additionally, we incorporated an independent external validation cohort consisting of 113 patients to verify the model's effectiveness. Results:XGBoost consistently outperformed other algorithms in both tasks. Eight features emerged as most informative: age, ultrasound BI-RADS category, maximum and minimum ultrasound diameters, ultrasound margin characteristics, biopsy procedure, mammographic density, and microcalcifications. For binary classification (SA-only vs. SA + BC), XGBoost achieved an AUC of 0.925, accuracy of 0.883, and C-index of 0.844. For ternary classification (SA-only, SA-DCIS, SA-IBC), the model achieved an AUC of 0.888, accuracy of 0.811, and C-index of 0.813. Age, ultrasound BI-RADS, and minimum lesion diameter were consistently top predictors. We further proposed a three-tier interpretability framework (global, cohort-level; local, subgroup-level; and individual, case-level) to facilitate clinical translation. Conclusion:Given the substantial risk of coexisting of SA with DCIS or IBC, and the potential for CNB to underestimate disease due to limited sampling, lesions diagnosed as SA on CNB should be evaluated with additional modalities before determining the need for surgical excision. The proposed interpretable AI model enhances discrimination between SA-only and SA with concomitant breast cancer (SA + BC), thereby supporting more informed clinical decision-making in breast disease management.
While the MammaPrint 70-gene assay has proven valuable for identifying ultra-low-risk breast cancer, its widespread adoption-particularly in primary care settings-remains limited by high costs and technical requirements. In response, there is growing interest in developing simpler, more accessible tools for genetic risk assessment that can complement existing genomic tests. This study introduces a multi-association-driven graph convolutional network (GCN) framework designed to address the challenges of high-dimensional feature spaces and limited sample sizes inherent in 70-gene risk stratification modeling. The proposed approach mitigates these issues by constructing a neighborhood information propagation mechanism that enhances learning across samples. From a local feature subset perspective, a multi-weighted patient graph construction algorithm is developed to capture cross-instance relationships across different dimensions. This graph is then integrated into a GCN to enable multi-dimensional association modeling. Experimental results demonstrate that the proposed framework achieves a classification accuracy of 80.10% and an area under the curve (AUC) of 0.8311. These findings suggest that multi-association modeling and the integration of multi-dimensional sample relationships contribute meaningfully to prediction accuracy. Overall, this work offers a potential technical pathway toward reducing reliance on cost-prohibitive genomic assays and supporting the broader implementation of precision medicine within a tiered healthcare system.
Fragmentary data are prevalent in credit risk scoring due to the unobservable nature of numerous covariates. As a result, models often fail to capture borrowers’ credit risk profiles comprehensively. While prevailing approaches typically rely on deletion or imputation, these techniques can lead to either information loss or distortion. To address this problem, this paper proposes a novel two-stage method for handling fragmentary credit data. In the first stage, we fit candidate models and construct model groups based on the missing patterns of the fragmentary data. In the second stage, the weights of candidate models in each model group are optimized by cross-validation-based Jackknife criterion, which is further validated through simulation experiments. Empirical results from a personal credit data analysis demonstrate that our proposed method significantly outperforms alternative approaches in predictive performance. Specifically, compared to baseline models, our method achieves a peak AUC score of 0.722 on real data. By leveraging the largest possible set of observed cases of fragmentary data without resorting to deletion or imputation, our proposed method provides a more accurate assessment of creditworthiness, ultimately enhancing decision-making processes in lending.
Credit scoring models play a significant role in risk management and decision-making processes within the credit industry. Nevertheless, these models often encounter challenges in dealing with fragmentary data characterized by missing values and varying patterns prevalent in credit datasets, which significantly impedes their predictive accuracy. This study proposed a novel two-stage credit scoring method to effectively handle this fragmentary data without resorting to deletion or imputation tactics. In the first stage, we fit candidate models and construct model groups for each missing pattern - the models within different groups are tailored according to their respective missing patterns. In the second stage, for each missing pattern, the results from the corresponding model group are incorporated into the final prediction through a flexibly chosen link function. The link function, chosen from linear discriminant analysis (LDA), logistic regression (LR), support vector machine (SVM), random forest (RF), and XGBoost, is adept at capturing the intricate nonlinear relationships among the models. To enhance interpretability, the influence of candidate models is elucidated using the Shapley Additive exPlanations (SHAP) method which measures the contribution of each variable. Both simulated and real-world data application results demonstrate that the proposed method surpasses alternative methods in predicting fragmentary data.
An explainable ensemble decision support framework was developed for forecasting ESG stock market prices from the perspective of geopolitical risk. Considering the limitation of single-dimensional geopolitical risk measurement, this paper characterizes geopolitical risk from eight aspects, and systematically study their heterogeneity. The results show that the prediction performance of ESG stock market volatility is worse than that of the return due to geopolitical risk. In detail, the empirical results show that the geopolitical risk has a larger predictive power for the return of SPLAELUT, while the volatility of SPEELMUT. The most prominent and largest predictive contribution is the terror acts (the military buildups) on the ESG stock return (volatility). Additionally, the feature contribution of geopolitical risk on ESG stock market price prediction changes dynamically with different periods, and their morphological characteristics have typical category patterns, which could be clustered into two patterns. Whether return or volatility prediction, the role of geopolitical risk on ESG stock market is similar among SPAPDLUT, SPAPELUT, and SPEELMUT markets. This paper expands the research perspective of ESG stock market similarity and geopolitical risk factor analysis, providing empirical support for the formulation of management strategies for financial regulators and portfolio allocation in securities markets
This study explores the application of multimodal data (online review text, search engine data, holiday calendars, weather, and historical arrivals) to forecast tourism demand across five destinations (Jiuzhaigou, Macau, and Mount Siguniang in China; Hawaii in the United States; and Singapore) during stable and turbulent periods. We develop three multimodal fusion strategies (early, intermediate, and late fusion) and compare their forecasting performance. Early fusion yields the best performance with higher accuracy, and it also achieves better performance compared with traditional unimodal feature extraction methods. Additionally, we extend the Mean Impact Value method to improve the interpretability of multimodal models. This interpretability allows us to understand how different types of information impact prediction results. Beyond providing model transparency, it also offers valuable references for tourism management regarding which types of information are more significant when the environment is uncertain.
Predicting public opinion trends during major infectious disease outbreaks is critical for guiding effective public health responses. However, predicting public opinion remains challenging because it is influenced by socioeconomic, psychological, and media factors. This paper presents a novel framework for predicting public opinion trends related to significant infectious diseases, with a focus on COVID-19 as a case study. The proposed framework identifies the key factors influencing public opinion development and enables both point and interval predictions. The framework uses information ecology theory and applies the NSGA-II algorithm to select the features that best drive public opinion trends. By incorporating this framework, accurate point forecasts are produced alongside prediction intervals, effectively quantifying the uncertainty inherent in public opinion dynamics. This approach minimizes the quality-driven loss function to generate precise prediction intervals, providing decision-makers with critical insights into public opinion fluctuations during epidemics. The results offer valuable, real-time public sentiment warnings, supporting timely and effective interventions in epidemic prevention and control efforts.
This study investigates the predictive value of soft information for consumer loan defaults. We propose a novel framework to address class imbalance by utilizing the concept of Bayesian model averaging. Specifically, we assign unequal weights to machine learning sub-models that incorporate different combinations of variables, thereby creating an accurate and robust model for predicting consumer loan defaults. Additionally, this framework incorporates the Shapley additive explanations (SHAP) method to estimate individual contributions and employs the Bayesian information criterion to assess the variable contributions of the sub-models. We validate the effectiveness and robustness of our proposed method using authentic loan data and publicly available credit default records from a prominent consumer platform in China. Our empirical research suggests that the characteristics of user online behavior are significantly predictive of loan defaults, demonstrating asymmetry at different stages of default.
Jaundice, caused by elevated bilirubin levels, manifests as yellow discoloration of the eyes, mucous membranes, and skin, often serving as a clinical indicator of conditions such as hepatitis or liver cancer. This study introduces a non-invasive, multi-class jaundice detection framework that utilizes weakly supervised pre-training on largescale medical images, followed by transfer learning and fine-tuning on 450 collected jaundice cases. Compared to existing studies, our classification approach is more detailed, encompassing a wider range of jaundice samples, including cases of occult jaundice, thereby enabling the accurate detection of more complex and subtle forms of the condition. Our model demonstrates exceptional performance on an independent test set, achieving an accuracy of 98.9 %, sensitivity of 0.991, specificity of 0.999, AUC of 0.999, and an F1-score of 0.990. Notably, the model's computational efficiency is optimized for mobile deployment, requiring only 0.128 GFLOPs per image. Furthermore, the reliability of the model in identifying nuanced pathological features is validated through SHAP-based interpretability analyses. These findings highlight that weakly supervised pretraining outperforms methods reliant on detailed annotations, providing profound insights into small-sample deep learning applications in medical imaging and paving the way for more precise and scalable diagnostic tools.
Forecasting tourism demand is crucial but challenging, especially with irregular and non-periodic holidays due to mismatches between lunar and Gregorian calendars and the transfer system. Current methods simplify holidays as dummy variables, overlooking their complex impacts on travel demand. This study introduces an H-temporal embedding technique to incorporate holiday schedules and timestamps and integrates it into the Transformer-based Holiformer model. Using multidimensional data, including holidays, weather, historical arrivals, and search engines, we forecast demand for three destinations before and during the COVID-19 pandemic. The experimental results demonstrate the high accuracy and stability of the Holiformer model. Furthermore, we conducted an in-depth analysis of the relationships between various influencing factors in the Holiformer model and tourist arrivals, revealing that the holiday effect in China has a more pronounced impact on tourist numbers than the holiday effect in the United States. This finding provides a new perspective for tourism demand forecasting.
The COVID-19 pandemic continues to destroy the carbon market. To alleviate the situation, governments launched vaccination program campaigns. This study aims to predict two carbon pricing features––return and volatility––considering the impacts of the COVID-19 vaccination program. The present study applies the SHAPley Additive exPlanations method of model analysis and interpretability to determine the forces that predict carbon pricing. Our results show that compared with the volatility of the carbon market, the number of daily vaccinations has better predictive performance in terms of carbon pricing. However, compared with other related control factors, the predictive contribution of the COVID-19 vaccination program to volatility is greater than the return of the carbon market. In addition, a smaller number of daily vaccinations correspond to higher carbon market volatility and lower returns. Our results have crucial implications for investors and policymakers in stabilizing and promoting the carbon market during the COVID-19 pandemic; moreover, our results provide a reference for formulating new COVID-19 vaccination-related policies.
Coastal marine areas are frequently affected by human activities and face ecological and environmental threats, such as algal blooms and climate change. The community structure of phytoplankton-primary producers in marine ecosystems-is highly sensitive to environmental factors, such as temperature, salinity, and nutrients. However, traditional methods for exploring the relationship between phytoplankton communities and environmental factors in eutrophic marine areas are limited by various factors. Therefore, this study employed interpretable machine learning models, integrating high-dimensional data analysis and complex system modeling, to quantitatively and thoroughly analyze the dynamic relationship between phytoplankton communities and environmental variables in high-frequency samples collected over 53 weeks from eutrophic marine areas. The cell abundance of phytoplankton exhibited a distinct "two-peak pattern" variation. Interpretable machine learning model analysis revealed the dynamic contributions of different environmental factors during changes in the phytoplankton community structure. The results showed that temperature was a key environmental factor that affected phytoplankton growth during peak periods. In addition, the contribution of salinity increased during the second peak in phytoplankton abundance, highlighting its central role in the ecological dynamics of this phase. During green tide outbreaks, particularly in Area 01, the contributions of factors such as temperature and salinity increased, whereas those of phosphates and silicates decreased, indicating that green tide outbreaks substantially altered the nutritional dynamics of the ecosystem. Furthermore, different phytoplankton species, such as Skeletonema costatum, , Thalassiosira spp., and Nitzschia spp., exhibit varying responses to environmental factors. Hence, the predictions made using random forest and generalized additive models for phytoplankton cell abundance in two marine areas revealed complex nonlinear relationships between environmental factors, such as temperature, salinity, and phytoplankton abundance.
This study analyzes the effect of governance quality (six aspects: government effectiveness; control of corruption; voice and accountability; regulatory quality; political stability and absence of violence; and rule of law) on the renewable and nonrenewable energy consumption prediction based on the SHapely Additive exPlanations method for model analysis and interpretability. The empirical findings indicate that the time-varying contributions of six aspects of governance quality on nonrenewable (renewable) energy consumption predicting vary greatly in E-7 and G-7 countries. The time-varying contribution of governance quality within countries is heterogeneous and asymmetrical, especially India (Germany) in E-7 countries (G-7 countries). The prediction contribution distribution of governance quality between countries is more discrete in G-7 countries than E-7 countries. Our results are of great importance to policymakers and investors for enhancing the renewable energy consumption level in overcoming environmental challenges based on the country itself through governance quality.
Default risk prediction presents a significant challenge for micro and small enterprises due to the unavailability of comprehensive information databases. This paper develops a default risk management tool based on user portrait theory, utilizing common and objective indicators of micro and small enterprises, such as basic information about entrepreneurs, enterprises, and loans. The Shapley Additive exPlanations (SHAP) method is employed to analyze the contribution of each indicator to default prediction. Empirical results show that household income and personal income are the two most important variables in general, with higher household income associated with a lower probability of default. However, a higher personal income is associated with a higher probability of default. Moreover, the importance of variables and the direction of their relationship with default prediction vary across samples. These findings provide significant insights for developing an accurate default prediction warning system for financial institution managers and policymakers, using the proposed methodology and technical framework.
The identification of spatial layout and functional characteristics among industrial clusters is vital to support the development of regional industries. Based on the industrial registration data of more than 330,000 companies in "The First Industrial Clusters in China," and natural language processing methods (NLP), a set of identification framework for urban industrial functional zones is constructed by introducing commercial registration data and electronic map location data to deal with complex industrial big data, realizing the recognition of industrial spatial layout and functional characteristics of Nanshan. The results show that the industrial coverage of Nanshan is as high as 79.07%, and the wholesale and retail enterprises are its main bodies. Among the nine industrial functional zones, emerging enterprises accounted for the majority and diversified agglomeration areas were more than specialized agglomerations in capital scale and employability. Thus, in future industrial planning, maintaining a diversified industrial agglomeration, while giving more policy favor to functional zones characterized by wholesale and retail, can better stimulate consumption and promote economic development in the region.
Evaluating and understanding the financial impacts of COVID-19 has emerged as an urgent research agenda. Nevertheless, the impacts of government interventions on stock markets remain poorly understood. This study explores, for the first time, the impact of COVID-19 related government intervention policies on different stock market sectors using explainable machine learning-based prediction models. The empirical findings suggest that the LightGBM model provides excellent prediction accuracy while preserving computationally efficient and easy explainability of the model. We also find that COVID-19 government interventions are better predictors of stock market volatility than stock market returns. We further show that the observed effects of government intervention on the volatility and returns of ten stock market sectors are heterogeneous and asymmetrical. Our findings have important implications for policymakers and investors in terms of promoting balance and sustaining prosperity across industry sectors through government interventions.
Significant and composite indices for infectious disease can have implications for developing interventions and public health. This paper presents an investment for developing access to further analysis of the incidence of individual and multiple diseases. This research mainly comprises two steps: first, an automatic and reproducible procedure based on functional data analysis techniques was proposed for analyzing the dynamic properties of each disease; second, orthogonal transformation was adopted for the development of composite indices. Between 2000 and 2019, nineteen class B notifiable diseases in China were collected for this study from the National Bureau of Statistics of China. The study facilitates the probing of underlying information about the dynamics from discrete incidence rates of each disease through the procedure, and it is also possible to obtain similarities and differences about diseases in detail by combining the derivative features. There has been great success in intervening in the majority of notifiable diseases in China, like bacterial or amebic dysentery and epidemic cerebrospinal meningitis, while more efforts are required for some diseases, like AIDS and virus hepatitis. The composite indices were able to reflect a more complex concept by combining individual incidences into a single value, providing a simultaneous reflection for multiple objects, and facilitating disease comparisons accordingly. For the notifiable diseases included in this study, there was superior management of gastro-intestinal infectious diseases and respiratory infectious diseases from the perspective of composite indices. This study developed a methodology for exploring the prevalent properties of infectious diseases. The development of effective and reliable analytical methods provides special insight into infectious diseases’ common dynamics and properties and has implications for the effective intervention of infectious diseases.
In this study, we investigate a new neural network method to solve Volterra and Fredholm integral equations based on the sine-cosine basis function and extreme learning machine (ELM) algorithm. Considering the ELM algorithm, sine-cosine basis functions, and several classes of integral equations, the improved model is designed. The novel neural network model consists of an input layer, a hidden layer, and an output layer, in which the hidden layer is eliminated by utilizing the sine-cosine basis function. Meanwhile, by using the characteristics of the ELM algorithm that the hidden layer biases and the input weights of the input and hidden layers are fully automatically implemented without iterative tuning, we can greatly reduce the model complexity and improve the calculation speed. Furthermore, the problem of finding network parameters is converted into solving a set of linear equations. One advantage of this method is that not only we can obtain good numerical solutions for the first- and second-kind Volterra integral equations but also we can obtain acceptable solutions for the first- and second-kind Fredholm integral equations and Volterra-Fredholm integral equations. Another advantage is that the improved algorithm provides the approximate solution of several kinds of linear integral equations in closed form (i.e., continuous and differentiable). Thus, we can obtain the solution at any point. Several numerical experiments are performed to solve various types of integral equations for illustrating the reliability and efficiency of the proposed method. Experimental results verify that the proposed method can achieve a very high accuracy and strong generalization ability.
This study proposed a two-stage dual-game model methodology to evaluate the existing difficulty of healthcare accessibility in China. First, we analyzed a multi-player El Farol bar game with incomplete information by mixed strategy to explore the Nash equilibrium, and then a weighted El Farol bar game was discussed to identify the existence of a contradiction between supply and demand sides in a tertiary hospital. Second, the overall payoff based on healthcare quality was calculated. In terms of the probability of medical experience reaching that expected level, residents are not optimistic about going to the hospital, and the longer the observation period is, the more pronounced this trend becomes. By adjusting the threshold value to observe the change in the probability of being able to obtain the expected medical experience, it is found that the median number of hospital visits is a key parameter. Going to the hospital did bring benefits to people with consideration of the payoffs, while the benefits varied significantly with the observation period among different months. This study is recommended as a new method and approach to quantitatively assess the tense relationship in access to medical care between the demand and supply sides and a foundation for policy and practice improvements to ensure the efficient delivery of healthcare.
Infectious diseases can cause a sudden and serious spread of public opinion, making it a popular topic for analysis. However, current research on public sentiment mostly relies on traditional machine learning methods, which are limited by sample size and labor costs. This paper presents a novel deep learning model called PCA-BERT, an improved BERT model that utilizes principal component analysis (PCA) to extract and fuse the effective features of each layer of the BERT model. This approach offers a more accurate measurement of public sentiment. Furthermore, this paper proposes an analytical framework to comprehensively study the characteristics of network public opinion evolution in major infectious diseases from three perspectives: content, structure, and behavior. To validate the proposed model and framework, we analyze the COVID-19 pandemic as a case study and collect social media data from the past three years since the outbreak. We calculate public emotions using the PCA-BERT model and combine the obtained emotional values to summarize the temporal and spatial laws of the evolution of network public opinion in terms of content, structure, and behavior. This study can help guide the government to identify public demands during the epidemic and carry out epidemic prevention and control more effectively.