Background:Sclerosing adenosis (SA) and breast cancer (BC) often exhibit overlapping clinical, imaging, and pathological characteristics, making them difficult to differentiate. SA may also coexist with BC (SA + BC), including ductal carcinoma in situ (SA-DCIS) and invasive breast cancer (SA-IBC), which complicates diagnosis even when core-needle biopsy (CNB) suggests SA. This study aimed to develop interpretable AI-based binary and ternary classification models that leverage clinical and imaging features to distinguish SA-only from SA + BC and to further differentiate among SA-only, SA-DCIS, and SA-IBC. Methods:We retrospectively analyzed a cohort of 726 patients with SA (January 2006 to December 2021), comprising 537 SA-only and 189 SA + BC cases (90 SA-DCIS, 99 SA-IBC). Multiple machine learning algorithms-logistic regression, support vector machine, decision tree, XGBoost, and random forest-were compared using AUC, accuracy, F1-score, and C-index. Model interpretability was assessed with SHAP to elucidate feature contributions and identify key predictors. Additionally, we incorporated an independent external validation cohort consisting of 113 patients to verify the model's effectiveness. Results:XGBoost consistently outperformed other algorithms in both tasks. Eight features emerged as most informative: age, ultrasound BI-RADS category, maximum and minimum ultrasound diameters, ultrasound margin characteristics, biopsy procedure, mammographic density, and microcalcifications. For binary classification (SA-only vs. SA + BC), XGBoost achieved an AUC of 0.925, accuracy of 0.883, and C-index of 0.844. For ternary classification (SA-only, SA-DCIS, SA-IBC), the model achieved an AUC of 0.888, accuracy of 0.811, and C-index of 0.813. Age, ultrasound BI-RADS, and minimum lesion diameter were consistently top predictors. We further proposed a three-tier interpretability framework (global, cohort-level; local, subgroup-level; and individual, case-level) to facilitate clinical translation. Conclusion:Given the substantial risk of coexisting of SA with DCIS or IBC, and the potential for CNB to underestimate disease due to limited sampling, lesions diagnosed as SA on CNB should be evaluated with additional modalities before determining the need for surgical excision. The proposed interpretable AI model enhances discrimination between SA-only and SA with concomitant breast cancer (SA + BC), thereby supporting more informed clinical decision-making in breast disease management.
While the MammaPrint 70-gene assay has proven valuable for identifying ultra-low-risk breast cancer, its widespread adoption-particularly in primary care settings-remains limited by high costs and technical requirements. In response, there is growing interest in developing simpler, more accessible tools for genetic risk assessment that can complement existing genomic tests. This study introduces a multi-association-driven graph convolutional network (GCN) framework designed to address the challenges of high-dimensional feature spaces and limited sample sizes inherent in 70-gene risk stratification modeling. The proposed approach mitigates these issues by constructing a neighborhood information propagation mechanism that enhances learning across samples. From a local feature subset perspective, a multi-weighted patient graph construction algorithm is developed to capture cross-instance relationships across different dimensions. This graph is then integrated into a GCN to enable multi-dimensional association modeling. Experimental results demonstrate that the proposed framework achieves a classification accuracy of 80.10% and an area under the curve (AUC) of 0.8311. These findings suggest that multi-association modeling and the integration of multi-dimensional sample relationships contribute meaningfully to prediction accuracy. Overall, this work offers a potential technical pathway toward reducing reliance on cost-prohibitive genomic assays and supporting the broader implementation of precision medicine within a tiered healthcare system.
This study explores the application of multimodal data (online review text, search engine data, holiday calendars, weather, and historical arrivals) to forecast tourism demand across five destinations (Jiuzhaigou, Macau, and Mount Siguniang in China; Hawaii in the United States; and Singapore) during stable and turbulent periods. We develop three multimodal fusion strategies (early, intermediate, and late fusion) and compare their forecasting performance. Early fusion yields the best performance with higher accuracy, and it also achieves better performance compared with traditional unimodal feature extraction methods. Additionally, we extend the Mean Impact Value method to improve the interpretability of multimodal models. This interpretability allows us to understand how different types of information impact prediction results. Beyond providing model transparency, it also offers valuable references for tourism management regarding which types of information are more significant when the environment is uncertain.
Interaction of cells and the surrounding lumen drives the formation of tubular system that plays the transport and exchange functions within an organism. The physical and biological mechanisms of lumen expansion have been explored. However, how cells communicate and coordinate with the surrounding lumen, leading to continuous tube expansion to a defined geometry, is crucial but remains elusive. In this study, we utilized the ascidian notochord tube as a model to address the underlying mechanisms. We first quantitatively measured and calculated the geometric parameters and found that tube expansion experienced three distinct phases. During the growth processes, we identified and experimentally demonstrated that both Rho GTPase Cdc42 signaling-mediated cell cortex distribution and the stability of tight junctions (TJs) were essential for lumen opening and tube expansion. Based on these experimental data, a conservation-laws-based tube expansion theory was developed, considering critical cell communication pathways, including secretory activity through vesicles, asymmetric cortex tension driven anisotropic lumen geometry, as well as the TJs gate barrier function. Moreover, by estimating the critical tube expansion parameters from experimental observation, we successfully predicted tube growth kinetics under different conditions through the combination of computational and experimental approaches, highlighting the coupling between actomyosin-based active mechanics and hydraulic processes. Taken together, our findings identify the critical cellular regulatory factors that drive the biological tube expansion and maintain its stability.
Jaundice, caused by elevated bilirubin levels, manifests as yellow discoloration of the eyes, mucous membranes, and skin, often serving as a clinical indicator of conditions such as hepatitis or liver cancer. This study introduces a non-invasive, multi-class jaundice detection framework that utilizes weakly supervised pre-training on largescale medical images, followed by transfer learning and fine-tuning on 450 collected jaundice cases. Compared to existing studies, our classification approach is more detailed, encompassing a wider range of jaundice samples, including cases of occult jaundice, thereby enabling the accurate detection of more complex and subtle forms of the condition. Our model demonstrates exceptional performance on an independent test set, achieving an accuracy of 98.9 %, sensitivity of 0.991, specificity of 0.999, AUC of 0.999, and an F1-score of 0.990. Notably, the model's computational efficiency is optimized for mobile deployment, requiring only 0.128 GFLOPs per image. Furthermore, the reliability of the model in identifying nuanced pathological features is validated through SHAP-based interpretability analyses. These findings highlight that weakly supervised pretraining outperforms methods reliant on detailed annotations, providing profound insights into small-sample deep learning applications in medical imaging and paving the way for more precise and scalable diagnostic tools.
Accurate copper price forecasting is crucial and challenging due to the uncertainty and complex fluctuations caused by various factors of financial markets. In this area, the single-factor point prediction methods have made significant contributions but do not fully consider the influence of multiple factors and the robustness of the predictions. This study develops a novel hybrid interval prediction framework that combines multi-objective optimization with quantile deep learning for copper price prediction. The framework holistically evaluates the fluctuation range by assessing the distribution of copper prices and incorporates multiple variables chosen through diverse feature selection methods, which are crucial for accurate copper price prediction. The proposed framework encompasses two sub-stages: (1) initial interval prediction and quantile deep learning models of copper price; (2) multi-objective optimization procedure. In the first phase, four probabilistic forecasting algorithms are employed to sharpen prediction accuracy and provide a comprehensive picture of the interpretation of the outcome parameters by evaluating the distribution. The subsequent phase delves deeper to enhance prediction precision. Four multi-objective optimization algorithms are harnessed to refine the predictions, aiming to boost their reliability and resolution. The experiment findings underscore the superior predicted capabilities of the Quantile Regression Long Short-Term Memory (QRLSTM) model when optimized using the Multi-Objective Salp Swarm Algorithm (MOSSA), achieving a Prediction Interval Coverage Probability of 94.5205%, a Prediction Interval Normalized Average Width of 0.0066, and an Average Interval Score of -373.9687 at 95% confidence levels. The probabilistic forecasting framework developed in this research is reliable and comprehensive, considering many factors influencing the copper price.
Crude oil price forecasting has been one of the research hotspots in the field of energy economics, which plays a crucial role in energy supply and economic development. However, numerous influencing factors bring serious challenges to crude oil price forecasting, and existing research has room for further improvement in terms of an integrated research roadmap that combines impact factor analysis with predictive modelling. This study aims to examine the impact of financial market factors on the crude oil market and to propose a nonlinear combined forecasting framework based on common variables. Four types of daily exogenous financial market variables are introduced: commodity prices, exchange rates, stock market indices, and macroeconomic indicators for ten indicators. First, various variable selection methods generate different variable subsets, providing more diversity and reliability. Next, common variables in the subset of variables are selected as key features for subsequent models. Then, four models predict crude oil prices using common features as inputs and obtain the prediction results for each model. Finally, the nonlinear mechanism of the deep learning technology is introduced to combine above single prediction results. Experimental results reveal that commodity and foreign exchange factors in financial markets are critical determinants of crude oil market volatility over the long term, as observed in experiments conducted on the West Texas Intermediate and Brent oil price datasets. The proposed model demonstrates strong performance regarding average absolute percentage error, recorded at 2.9962% and 2.4314%, respectively, indicating high forecasting accuracy and robustness. This forecasting framework offers an effective methodology for predicting crude oil prices and enhances understanding the crude oil market.
Forecasting tourism demand is crucial but challenging, especially with irregular and non-periodic holidays due to mismatches between lunar and Gregorian calendars and the transfer system. Current methods simplify holidays as dummy variables, overlooking their complex impacts on travel demand. This study introduces an H-temporal embedding technique to incorporate holiday schedules and timestamps and integrates it into the Transformer-based Holiformer model. Using multidimensional data, including holidays, weather, historical arrivals, and search engines, we forecast demand for three destinations before and during the COVID-19 pandemic. The experimental results demonstrate the high accuracy and stability of the Holiformer model. Furthermore, we conducted an in-depth analysis of the relationships between various influencing factors in the Holiformer model and tourist arrivals, revealing that the holiday effect in China has a more pronounced impact on tourist numbers than the holiday effect in the United States. This finding provides a new perspective for tourism demand forecasting.
Ground-level ozone has emerged as a major pollutant in China, where maintaining data integrity and prediction accuracy is crucial for effective environmental governance and policymaking. Nevertheless, the inevitable missingness in ozone time series data compromises dataset completeness, thereby posing a significant obstacle to downstream prediction and advanced analysis. In this paper, we propose a generative imputation and forecasting method, SSSD-Transformer, that injects empirical knowledge as conditional information and combines structured state space diffusion (SSSD) and Transformer to capture the temporal trends and detailed patterns exhibited by ozone pollution for generating the clean ozone sequences. The experiments and case studies on three different missingness scenarios clearly demonstrate that SSSD-Transformer generates more accurate and robust imputations, exhibiting greater reliability. In forecasting task, the model also demonstrated strong capabilities, with performance remaining relatively stable as the prediction horizon increased. Empowered by the potent empirical knowledge and weighted combination, the presented method successfully achieves excellent performance in ozone concentration imputation and forecasting. The simultaneous achievements provide a new perspective for ensuring data integrity and enhancing prediction credibility, and will make significant contributions to effective air pollution control.
Coastal marine areas are frequently affected by human activities and face ecological and environmental threats, such as algal blooms and climate change. The community structure of phytoplankton-primary producers in marine ecosystems-is highly sensitive to environmental factors, such as temperature, salinity, and nutrients. However, traditional methods for exploring the relationship between phytoplankton communities and environmental factors in eutrophic marine areas are limited by various factors. Therefore, this study employed interpretable machine learning models, integrating high-dimensional data analysis and complex system modeling, to quantitatively and thoroughly analyze the dynamic relationship between phytoplankton communities and environmental variables in high-frequency samples collected over 53 weeks from eutrophic marine areas. The cell abundance of phytoplankton exhibited a distinct "two-peak pattern" variation. Interpretable machine learning model analysis revealed the dynamic contributions of different environmental factors during changes in the phytoplankton community structure. The results showed that temperature was a key environmental factor that affected phytoplankton growth during peak periods. In addition, the contribution of salinity increased during the second peak in phytoplankton abundance, highlighting its central role in the ecological dynamics of this phase. During green tide outbreaks, particularly in Area 01, the contributions of factors such as temperature and salinity increased, whereas those of phosphates and silicates decreased, indicating that green tide outbreaks substantially altered the nutritional dynamics of the ecosystem. Furthermore, different phytoplankton species, such as Skeletonema costatum, , Thalassiosira spp., and Nitzschia spp., exhibit varying responses to environmental factors. Hence, the predictions made using random forest and generalized additive models for phytoplankton cell abundance in two marine areas revealed complex nonlinear relationships between environmental factors, such as temperature, salinity, and phytoplankton abundance.
The COVID-19 epidemic and the Russian-Ukrainian conflict have created significant uncertainty in the crude oil market, greatly increasing the difficulty of crude oil futures price volatility forecasting. The objective of this study is to investigate the time-varying nature of the impact of major events on the crude oil market and to examine how to accurately predict crude oil futures price volatility during major shocks. To this end, we propose a novel crude oil futures price volatility forecasting framework based on the Bidirectional Long and Short-Term Memory Neural Network-Attention Mechanism Model (Bi-LSTM-Attention), which contains period division, variable screening, model prediction, and assessment indicators, to further analyze in-depth the impacts of COVID-19 and the Russia-Ukraine conflict on the volatility of crude oil futures price. In addition, we also investigate the impact of external factors on the crude oil market, including economic status, geopolitical events, and the COVID-19 pandemic. Empirical evidence demonstrates that the impact of major events on crude oil futures price volatility is significantly time-varying and differentiated. Particularly during high shock periods, the Fake News Index plays an extremely influential role in driving crude oil futures price volatility. Moreover, our proposed model reliably captures the trend of crude oil futures price volatility during these violent shocks. The contributions of this study are to provide valuable insights into the impact of significant events on crude oil price volatility and contribute to a deeper understanding of crude oil market dynamics.
In the electricity market, the accuracy of electricity price forecasting is significant for real-time control; however, the complexity and volatility of electricity prices make this a challenge. Existing forecasting models focus on deterministic forecasting and rarely address the uncertainty in electricity price forecasting. Therefore, this study fills this knowledge gap by introducing a novel combined probability forecasting system (CPFS) and creatively incorporating probability density estimation based on kernel functions in a multi-objective optimization algorithm. In addition, to effectively integrate the forecasting components, the tuna optimization algorithm was enhanced to overcome the limitations of traditional multi-objective optimization algorithms. Finally, the validity of the CPFS is confirmed through two electricity price cases, considering three equally important aspects: reliability, resolution, and sharpness. From a comprehensive perspective, CPFS outperformed the most advanced benchmark by more than 5.66% and 38.93% in AIS and by more than 13.41% and 3.55% in quantile loss on the NSW and Singapore datasets, respectively. The experimental results demonstrate that the CPFS provides an effective range for electricity price fluctuations. Furthermore, given that probabilistic forecasting is essential for risk management, it offers important implications for the electricity market.
Crude oil plays a vital role in industrial and social development and has become an integral part of the economic development. However, influenced by policies, wars, etc., it is hard to capture the trend of complex and volatile crude oil price if only one model is used, which often leads to poor forecasts. To enhance the prediction accuracy and robustness of forecasting models, a novel combined forecasting method with time-varying weights, i.e., Jaynes weight hybrid ( JWH ) model incorporating the Shannon information entropy and several forecasting methods is proposed in this paper. In the selection of the baseline models, the autoregressive integrated moving average in classical statistical forecasting strategies, back propagation neural network, extreme learning machine in neural network and long short-term memory neural network in deep learning models are chosen to fit the crude oil price. Four datasets and three experiments are constructed to verify the prediction ability of the novel combined forecasting model. Empirical results are calculated by five measurement criteria, suggesting that the prediction accuracy of the novel combined method is significantly higher than several comparison models and the mean absolute percentage error of the model has arrived 2.81 % in detail. Particularly, the proposed model has achieved satisfactory performances in Covid-19 and the War in Ukraine, further verifying the robustness of the combined methodology.
Accurately predicting residential solar energy consumption is crucial for efficient electricity production, supply, and power dispatch. However, conventional forecasting methods often struggle to handle complex energy consumption data. In response to this challenge, this study develops a pioneering two-stage error-corrected combined forecasting model that integrates traditional linear methods, seasonal processing techniques, deep learning models, and intelligent optimization algorithms to outperform other combined forecasting methods in terms of performance. This research analyzes the combined weight values, shedding light on why the proposed model consistently outperforms its counterparts. To confirm its superiority, the proposed model and five benchmark models are rigorously tested in this paper using four evaluation metrics and a hypothesis testing method. The empirical results show that the proposed combined model performs well in terms of accuracy and stability. Notably, the average absolute percentage error of its 24-step ahead prediction is 2.9053 %, which outperforms all comparative models, both single and combined model. These results fully illustrate the advantages of the combined model and reaffirm the excellence of its prediction performance in predicting energy consumption.
Accurate prediction of coal consumption is crucial to the structural adjustment and high-quality development of the coal industry, and provides an effective basis for the government to formulate energy strategies. To enhance the prediction accuracy, a hybrid prediction model combining grey wolf optimization algorithm and grey Markov model (GWO_Markov_DNGM) is proposed in this study, which introduces the idea of Markov interval division in order to make up the defect of obtaining interval by experience in the past. In this research, coal consumption in Hebei Province from 2001 to 2020 is used as an example to verify the accuracy of the model and the results show that the proposed model has higher accuracy than the comparison models, including ARIMA, GM (1, 1), DGM (1, 1), DNGM (1, 1) and the grey Markov model with equidistant division of state interval. In addition, this study also discusses the influence of numbers of state intervals on the prediction accuracy of the proposed model, which further illustrates the high prediction accuracy of the proposed hybrid model. Finally, the proposed model is employed to predict the coal consumption of Hebei Province under three scenarios from 2021 to 2025, and several suggestions are put forward.
In the current situation of energy supply shortage and surging demand, effective and stable load forecasting is essential to ensure reliable power supply and the security of the power system. However, due to some factors such as periodicity and seasonality, the power load sequence shows complex nonlinear characteristics. Meanwhile, the current load forecasting lacks the ability to explore the data deeply, and it is difficult to accurately predict the short-term trend and fluctuation range. To remedy these limitations, this study proposes a hybrid point-interval prediction system (HPILS). The system integrates data preprocessing, optimal model selection, multi-objective optimization combination and interval prediction modules. To verify the performance of the proposed system, four load data sets in Australia are used as examples to conduct experiments. The experimental results demonstrate that HPILS can effectively provide the predicting power load trend changes and fluctuation ranges. Specifically, compared with the benchmark model, HPILS has a 13.47% ∼ 67.89% improvement in point prediction and a 1.67% ∼ 72.08% improvement in interval prediction. In addition, a series of discussion tests are performed to verify the superiority of the proposed system and further confirm the validity of our proposed system.
Nowadays, construction waste has become a global challenge that needs to be addressed, and its accurate prediction is crucial for subsequent treatment and policy development. However, current popular big data methods struggle to be effectively applied due to sample size limits. Meanwhile, the key issues of data uncertainty and time-delayed nature must be considered, the former of which can be addressed by interval prediction. Therefore, a novel interval time-delayed threeparameter discrete grey model for construction waste prediction is proposed in this study, whose time-delayed coefficient is optimized via optimization algorithms. Moreover, the relevant mathematical derivation and proof of this model are also described and discussed in detail in the research. Two case studies of the construction waste dataset are used to examine the effectiveness of the proposed model. Through a series of comparisons with other state-of-the-art models, the results show that the proposed model can improve the prediction performance with excellent prediction results. Meanwhile, this study also presents scenario analysis and discussions of the construction waste prediction results during the 14th Five-Year Plan period, and finds that the current policies and measures are difficult to achieve the future planning goals, and further policy development and implementation of measures are imminent.
Accurate small-sample prediction is an urgent, very difficult, and challenging task due to the quality of data storage restricted in most realistic situations, especially in developing countries. The grey model performs well in small-sample prediction. Therefore, a novel multivariate grey model is proposed in this study, called FBNGM (1, N, r), with a fractional order operator, which can increase the impact of new information and background value coefficient to achieve high prediction accuracy. The utilization of an intelligence optimization algorithm to tune the parameters of the multivariate grey model is an improvement over the conventional method, as it leads to superior accuracy. This study conducts two sets of numerical experiments on CO2 emissions to evaluate the effectiveness of the proposed FBNGM (1, N, r) model. The FBNGM (1, N, r) model has been shown through experiments to effectively leverage all available data and avoid the problem of overfitting. Moreover, it can not only obtain higher prediction accuracy than comparison models but also further confirm the indispensable importance of various influencing factors in CO2 emissions prediction. Additionally, the proposed FBNGM (1, N, r) model is employed to forecast CO2 emissions in the future, which can be taken as a reference for relevant departments to formulate policies.
The identification of spatial layout and functional characteristics among industrial clusters is vital to support the development of regional industries. Based on the industrial registration data of more than 330,000 companies in "The First Industrial Clusters in China," and natural language processing methods (NLP), a set of identification framework for urban industrial functional zones is constructed by introducing commercial registration data and electronic map location data to deal with complex industrial big data, realizing the recognition of industrial spatial layout and functional characteristics of Nanshan. The results show that the industrial coverage of Nanshan is as high as 79.07%, and the wholesale and retail enterprises are its main bodies. Among the nine industrial functional zones, emerging enterprises accounted for the majority and diversified agglomeration areas were more than specialized agglomerations in capital scale and employability. Thus, in future industrial planning, maintaining a diversified industrial agglomeration, while giving more policy favor to functional zones characterized by wholesale and retail, can better stimulate consumption and promote economic development in the region.
This study attempts to explore the bond between Pakistani exports, gross capital formation, energy use and carbon dioxide emission. It uses the data from Pakistan spanning over a longer time horizon of 40 years ranging from 1981 to 2020. The results of ARDL analysis show that Pakistani exports have inverse relationship with CO2 emission in both short as well as long run and the carbon emission reverts to equilibrium at the speed of 54.9%. Increase in carbon emission also lowers export but the causal relationship is only from exports to carbon emission. Energy utilization results in higher carbon emission both in short as well as long run. Based on above findings this study suggests that Pakistan should increase its exports to improve the position of its balance of payment because higher exports do not harm environment in case of Pakistan.