By capturing the prevailing sentiment and market mood, textual data has become increasingly vital for forecasting commodity prices, particularly in metal markets. However, the effectiveness of lightweight, finetuned large language models (LLMs) in extracting predictive signals for aluminum prices, and the specific market conditions under which these signals are most informative, remains under-explored. This study generates monthly sentiment scores from English and Chinese news headlines (Reuters, Dow Jones Newswires, and China News Service) and integrates them with traditional tabular data, including base metal indices, exchange rates, inflation rates, and energy prices. We evaluate the predictive performance and economic utility of these models through long-short simulations on the Shanghai Metal Exchange from 2007 to 2024. Our results demonstrate that during periods of high volatility, Long Short-Term Memory (LSTM) models incorporating sentiment data from a finetuned Qwen3 model (Sharpe ratio 1.04) significantly outperform baseline models using tabular data alone (Sharpe ratio 0.23). Subsequent analysis elucidates the nuanced roles of news sources, topics, and event types in aluminum price forecasting.
BackgroundLarge language models (LLMs) can generate outputs understandable by humans, such as answers to medical questions and radiology reports. With the rapid development of LLMs, clinicians face a growing challenge in determining the most suitable algorithms to support their work. ObjectiveWe aimed to provide clinicians and other health care practitioners with systematic guidance in selecting an LLM that is relevant and appropriate to their needs and facilitate the integration process of LLMs in health care. MethodsWe conducted a literature search of full-text publications in English on clinical applications of LLMs published between January 1, 2022, and March 31, 2025, on PubMed, ScienceDirect, Scopus, and IEEE Xplore. We excluded papers from journals below a set citation threshold, as well as papers that did not focus on LLMs, were not research based, or did not involve clinical applications. We also conducted a literature search on arXiv within the same investigated period and included papers on the clinical applications of innovative multimodal LLMs. This led to a total of 270 studies. ResultsWe collected 330 LLMs and recorded their application frequency in clinical tasks and frequency of best performance in their context. On the basis of a 5-stage clinical workflow, we found that stages 2, 3, and 4 are key stages in the clinical workflow, involving numerous clinical subtasks and LLMs. However, the diversity of LLMs that may perform optimally in each context remains limited. GPT-3.5 and GPT-4 were the most versatile models in the 5-stage clinical workflow, applied to 52% (29/56) and 71% (40/56) of the clinical subtasks, respectively, and they performed best in 29% (16/56) and 54% (30/56) of the clinical subtasks, respectively. General-purpose LLMs may not perform well in specialized areas as they often require lightweight prompt engineering methods or fine-tuning techniques based on specific datasets to improve model performance. Most LLMs with multimodal abilities are closed-source models and, therefore, lack of transparency, model customization, and fine-tuning for specific clinical tasks and may also pose challenges regarding data protection and privacy, which are common requirements in clinical settings. ConclusionsIn this review, we found that LLMs may help clinicians in a variety of clinical tasks. However, we did not find evidence of generalist clinical LLMs successfully applicable to a wide range of clinical tasks. Therefore, their clinical deployment remains challenging. On the basis of this review, we propose an interactive online guideline for clinicians to select suitable LLMs by clinical task. With a clinical perspective and free of unnecessary technical jargon, this guideline may be used as a reference to successfully apply LLMs in clinical settings.
The recent discovery of genetic mutations in Plasmodium falciparum —the most lethal malaria parasite—that enable it to overcome the protective effects of sickle cell trait, raises fundamental questions about the underlying biological and evolutionary interactions. Here we develop a geostatistical model to compare sickle haemoglobin genotype frequencies to the Plasmodium falciparum sickle-associated alleles across global populations, and find a robust association at multiple geographical scales, implying that sickle drives positive selection for these parasite mutations. A model of parasite evolution and an analysis of local haplotype patterns suggest that key features of these mutations – that they are polymorphic in all African populations and are mutually correlated despite lying in different genome regions - are caused by geographical variation in selection pressure, and that the alleles may have been maintained by balancing selection over timescales comparable to the age of the sickle mutation itself. The predicted impact of this host-parasite interaction on disease outcomes varies widely across populations, and functional data are needed to discover the biological mechanisms involved. ### Competing Interest Statement The authors have declared no competing interest. Wellcome Trust, https://ror.org/029chgv08, 304926/Z/23/Z National Natural Science Foundation of China, T2350610281, 82273731 Zhejiang University Education Foundation Global Partnership Fund, 188170-11103
Surface soil moisture is projected to decrease under global warming. Such projections are mostly based on climate models, which show large uncertainty (i.e., inter-model spread) partly due to inadequate observational constraint. Here we identify strong physically-based emergent relationships between soil moisture change (2070–2099 minus 1980–2014) and recent air temperature and precipitation trends across an ensemble of climate models. We extend the commonly used univariate Emergent Constraints to a bivariate method and use observed temperature and precipitation trends to constrain global soil moisture changes. Our results show that the bivariate emergent constraints can reduce soil moisture change uncertainty by 7.87
In spatial regression models, spatial heterogeneity may be considered with either continuous or discrete specifications. The latter is related to delineation of spatially connected regions with homogeneous relationships between variables (spatial regimes). Although various regionalization algorithms have been proposed and studied in the field of spatial analytics, methods to optimize spatial regimes have been largely unexplored. In this paper, we propose two new algorithms for spatial regime delineation, two-stage K-Models and Regional-K-Models. We also extend the classic Automatic Zoning Procedure to spatial regression context. The proposed algorithms are applied to a series of synthetic datasets and two real-world datasets. Results indicate that all three algorithms achieve superior or comparable performance to existing approaches, while the two-stage K-Models algorithm largely outperforms existing approaches on model fitting, region reconstruction, and coefficient estimation. Our work enriches the spatial analytics toolbox to explore spatial heterogeneous processes.
BACKGROUND:Reliable and detailed data on the prevalence of tuberculosis (TB) with sub-national estimates are scarce in Ethiopia. We address this knowledge gap by spatially predicting the national, sub-national and local prevalence of TB, and identifying drivers of TB prevalence across the country.METHODS:TB prevalence data were obtained from the Ethiopia national TB prevalence survey and from a comprehensive review of published reports. Geospatial covariates were obtained from publicly available sources. A random effects meta-analysis was used to estimate a pooled prevalence of TB at the national level, and model-based geostatistics were used to estimate the spatial variation of TB prevalence at sub-national and local levels. Within the MBG Plugin Framework, a logistic regression model was fitted to TB prevalence data using both fixed covariate effects and spatial random effects to identify drivers of TB and to predict the prevalence of TB.RESULTS:The overall pooled prevalence of TB in Ethiopia was 0.19% [95% confidence intervals (CI): 0.12%-0.28%]. There was a high degree of heterogeneity in the prevalence of TB (I2 96.4%, P <0.001), which varied by geographical locations, data collection periods and diagnostic methods. The highest prevalence of TB was observed in Dire Dawa (0.96%), Gambela (0.88%), Somali (0.42%), Addis Ababa (0.28%) and Afar (0.24%) regions. Nationally, there was a decline in TB prevalence from 0.18% in 2001 to 0.04% in 2009. However, prevalence increased back to 0.29% in 2014. Substantial spatial variation of TB prevalence was observed at a regional level, with a higher prevalence observed in the border regions, and at a local level within regions. The spatial distribution of TB prevalence was positively associated with population density.CONCLUSION:The results of this study showed that TB prevalence varied substantially at sub-national and local levels in Ethiopia. Spatial patterns were associated with population density. These results suggest that targeted interventions in high-risk areas may reduce the burden of TB in Ethiopia and additional data collection would be required to make further inferences on TB prevalence in areas that lack data.
Topic models are a useful and popular method to find latent topics of documents. However, the short and sparse texts in social media micro-blogs such as Twitter are challenging for the most commonly used Latent Dirichlet Allocation (LDA) topic model. We compare the performance of the standard LDA topic model with the Gibbs Sampler Dirichlet Multinomial Model (GSDMM) and the Gamma Poisson Mixture Model (GPM), which are specifically designed for sparse data. To compare the performance of the three models, we propose the simulation of pseudo-documents as a novel evaluation method. In a case study with short and sparse text, the models are evaluated on tweets filtered by keywords relating to the Covid-19 pandemic. We find that standard coherence scores that are often used for the evaluation of topic models perform poorly as an evaluation metric. The results of our simulation-based approach suggest that the GSDMM and GPM topic models may generate better topics than the standard LDA model.
A rapid response to global infectious disease outbreaks is crucial to protect public health. Ex ante information on the spatial probability distribution of early infections can guide governments to better target protection efforts. We propose a two-stage statistical approach to spatially map the ex ante importation risk of COVID-19 and its uncertainty across Indonesia based on a minimal set of routinely available input data related to the Indonesian flight network, traffic and population data, and geographical information. In a first step, we use a generalised additive model to predict the ex ante COVID-19 risk for 78 domestic Indonesian airports based on data from a global model on the disease spread and covariates associated with Indonesian airport network flight data prior to the global COVID-19 outbreak. In a second step, we apply a Bayesian geostatistical model to propagate the estimated COVID-19 risk from the airports to all of Indonesia using freely available spatial covariates including traffic density, population and two spatial distance metrics. The results of our analysis are illustrated using exceedance probability surface maps, which provide policy-relevant information accounting for the uncertainty of the estimates on the location of areas at risk and those that might require further data collection.
Instance segmentation is an important method for high-resolution remote sensing images (HRRSIs) analysis. Traditional instance segmentation algorithms are not suitable to analyze complex HRRSIs that exhibit: 1) various shapes and sizes of targets; 2) a large number of small targets; and 3) data with long tail distribution. Here we introduce DB-BlendMask, an efficient and accurate instance segmentation method that can accommodate complex HRRSIs. It is composed of size balance coefficient (SBC), class balance module (CBM), and decomposed attention blender module (DA-Blender module). SBC consists of a fair weight allocation strategy for positive samples in object detection. CBM combines classification obtained in object detection stage to guide the semantic feature extraction. Complementary to a traditional convolutional neural network (CNN) architecture, DA-Blender module has the ability to considerably compress space complexity of attention and merge attention with semantic feature to generate the instance mask. We compare the performance of DB-BlendMask with a benchmark Mask R-CNN on two typical datasets, iSAID, and ISPRS Postdam. We obtain an average detection precision of 39.2% on iSAID and 63.6% on ISPRS Postdam, which corresponds to an improvement of 2.5% and 2.7%, respectively, compared to the benchmark in a real-time scenario.
Models applied to geographic data face a trade-off between producing general results and capturing local variations due to spatial heterogeneity. Spatial modelling within carefully defined regions offers an intermediate position between global and local models. However, current spatial optimization approaches to delineate homo- geneous regions consider the similarity of attribute values, thus unable to identify regions with similar data generation processes described by geographical models. We propose a generalized regionalization framework, which optimizes region delineation corresponding to a model with region-specific parameters. Within this framework, we introduce three regionalization algorithms, namely automatic zoning procedure (AZP), K-Models, and Regional-K-Models. We adopt an objective function that jointly minimizes modelling errors and the complexity of the region scheme. Results from regression experiments indicate that the K-Models algorithm reconstructs the regions better than the baseline, according to Rand index and mutual information measures. Our suggested framework contributes to better capturing processes ex- hibiting spatial heterogeneity and may be applied to a wide range of modelling scenarios.
This study provides a comprehensive validation of Integrate Multi-SatellitE Retrievals of (IMERG) Global Precipitation Measurement (GPM) products in detecting extremes and drought over mainland China. The estimated values of extreme precipitation and drought are provided by three runs (early, late and final) of the latest IMERG products (V06) and 696 in-situ gauges over mainland China during 2008–2017. The results demonstrate that the three runs of IMERG V06 exhibit a relatively good performance in detecting the spatial patterns of extreme precipitation volume. The early and late runs present limited capability to capture extreme precipitation events, while the final run performs slightly better. Based on the extreme value theory, the three runs show better performances in estimating extreme precipitation within short return period (10-yr) compared to 50-yr and 100-yr return periods. All runs consistently underestimate extreme precipitation in all investigated return periods over eastern China. The standardized precipitation index (SPI) is used as a drought monitoring tool. The SPI of three runs of IMERG is well aligned with the in-situ data in southern and eastern China at both time and space scales, which demonstrates that the great potential of IMERG V06 in monitoring drought in these regions. However, the three runs of IMERG V06 products have inferior performances in monitoring drought over western China at four investigated timescales (1-, 3-, 6- and 12-month). This study highlights discrepancies in the capability of the IMERG products to estimate major extreme climatic variables and suggests potential avenues of research to improve the algorithm associated with the products. Meanwhile, the results of this study can benefit both developers and users of the new generation GPM IMERG products.
Plain Language Summary Clinopyroxene is a major mineral in Earth's upper mantle. Previous studies have attempted to discriminate between reactions modifying the mantle by plotting clinopyroxene major and trace element compositions in two‐dimensional (2‐D) diagrams. However, these 2‐D methods show poor accuracy when applied to global datasets. Therefore, we suggest a machine learning approach to evaluate clinopyroxene compositional data in higher dimensions. Our results demonstrate that machine learning can significantly improve the accuracy of clinopyroxene compositional predictions over classical methods utilizing elemental ratios. Furthermore, the application of our algorithm to a global clinopyroxene dataset suggests that mantle metasomatism is globally widespread.
As malaria incidence decreases and more countries move towards elimination, maps of malaria risk in low-prevalence areas are increasingly needed. For low-burden areas, disaggregation regression models have been developed to estimate risk at high spatial resolution from routine surveillance reports aggregated by administrative unit polygons. However, in areas with both routine surveillance data and prevalence surveys, models that make use of the spatial information from prevalence point-surveys might make more accurate predictions. Using case studies in Indonesia, Senegal and Madagascar, we compare the out-of-sample mean absolute error for two methods for incorporating point-level, spatial information into disaggregation regression models. The first simply fits a binomial-likelihood, logit-link, Gaussian random field to prevalence point-surveys to create a new covariate. The second is a multi-likelihood model that is fitted jointly to prevalence point-surveys and polygon incidence data. We find that in most cases there is no difference in mean absolute error between models. In only one case, did the new models perform the best. More generally, our results demonstrate that combining these types of data has the potential to reduce absolute error in estimates of malaria incidence but that simpler baseline models should always be fitted as a benchmark.
The Arctic warming rate is triple the global average, which is partially caused by surface albedo feedback (SAF). Understanding the varying pattern of SAF and the mechanisms is therefore critical for predicting future Arctic climate under anthropogenic warming. To date, however, how the spatial pattern of seasonal SAF is influenced by various land surface factors remains unclear. Here, we aim to quantify the strengths of seasonal SAF across the Arctic and to attribute its spatial heterogeneity to the dynamics of vegetation, snow and soil as well as their interactions. The results show a large positive SAF above −5% K−1 across Baffin Island in January and eastern Yakutia in June, while a large negative SAF beyond 5% K−1 is observed in Canada, Chukotka and low latitudes of Greenland in January and Nunavut, Baffin Island and Krasnoyarsk Krai in July. Overall, a great spatial heterogeneity of Arctic land warming induced by positive SAF is found with a coefficient of variation (CV) larger than 61.5%, and the largest spatial difference is detected in wintertime with a CV > 643.9%. Based on the optimal parameter-based geographic detector model, the impacts of snow cover fraction (SCF), land cover type (LC), normalized difference vegetation index (NDVI), soil water content (SW), soil substrate chemistry (SC) and soil type (ST) on the spatial pattern of positive SAF are quantified. The rank of determinant power is SCF > LC > NDVI > SW > SC > ST, which indicates that the spatial patterns of snow cover, land cover and vegetation coverage dominate the spatial heterogeneity of positive SAF in the Arctic. The interactions between SCF, LC and SW exert further influences on the spatial pattern of positive SAF in March, June and July. This work could provide a deeper understanding of how various land factors contribute to the spatial heterogeneity of Arctic land warming at the annual cycle.
As the COVID-19 pandemic continues to threaten various regions around the world, obtaining accurate and reliable COVID-19 data is crucial for governments and local communities aiming at rigorously assessing the extent and magnitude of the virus spread and deploying efficient interventions. Using data reported between January and February 2020 in China, we compared counts of COVID-19 from near-real-time spatially disaggregated data (city level) with fine-spatial scale predictions from a Bayesian downscaling regression model applied to a reference province-level data set. The results highlight discrepancies in the counts of coronavirus-infected cases at the district level and identify districts that may require further investigation.
Research on the effects of terrorism mostly focuses on the coercive effects of violence on the macrolevel, while other effects like provocation, particularly on the microlevel, do not receive the same attention. In this article, we seek to address previous omissions. We argue that terrorism can provoke ordinary people into a violent reaction. By reducing perceived security and creating a desire for revenge terrorism may lead civilians to attack uninvolved members of the terrorists’ constituency. Using geo-referenced data on terrorism (Global Terrorism Database) and violent riots (Social Conflict Analysis Database), we assess with a matched wake analysis if the treatment of terrorist violence against civilians causes an increase in violent behavior. The results of our analyses show that terrorism significantly increases violent riots. We thus conclude that terrorism can not only provoke governments but also civilians into an overreaction.
Several thousand people die every year worldwide because of terrorist attacks perpetrated by non-state actors. In this context, reliable and accurate short-term predictions of non-state terrorism at the local level are key for policy makers to target preventative measures. Using only publicly available data, we show that predictive models that include structural and procedural predictors can accurately predict the occurrence of non-state terrorism locally and a week ahead in regions affected by a relatively high prevalence of terrorism. In these regions, theoretically informed models systematically outperform models using predictors built on past terrorist events only. We further identify and interpret the local effects of major global and regional terrorism drivers. Our study demonstrates the potential of theoretically informed models to predict and explain complex forms of political violence at policy-relevant scales.
As malaria incidence decreases and more countries move towards elimination, maps of malaria risk in low-prevalence areas are increasingly needed. For low-burden areas, disaggregation regression models have been developed to estimate risk at high spatial resolution from routine surveillance reports aggregated by administrative unit polygons. However, in areas with both routine surveillance data and prevalence surveys, models that make use of the spatial information from prevalence point-surveys might make more accurate predictions. Using case studies in Indonesia, Senegal and Madagascar, we compare the out-of-sample mean absolute error for two methods for incorporating point-level, spatial information into disaggregation regression models. The first simply fits a binomial-likelihood, logit-link, Gaussian random field to prevalence point-surveys to create a new covariate. The second is a multi-likelihood model that is fitted jointly to prevalence point-surveys and polygon incidence data. We find that in most cases there is no difference in mean absolute error between models. In only one case, did the new models perform the best. More generally, our results demonstrate that combining these types of data has the potential to reduce absolute error in estimates of malaria incidence but that simpler baseline models should always be fitted as a benchmark.