Large language models (LLMs) play an increasingly important role in finan- cial markets analysis by capturing signals from complex and heterogeneous textual data sources, such as tweets, news articles, reports, and microblogs. However, their performance is dependent on large computational resources and proprietary datasets, which are costly, restricted, and therefore inacces- sible to many researchers and practitioners. To reflect realistic situations we investigate the ability of lightweight open-source LLMs - smaller and publicly available models designed to operate with limited computational resources - to generalize sentiment understanding from financial datasets of varying sizes, sources, formats, and languages. We compare the benchmark finance natural language processing (NLP) model, FinBERT, and three open-source lightweight LLMs, DeepSeek-LLM 7B, Llama3 8B Instruct, and Qwen3 8B on five publicly available datasets: FinancialPhraseBank, Financial Question Answering, Gold News Sentiment, Twitter Sentiment and Chinese Finance Sentiment. We find that LLMs, specially Qwen3 8B and Llama3 8B, perform best in most scenarios, even from using only 5
BACKGROUND:Worldwide, tuberculosis (TB) remains the leading cause of death from infectious diseases. Africa is the second most-affected region, accounting for a quarter of the global TB burden, but there is limited evidence whether there is subnational variation of TB prevalence across the continent. Therefore, this study aimed to estimate sub-national and local TB prevalence across Africa. METHODS:We compiled geolocated data from 50 population-based surveys across 14 African countries. A total of 212 data points were identified and linked to covariates assembled from publicly available sources. Bayesian geostatistical modelling was used to predict TB prevalence across Africa, and results were aggregated to estimate number of TB cases at national and subnational levels. RESULTS:Here we estimate 1.28 million TB cases (95% uncertainty interval [UI] 0.14-4.87) across 14 countries, with marked spatial variations. The highest cases are estimated in Nigeria (460,247 95% UI 7954-1,783,106), and Mozambique (120,622 95%UI 20,027-321,177) while the lowest in Guinea-Bissau (1952 95%UI 154-7365) and Rwanda (2207 95% UI 1050-9225). National TB prevalence range from 0.25 to 7.32 per 1000 with significant variation at higher spatial resolution. Temperature (°C) (OR = 1.27; 95% CrI: 1.20-1.35), precipitation (mm) (OR = 1.34; 95% CrI: 1.26-1.40), and access to city (minute) (OR = 1.21; 95% CrI: 1.14-1.25) are positively associated with TB prevalence, while altitude (m) (OR = 0.83; 95% CrI: 0.78-0.87) is negatively associated. CONCLUSIONS:We find substantial variations in TB prevalence at national, sub-national, and local levels in Africa. These considerable spatial variations suggest the need for geographically targeted interventions to control TB in Africa.
In the current Noisy Intermediate-Scale Quantum era, quantum circuit analysis is an essential technique for designing high-performance quantum programs. Current analysis methods exhibit either accuracy limitations or high computational complexity for obtaining precise results. To reduce this tradeoff, we propose QuCT, a unified framework for extracting, analyzing, and optimizing quantum circuits. The main innovation of QuCT is to vectorize each gate with each element, quantitatively describing the degree of the interaction with neighboring gates. Extending from the vectorization model, we propose two representative downstream models for fidelity prediction and unitary decomposition. The fidelity prediction model performs a linear transformation on all gate vectors and aggregates the results to estimate the overall circuit fidelity. By identifying critical weights in the transformation matrix, we propose two optimizations to improve the circuit fidelity. In the unitary decomposition model, we significantly reduce the search space by bridging the gap between unitary and circuit via gate vectors. Experiments show that QuCT improves the accuracy of fidelity prediction by 4.2x on 5-qubit and 18-qubit quantum devices and achieves 2.5x fidelity improvement compared to existing quantum compilers [19, 55]. In unitary decomposition, QuCT achieves 46.3x speedup for 5-qubit unitary and more than hundreds of speedup for 8-qubit unitary, compared to the state-of-the-art method [87].
Propositional satisfiability problem (SAT) is represented in a conjunctive normal form with multiple clauses, which is an important non-deterministic polynomial-time (NP) complete problem that plays a major role in various applications including artificial intelligence, graph colouring, and circuit analysis. Quantum annealing (QA) is a promising methodology for solving complex SAT problems by exploiting the parallelism of quantum entanglement, where the SAT variables are embedded to the qubits. However, the long embedding time fundamentally limits existing QA-based methods, leading to inefficient hardware implementation and poor scalability.In this paper, we propose HyQSAT, a hybrid approach that integrates QA with the classical Conflict-Driven Clause Learning (CDCL) algorithm to enable end-to-end acceleration for solving SAT problems. Instead of embedding all clauses to QA hardware, we quantitatively estimate the conflict frequency of clauses and apply breadth-first traversal to choose their embedding order. We also consider the hardware topology to maximize the utilization of physical qubits in embedding to QA hardware. Besides, we adjust the embedding coefficients to improve the computation accuracy under qubit noise. Finally, we present how to interpret the satisfaction probability based on QA energy distribution and use this information to guide the CDCL search. Our experiments demonstrate that HyQSAT can effectively support larger-scale SAT problems that are beyond the capability of existing QA approaches, achieve up to 12.62X end-to-end speedup using D-Wave 2000Q compared to the classic CDCL algorithm on Intel E5 CPU, and considerably reduce the QA embedding time from 17.2s to 15.7µs compared to the D-Wave Minorminer algorithm [11].
This study evaluated the performance of the early, late and final runs of IMERG version 06 precipitation products at various spatial and temporal scales in China from 2008 to 2017, against observations from 696 rain gauges. The results suggest that the three IMERG products can well reproduce the spatial patterns of precipitation, but exhibit a gradual decrease in the accuracy from the southeast to the northwest of China. Overall, the three runs show better performances in the eastern humid basins than the western arid basins. Compared to the early and late runs, the final run shows an improvement in the performance of precipitation estimation in terms of correlation coefficient, Kling–Gupta Efficiency and root mean square error at both daily and monthly scales. The three runs show similar daily precipitation detection capability over China. The biases of the three runs show a significantly positive (p < 0.01) correlation with elevation, with higher accuracy observed with an increase in elevation. However, the categorical metrics exhibit low levels of dependency on elevation, except for the probability of detection. Over China and major river basins, the three products underestimate the frequency of no/tiny rain events (P < 0.1 mm/day) but overestimate the frequency of light rain events (0.1 ≤ P < 10 mm/day). The three products converge with ground-based observation with regard to the frequency of rainstorm (P ≥ 50 mm/day) in the southern part of China. The revealed uncertainties associated with the IMERG products suggests that sustaining efforts are needed to improve their retrieval algorithms in the future.
Abstract The increasingly growing availability of large disaggregated datasets on terrorism along with advances made in the fields of statistics and data science has allowed scholars to better understand the patterns of non‐state terrorism at fine spatial and temporal scales. While acknowledging important limitations that stem from the analysis of data on terrorism, we present a Bayesian hierarchical modeling approach that contributes to explaining fine‐scale patterns of the lethality of worldwide non‐state terrorism.
We present a Bayesian geospatial modeling framework developed for the synthesis of point prevalence and health facility catchment data of mixed types: Plasmodium parasite prevalence and malaria febrile incidence. Since the clinical case definition for health facility record keeping is less strict than that used in cohort studies used to construct previous parasite prevalence to clinical incidence relationships (the latter usually incorporating a parasite density threshold to remove background fevers accompanied by coincidence asymptomatic parasite infections) our model learns a smooth prevalence-to-incidence conversion during posterior sampling. Also jointly fitted are a catchment model based on the relative travel times between each pixel location and its nearby health facilities, as well as a distribution regression-based covariance structure for explaining the residual errors at health facility level based on the similarity between the ‘bags’ of covariate values in their respective catchments.
Approximately 150 triatomine species are suspected to be infected with the Chagas parasite, Trypanosoma cruzi, but they differ in the risk they pose to human populations. The largest risk comes from species that have a domestic life cycle and these species have been targeted by indoor residual spraying campaigns, which have been successful in many locations. It is now important to consider residual transmission that may be linked to persistent populations of dominant vectors, or to secondary or minor vectors. The aim of this project was to define the geographical distributions of the community of triatomine species across the Chagas endemic region. Presence-only data with over 12, 000 observations of triatomine vectors were extracted from a public database and target-group background data were generated to account for sampling bias in the presence data. Geostatistical regression was then applied to estimate species distributions and fine-scale distribution maps were generated for thirty triatomine vector species including those found within one or two countries and species that are more widely distributed from northern Argentina to Guatemala, Bolivia to southern Mexico, and Mexico to the southern United States of America. The results for Rhodnius pictipes, Panstrongylus geniculatus, Triatoma dimidiata, Triatoma gerstaeckeri, and Triatoma infestans are presented in detail, including model predictions and uncertainty in these predictions, and the model validation results for each of the 30 species are presented in full. The predictive maps for all species are made publicly available so that they can be used to assess the communities of vectors present within different regions of the endemic zone. The maps are presented alongside key indicators for the capacity of each species to transmit T. cruzi to humans. These indicators include infection prevalence, evidence for human blood meals, and colonisation or invasion of homes. A summary of the published evidence for these indicators shows that the majority of the 30 species mapped by this study have the potential to transmit T. cruzi to humans.
Can Bayesian models reveal the underlying processes that drive the lethality of non‐state terrorism at a local level? Andre Python, Janine B. Illian, Charlotte M. Jones‐Todd and Marta Blangiardo investigate.
Maps of infection risk are a vital tool for the elimination of malaria. Routine surveillance data of malaria case counts, often aggregated over administrative regions, is becoming more widely available and can better measure low malaria risk than prevalence surveys. However, aggregation of case counts over large, heterogeneous areas means that these data are often underpowered for learning relationships between the environment and malaria risk. A model that combines point surveys and aggregated surveillance data could have the benefits of both but must be able to account for the fact that these two data types are different malariometric units. Here, we train multiple machine learning models on point surveys and then combine the predictions from these with a geostatistical disaggregation model that uses routine surveillance data. We find that, in tests using data from Colombia and Madagascar, using a disaggregation regression model to combine predictions from machine learning models trained on point surveys improves model accuracy relative to using the environmental covariates directly.
SummaryTerrorism persists as a worldwide threat, as exemplified by the on-going lethal attacks perpetrated by Islamic State in Iraq and Syria, Al Qaeda in Yemen and Boko Haram in Nigeria. In response, states deploy various counterterrorism policies, the costs of which could be reduced through efficient preventive measures. Statistical models that can account for complex spatiotemporal dependences have not yet been applied, despite their potential for providing guidance to explain and prevent terrorism. To address this shortcoming, we employ hierarchical models in a Bayesian context, where the spatial random field is represented by a stochastic partial differential equation. Our main findings suggest that lethal terrorist attacks tend to generate more deaths in ethnically polarized areas and in locations within democratic countries. Furthermore, the number of lethal attacks increases close to large cities and in locations with higher levels of population density and human activity.