Drug overdose deaths, including from opioids, remain a significant public health threat to the United States (US). To abate the harms of opioid misuse, understanding its prevalence at the local level is crucial for stakeholders in communities to develop response strategies that effectively use limited resources. Although there exist several state-specific studies that provide county-level prevalence estimates, such estimates are not widely available across the country, as the datasets used in these studies are not always readily available in other states, which, therefore, has limited the wider applications of existing models. To fill this gap, we propose a Bayesian multi-state data integration approach that fully utilizes publicly available data sources to estimate county-level opioid misuse prevalence for all counties in the US. The hierarchical structure jointly models opioid misuse prevalence and overdose death outcomes, leverages existing county-level prevalence estimates in limited states and state-level estimates from national surveys, and accounts for heterogeneity across counties and states with counties' covariates and mixed effects. Furthermore, our parsimonious and generalizable modeling framework employs horseshoe+ prior to flexibly shrink coefficients and prevent overfitting, ensuring adaptability as new county-level prevalence data in additional states become available. Using real-world data, our model shows high estimation accuracy through cross-validation and provides nationwide county-level estimates of opioid misuse for the first time.
This Element introduces the basics of Bayesian regression modeling using modern computational tools. This Element only assumes that the reader has taken a basic statistics course and has seen Bayesian inference at the introductory level of Gill and Bao (2024). Some matrix algebra knowledge is assumed but the authors walk carefully through the necessary structures at the start of this Element. At the end of the process readers will fully understand how Bayesian regression models are developed and estimated, including linear and nonlinear versions. The sections cover theoretical principles and real-world applications in order to provide motivation and intuition. Because Bayesian methods are intricately tied to software, code in R and Python is provided throughout.
Key populations at high risk of HIV infection are critical for understanding and monitoring HIV epidemics, but global estimation is hampered by sparse, uneven data. We analyse data from 199 countries for female sex workers (FSW), men who have sex with men (MSM), and people who inject drugs (PWID) over 2011-2021, and introduce a cross-population temporal hierarchical model that borrows strength across countries, years, and populations. The model combines region- and population-specific means with country random effects, temporal dependence, and cross-population correlations in a Gaussian Markov random-field formulation on the log-prevalence scale. In fivefold cross-validation, the approach outperforms a regional-median baseline and reduced variants (65% reduction in cross-validation mean squared error) with a well-calibrated posterior predictive coverage (93%). We map the 2021 prevalence and quantify the change between 2011 and 2021 using posterior prevalence ratios to identify countries with substantial increases or decreases. The framework yields globally comparable and uncertainty-quantified country-by-year prevalence estimates, enhancing evidence for resource allocation and targeted interventions for marginalized populations where routine data are limited.
This paper focuses on minimizing a smooth function combined with a nonsmooth regularization term on a compact Riemannian submanifold embedded in the Euclidean space under a decentralized setting. Typically, there are two types of approaches at present for tackling such composite optimization problems. The first, subgradient-based approaches, rely on subgradient information of the objective function to update variables, achieving an iteration complexity of $\mathcal{O}(\epsilon^{-4}\log^2(\epsilon^{-2}))$. The second, smoothing approaches, involve constructing a smooth approximation of the nonsmooth regularization term, resulting in an iteration complexity of $\mathcal{O}(\epsilon^{-4})$. This paper proposes a proximal gradient type algorithm that fully exploits the composite structure. The global convergence to a stationary point is established with a significantly improved iteration complexity of $\mathcal{O}(\epsilon^{-2})$. To validate the effectiveness and efficiency of our proposed method, we present numerical results in real-world applications, showcasing its superior performance.
Accurate assessment of adverse event (AE) incidence is critical in clinical research for drug safety. While meta-analysis serves as an essential tool to comprehensively synthesize the evidence across multiple studies, incomplete AE reporting in clinical trials remains a persistent challenge. In particular, AEs occurring below study-specific reporting thresholds are often omitted from publications, leading to left-censored data. Failure to account for these censored AE counts can result in biased AE incidence estimates. We present an R Shiny application that implements a Bayesian meta-analysis model specifically designed to incorporate censored AE data into the estimation process. This interactive tool provides a user-friendly interface for researchers to conduct AE meta-analyses and estimate the AE incidence probability using an unbiased approach. It also enables direct comparisons between models that either incorporate or ignore censoring, highlighting the biases introduced by conventional approaches. This tutorial demonstrates the Shiny application’s functionality through an illustrative example on meta-analysis of PD-1/PD-L1 inhibitor safety and highlights the importance of this tool in improving AE risk assessment. Ultimately, the new Shiny app facilitates more accurate and transparent drug safety evaluations. The Shiny-MAGEC app is available at: https://zihanzhou98.shinyapps.io/Shiny-MAGEC/ .
Background:Patients with rare cancers face substantial challenges due to limited evidence-based treatment options, resulting from sparse clinical trials. Advances in large language models (LLMs) and recommendation algorithms offer new opportunities to utilize all clinical trial information to improve clinical decisions. Methods:We used LLM to systematically extract and standardize more than 100,000 cancer trials from ClinicalTrials.gov. Each trial was annotated using a customized scoring system reflecting cancer-treatment interactions based on clinical outcomes and trial attributes. Using this structured data set, we implemented three state-of-the-art collaborative filtering algorithms to recommend potentially effective treatments across different cancer types. Results:The LLM-driven data extraction process successfully generated a comprehensive and rigorously curated database from fragmented clinical trial information, covering 78 cancer types and 5,315 distinct interventions. Recommendation models demonstrated high predictive accuracy (cross-validated RMSE: 0.49-0.62) and identified clinically meaningful new treatments for melanoma, independently validated by oncology experts. Conclusions:Our study establishes a proof of concept demonstrating that the combination of LLMs with sophisticated recommendation algorithms can systematically identify novel and clinically plausible cancer treatments. This integrated approach may accelerate the identification of effective therapies for rare cancers, ultimately improving patient outcomes by generating evidence-based treatment recommendations where traditional data sources remain limited.
Dynamic models have been successfully used in producing estimates of HIV epidemics at the national level due to their epidemiological nature and their ability to estimate prevalence, incidence, and mortality rates simultaneously. Recently, HIV interventions and policies have required more information at sub-national levels to support local planning, decision-making and resource allocation. Unfortunately, many areas lack sufficient data for deriving stable and reliable results, and this is a critical technical barrier to more stratified estimates. One solution is to borrow information from other areas within the same country. However, directly assuming hierarchical structures within the HIV dynamic models is complicated and computationally time-consuming. In this article, we propose a simple and innovative way to incorporate hierarchical information into the dynamical systems by using auxiliary data. The proposed method efficiently uses information from multiple areas within each country without increasing the computational burden. As a result, the new model improves predictive ability and uncertainty assessment.
Purpose of ReviewBig Data Science can be used to pragmatically guide the allocation of resources within the context of national HIV programs and inform priorities for intervention. In this review, we discuss the importance of grounding Big Data Science in the principles of equity and social justice to optimize the efficiency and effectiveness of the global HIV response.Recent FindingsSocial, ethical, and legal considerations of Big Data Science have been identified in the context of HIV research. However, efforts to mitigate these challenges have been limited. Consequences include disciplinary silos within the field of HIV, a lack of meaningful engagement and ownership with and by communities, and potential misinterpretation or misappropriation of analyses that could further exacerbate health inequities.SummaryBig Data Science can support the HIV response by helping to identify gaps in previously undiscovered or understudied pathways to HIV acquisition and onward transmission, including the consequences for health outcomes and associated comorbidities. However, in the absence of a guiding framework for equity, alongside meaningful collaboration with communities through balanced partnerships, a reliance on big data could continue to reinforce inequities within and across marginalized populations.
We introduce a novel procedure for obtaining cross-validated predictive estimates for Bayesian hierarchical regression models (BHRMs). Bayesian hierarchical models are popular for their ability to model complex dependence structures and provide probabilistic uncertainty estimates, but can be computationally expensive to run. Cross-validation (CV) is therefore not a common practice to evaluate the predictive performance of BHRMs. Our method circumvents the need to re-run computationally costly estimation methods for each cross-validation fold and makes CV more feasible for large BHRMs. By conditioning on the variance-covariance parameters, we shift the CV problem from probability-based sampling to a simple and familiar optimization problem. In many cases, this produces estimates which are equivalent to full CV. We provide theoretical results and demonstrate its efficacy on publicly available data and in simulations.
Estimating new HIV infections is significant yet challenging due to the difficulty in distinguishing between recent and long-term infections. We demonstrate that HIV recency status (recent v.s. long-term) could be determined from the combination of self-report testing history and biomarkers, which are increasingly available in bio-behavioral surveys. HIV recency status is partially observed, given the self-report testing history. For example, people who tested positive for HIV over one year ago should have a long-term infection. Based on the nationally representative samples collected by the Population-based HIV Impact Assessment (PHIA) Project, we propose a likelihood-based probabilistic model for HIV recency classification. The model incorporates both labeled and unlabeled data and integrates the mechanism of how HIV recency status depends on biomarkers and the mechanism of how HIV recency status, together with the self-report time of the most recent HIV test, impacts the test results, via a set of logistic regression models. We compare our method to logistic regression and the binary classification tree (current practice) on Malawi, Zimbabwe, and Zambia PHIA data, as well as on simulated data. Our model obtains more efficient and less biased parameter estimates and is relatively robust to potential reporting error and model misspecification.
Ending the HIV/AIDS pandemic is among the sustainable development goals for the next decade. To overcome the problem caused by the imbalances between the need for care and the limited resources, we shall improve our understanding of the local HIV epidemics, especially for key populations at high risk of HIV infection. However, HIV prevalence rates for key populations have been difficult to estimate because their HIV surveillance data are very scarce. This paper develops a multivariate spatial model for predicting unknown HIV prevalence rates among key populations. The proposed multivariate conditional auto-regressive model efficiently pools information from neighbouring locations and correlated populations. As the real data analysis illustrates, it provides more accurate predictions than independently fitting the sub-epidemic for each key population. Furthermore, we investigate how different pieces of surveillance data contribute to the prediction and offer practical suggestions for epidemic data collection.
Myeloproliferative neoplasms (MPN) including polycythemia vera (PV), essential thrombocythemia (ET) and myelofibrosis (MF) are clonal myeloid malignancies derived from mutated hematopoietic stem cells. The JAK2V617F mutation was found in ~95% cases of PV and 50-60% patients with ET and MF. Mutations in MPL and CALR were also detected in ET and MF. The JAK inhibitors, Ruxolitinib and Fedratinib, can reduce splenomegaly and alleviate constitutional symptoms but they are not sufficient to induce remission of MPN. So, there is an unmet need to identify new therapeutic targets and therapies for MPN/MF. The non-receptor protein tyrosine phosphatase 11 (PTPN11), which positively regulates hematopoietic cell signaling, is constitutively hyperphosphorylated in MPN patient cells and in hematopoietic cell lines expressing MPN driver mutants. So, we hypothesize that PTPN11 plays an important role in MPNs. To examine the role of PTPN11 in JAK2V617F-induced MPN, we generated inducible PTPN11-deficient heterozygous JAK2V617F mice by crossing PTPN11 floxed mouse with Mx1Cre and JAK2V617F knock-in mice. While expression of heterozygous JAK2V617F (JAK2 VF/+) induced a PV-like MPN characterized by increased WBC, neutrophil, RBC, hemoglobin and platelet counts in the peripheral blood and enlargement of spleen size, deletion of PTPN11 normalized the blood counts and spleen size in JAK2 VF/+ mice. PTPN11 deletion significantly reduced the expansion of hematopoietic stem/progenitors, erythroid, megakaryocytic, and granulocytic cells in the bone marrow (BM) and spleens of JAK2 VF/+ mice. Furthermore, deletion of PTPN11 abrogated Epo-independent CFU-colonies, a hallmark feature of PV disease, in the BM and spleens of JAK2 VF/+ mice. We also investigated the role of PTPN11 in the development of MF using homozygous JAK2V617F and MPLW515L mouse models. We found that deletion of PTPN11 significantly reduced WBC, neutrophil and platelet counts, normalized spleen size and abrogated BM fibrosis in JAK2V617F and MPLW515L mouse models. These results suggest that PTPN11 plays an important role in the development and progression of PV and MF. Data from our PTPN11 genetic deletion study led us to investigate the effects of PTPN11 inhibition against MPN/MF. We tested the efficacy of an allosteric PTPN11 inhibitor SHP099 in MPN hematopoietic cell lines and mouse models of MPN/MF. We observed that treatment of SHP099 significantly reduced proliferation of hematopoietic cells expressing JAK2V617F and MPLW515L. SHP099 treatment also significantly inhibited hematopoietic progenitor colony growth in PV and MF patient CD34+ cells. Next, we tested the in vivo efficacy of SHP099 against PV using heterozygous JAK2V617F mice. Treatment of SHP099 significantly inhibited the increase in WBC, neutrophil and RBC counts and markedly reduced splenomegaly in JAK2 VF/+ mice. SHP099 treatment also significantly reduced hematopoietic stem/progenitors and myeloid precursors in the BM and spleens of JAK2 VF/+ mice. Furthermore, SHP099 treatment abrogated Epo-independent CFU-E colony formation in the BM and spleens of JAK2 VF/+ mice. We also tested the efficacy of SHP099 alone or in combination with ruxolitinib against myelofibrosis using homozygous JAK2V617F (JAK2 VF/VF) and MPLW515L mouse models. SHP099 treatment alone significantly reduced WBC and neutrophil counts and markedly reduced splenomegaly in JAK2 VF/VF and MPLW515L mice. Combined treatment of SHP099 and ruxolitinib resulted in greater reduction of blood cell counts and spleen size in JAK2 VF/VF and MPLW515Lmice. Histopathologic analyses showed that SHP099 treatment significantly reduced BM fibrosis in JAK2 VF/VF and MPLW515L mice while combined treatment of SHP099and ruxolitinib resulted in greater inhibition of BM fibrosis in these mice. To gain insights into the mechanism of inhibition of MPN/MF by PTPN11 deletion or inhibition, we performed RNA-seq analysis on purified LSK cells from JAK2 VF/+ and PTPN11deleted-JAK2 VF/+ mice as well as from JAK2 VF/VF mice treated with ruxolitinib and SHP099/ruxolitinib. RNA-seq data analysis revealed that genes related to MYC targets, ribosome biogenesis and translation were significantly down-regulated by PTPN11 deletion or SHP099/ruxolitinib combination treatment. Overall, our results suggest that inhibition of PTPN11 alone or in combination with JAK2 inhibition might be useful for treatment of PV and MF.
With the emergence of many knowledge-based systems worldwide, there have been more and more applications using different kinds of data and solving significant daily problems. Among that, the issues of missing data in such systems have become more popular, especially in data-driven areas. Other research on the imputation problem has dealt with partial and missing data. This study aims to investigate the imputation techniques for sparse data using the Singular Value Decomposition technique, namely SVDI. We explore the application of the SVDI framework for image classification and text classification tasks that involve sparse data. The experimental results show that the proposed SVDI method improves the speed and accuracy of the imputation process when compared to the PCAI method. We aim to publish our codes related to the SVDI later for the relevant research community.
Accurate HIV incidence estimation based on individual recent infection status (recent vs long-term infection) is important for monitoring the epidemic, targeting interventions to those at greatest risk of new infection, and evaluating existing programs of prevention and treatment. Starting from 2015, the Population-based HIV Impact Assessment (PHIA) individual-level surveys are implemented in the most-affected countries in sub-Saharan Africa. PHIA is a nationally-representative HIV-focused survey that combines household visits with key questions and cutting-edge technologies such as biomarker tests for HIV antibody and HIV viral load which offer the unique opportunity of distinguishing between recent infection and long-term infection, and providing relevant HIV information by age, gender, and location. In this article, we propose a semi-supervised logistic regression model for estimating individual level HIV recency status. It incorporates information from multiple data sources -- the PHIA survey where the true HIV recency status is unknown, and the cohort studies provided in the literature where the relationship between HIV recency status and the covariates are presented in the form of a contingency table. It also utilizes the national level HIV incidence estimates from the epidemiology model. Applying the proposed model to Malawi PHIA data, we demonstrate that our approach is more accurate for the individual level estimation and more appropriate for estimating HIV recency rates at aggregated levels than the current practice -- the binary classification tree (BCT).
Female sex workers (FSW) are affected by individual, network, and structural risks, making them vulnerable to poor health and well-being. HIV prevention strategies and local community-based programs can rely on estimates of the number of FSW to plan and implement differentiated HIV prevention and treatment services. However, there are limited systematic assessments of the number of FSW in countries across sub-Saharan Africa to facilitate the identification of prevention and treatment gaps. Here we provide estimated population sizes of FSW and the corresponding uncertainties for almost all sub-national areas in sub-Saharan Africa. We first performed a literature review of FSW size estimates and then developed a Bayesian hierarchical model to synthesize these size estimates, resolving competing size estimates in the same area and producing estimates in areas without any data. We estimated that there are 2.5 million (95% uncertainty interval 1.9 to 3.1) FSW aged 15 to 49 in sub-Saharan Africa. This represents a proportion as percent of all women of childbearing age of 1.1% (95% uncertainty interval 0.8 to 1.3%). The analyses further revealed substantial differences between the proportions of FSW among adult females at the sub-national level and studied the relationship between these heterogeneities and many predictors. Ultimately, achieving the vision of no new HIV infections by 2030 necessitates dramatic improvements in our delivery of evidence-based services for sex workers across sub-Saharan Africa.
As the global HIV pandemic enters its fifth decade, increasing numbers of countries use routine HIV testing among pregnant women to monitor their epidemics, allowing governments to look into the epidemics at a finer scale, for example, at subnational levels. Currently, the epidemic model that describes the dynamics of the spread of HIV consists of a set of differential equations and is applied independently to each subnational area. However, the availability of the data varies widely which leads to biased and unreliable estimates for areas with very few data points. We propose to overcome this issue by introducing dependence in the parameters across areas. The proposed method better reconstructs the epidemic trajectories than the independent model as shown in multiple countries in Sub-Saharan Africa. We also offer an approximate method for parameter estimation that is much less computationally burdensome than direct parameter estimation. Compared to direct parameter estimation from the dependent model, the approximate method provides competitive parameter estimation in simulations and the application of HIV subepidemic estimation.
Air pollution continues to be a major environmental concern in China. The wind‐driven transmission poses difficulties in understanding the air pollution patterns at the local level. The main objective of this study is to offer a straightforward approach for investigating the temporal trends and meteorological effects on the air pollutant concentrations during the generation process without being confounded by the complex wind‐driven transmission effect. We focus on the hourly data of the three most common air pollutants: PM2.5, NO, and CO under air stagnation in Beijing, China, during 2014–2017. We find that the local pollution levels under air stagnation in Beijing have decreased over the years; winter is the severest month of the year; Sunday is the clearest day of the week. Our model also interpolates the air pollutant concentrations at sites without monitoring stations and provides a map of air pollution concentrations under air stagnation. The results could be used to identify locations where air pollutants easily accumulate.
Aggregated relational data (ARD), formed from "How many X's do you know?" questions, is a powerful tool for learning important network characteristics with incomplete network data. Compared to traditional survey methods, ARD is attractive as it does not require a sample from the target population and does not ask respondents to self-reveal their own status. This is helpful for studying hard-to-reach populations like female sex workers who may be hesitant to reveal their status. From December 2008 to February 2009, the Kiev International Institute of Sociology (KIIS) collected ARD from 10,866 respondents to estimate the size of HIV-related groups in Ukraine. To analyze this data, we propose a new ARD model which incorporates respondent and group covariates in a regression framework and includes a bias term that is correlated between groups. We also introduce a new scaling procedure utilizing the correlation structure to further reduce biases. The resulting size estimates of those most-at-risk of HIV infection can improve the HIV response efficiency in Ukraine. Additionally, the proposed model allows us to better understand two network features without the full network data: 1. What characteristics affect who respondents know, and 2. How is knowing someone from one group related to knowing people from other groups. These features can allow researchers to better recruit marginalized individuals into the prevention and treatment programs. Our proposed model and several existing NSUM models are implemented in the networkscaleup R package.
Model development often takes data structure, subject matter considerations, model assumptions, and goodness of fit into consideration. To diagnose issues with any of these factors, it can be helpful to understand regression model estimates at a more granular level. We propose a new method for decomposing point estimates from a regression model via weights placed on data clusters. The weights are informed only by the model specification and data availability and thus can be used to explicitly link the effects of data imbalance and model assumptions to actual model estimates. The weight matrix has been understood in linear models as the hat matrix in the existing literature. We extend it to Bayesian hierarchical regression models that incorporate prior information and complicated dependence structures through the covariance among random effects. We show that the model weights, which we call borrowing factors, generalize shrinkage and information borrowing to all regression models. In contrast, the focus of the hat matrix has been mainly on the diagonal elements indicating the amount of leverage. We also provide metrics that summarize the borrowing factors and are practically useful. We present the theoretical properties of the borrowing factors and associated metrics and demonstrate their usage in two examples. By explicitly quantifying borrowing and shrinkage, researchers can better incorporate domain knowledge and evaluate model performance and the impacts of data properties such as data imbalance or influential points.