The main infections that cause cancer in the European Union (EU) are Helicobacter pylori (H. pylori), human papillomavirus (HPV), hepatitis B virus (HBV), hepatitis C virus (HCV) and human immunodeficiency virus (HIV). Altogether, in 2022, these infections accounted for ~ 5% of all cancers in the EU, mainly of the stomach, cervix uteri and liver. The largest burden of infection‐caused cancers was found in the south of the EU and near the eastern border. Substantial progress in the efficacy of interventions against these infections has been made since the release of the 4th edition of the European Code Against Cancer in 2015. Cancers due to infections can increasingly be prevented by prophylactic vaccines (HPV and HBV) and/or prompt diagnosis and treatment that can either cure (HCV and H. pylori) or slow down the infection (HBV and HIV), thus substantially reducing disease risk. Tools to tackle carcinogenic infections are also increasingly accessible and affordable in the EU, but their implementation is slow. Public awareness, political will and cost‐effective protocols are necessary to establish large programmes of vaccination or testing and treatment. Progress monitoring, as well as avoiding disinformation and stigma, is crucial to ensure that advances in medical progress are fully leveraged. The recently published 5th edition of the European Code Against Cancer therefore recommends: (1) vaccinate girls and boys against HBV and HPV at the age recommended in your country; (2) take part in testing and treatment for HBV and HCV, HIV and H. pylori, as recommended in your country.
In this study, dietary and serum measurements of vitamin-B6 and folate from two nested case–control studies within the European Prospective Investigation into Cancer and Nutrition study were integrated in a Bayesian framework to explore the data measurement error structure and relate dietary exposures to cancer risk. A Bayesian hierarchical model was developed, including: an exposure model, to define the unknown true intake distribution; a measurement model, to relate true intake to observed assessments; and a disease model, to estimate exposure/cancer relationships. Serum and plasma levels of vitamin-B6 and folate were inversely associated with kidney and lung cancer risk, while dietary assessments of these vitamins were not associated with kidney and lung cancer risk. The Bayesian synthesis of these data suggests a protective effect but with substantial uncertainty in the effect size.
Leverage and influence are regression diagnostics that are used to measure the sensitivity of parameter estimates to changes in the data. In this article, Bayesian leverage and influence diagnostics are derived from a local sensitivity framework in which the case weights of individual observations are perturbed. The resulting diagnostics can be applied to any Bayesian model and are easy to estimate using Markov Chain Monte Carlo. Bayesian measures of leverage and influence are closely related to predictive information criteria that are commonly used for Bayesian model choice. The penalty terms for these information criteria can be reinterpreted in terms of local sensitivity diagnostics. This connection helps to understand differences between the various information criteria that have been proposed in the literature. A comparison between leverage and influence measures leads to a new diagnostic for outlier detection. By considering multivariate case weight perturbations, groups of observations can be highlighted that are collectively outliers. This may suggest ways to improve the model. A diagnostic for sensitivity to the learning rate is also proposed that may be interpreted as a measure of prior-data conflict. This diagnostic can be adapted to measure cross-conflict between different parts of the data.
BACKGROUND:Helicobacter pylori (H. pylori) is a bacterium that colonizes the stomach and is a major risk factor for gastric cancer, with an estimated 89% of non-cardia gastric cancer cases worldwide attributable to H. pylori. Prospective studies provide reliable evidence for quantifying the association between gastric cancer and H. pylori, as they circumvent the risk of a false negative due to possible reduction in antibody levels before cancer development. METHODS:In a large-scale prospective study within the China Kadoorie Biobank, H. pylori infection is being analysed as a risk factor for gastric cancer. The presence of infection is typically determined by serological tests. The immunoblot test, although well established, is more labour intensive and uses a larger amount of plasma than the alternative high-throughput multiplex serology test. Immunoblot outputs a binary positive/negative serostatus classification, while multiplex outputs a vector of continuous antigen measurements. When mapping such multidimensional continuous measurements onto a binary classification, statistical challenges arise in defining classification cut-offs and accounting for the differences in infection evidence provided by different antigens. We discuss these challenges and propose a novel solution to optimize the translation of the continuous measurements from multiplex serology into probabilities of H. pylori infection, using classification algorithms (Bayesian additive regressive trees (BART), multidimensional monotone BART, logistic regression, random forest and elastic net). We (i) calibrate and apply classification models to predict probabilities of H. pylori infection given multiplex measurements, (ii) compare the predictive performance of the models using immunoblot as reference, (iii) discuss reasons for the differences in predictive performance and (iv) apply the calibrated models to gain insights on the relative strengths of infection evidence provided by the various antigens. RESULTS:All models showed high discriminative ability with at least 95% area under the curve (AUC) estimates on the training and test data. There was no substantial difference between the performance of models on the training and test data. CONCLUSIONS:Classification algorithms can be used to calibrate the H. pylori multiplex serology test to the immunoblot test in the China Kadoorie Biobank. This study furthers our understanding of the applicability of classification algorithms to the context of serologic tests.
Software tools for Bayesian inference have undergone rapid evolution in the past three decades, following popularisation of the first generation MCMC-sampler implementations. More recently, exponential growth in the number of users has been stimulated both by the active development of new packages by the machine learning community and popularity of specialist software for particular applications. This review aims to summarize the most popular software and provide a useful map for a reader to navigate the world of Bayesian computation. We anticipate a vigorous continued development of algorithms and corresponding software in multiple research fields, such as probabilistic programming, likelihood-free inference and Bayesian neural networks, which will further broaden the possibilities for employing the Bayesian paradigm in exciting applications.
Supplementary Table 1 from Concurrent Infection with Multiple Human Papillomavirus Types: Pooled Analysis of the IARC HPV Prevalence Surveys
The popular Bayesian meta-analysis expressed by Bayesian normal-normal hierarchical model (NNHM) synthesizes knowledge from several studies and is highly relevant in practice. Moreover, NNHM is the simplest Bayesian hierarchical model (BHM), which illustrates problems typical in more complex BHMs. Until now, it has been unclear to what extent the data determines the marginal posterior distributions of the parameters in NNHM. To address this issue we computed the second derivative of the Bhattacharyya coefficient with respect to the weighted likelihood, defined the total empirical determinacy (TED), the proportion of the empirical determinacy of location to TED (pEDL), and the proportion of the empirical determinacy of spread to TED (pEDS). We implemented this method in the R package \texttt{ed4bhm} and considered two case studies and one simulation study. We quantified TED, pEDL and pEDS under different modeling conditions such as model parametrization, the primary outcome, and the prior. This clarified to what extent the location and spread of the marginal posterior distributions of the parameters are determined by the data. Although these investigations focused on Bayesian NNHM, the method proposed is applicable more generally to complex BHMs.
Infection by certain pathogens is associated with cancer development. We conducted a case-cohort study of ~2500 incident cases of esophageal, gastric and duodenal cancer, and gastric and duodenal ulcer and a randomly selected subcohort of ~2000 individuals within the China Kadoorie Biobank study of >0.5 million adults. We used a bead-based multiplex serology assay to measure antibodies against 19 pathogens (total 43 antigens) in baseline plasma samples. Associations between pathogens and antigen-specific antibodies with risks of site-specific cancers and ulcers were assessed using Cox regression fitted using the Prentice pseudo-partial likelihood. Seroprevalence varied for different pathogens, from 0.7% for Hepatitis C virus (HCV) to 99.8% for Epstein-Barr virus (EBV) in the subcohort. Compared to participants seronegative for the corresponding pathogen, Helicobacter pylori seropositivity was associated with a higher risk of non-cardia (adjusted hazard ratio [HR] 2.73 [95% CI: 2.09-3.58]) and cardia (1.67 [1.18-2.38]) gastric cancer and duodenal ulcer (2.71 [1.79-4.08]). HCV was associated with a higher risk of duodenal cancer (6.23 [1.52-25.62]) and Hepatitis B virus was associated with higher risk of duodenal ulcer (1.46 [1.04-2.05]). There were some associations of antibodies again some herpesviruses and human papillomaviruses with risks of gastrointestinal cancers and ulcers but these should be interpreted with caution. This first study of multiple pathogens with risk of gastrointestinal cancers and ulcers demonstrated that several pathogens are associated with risks of gastrointestinal cancers and ulcers. This will inform future investigations into the role of infection in the etiology of these diseases.
I consider the development of Markov chain Monte Carlo (MCMC) methods, from late-1980s Gibbs sampling to present-day gradient-based methods and piecewise-deterministic Markov processes. In parallel, I show how these ideas have been implemented in successive generations of statistical software for Bayesian inference. These software packages have been instrumental in popularizing applied Bayesian modeling across a wide variety of scientific domains. They provide an invaluable service to applied statisticians in hiding the complexities of MCMC from the user while providing a convenient modeling language and tools to summarize the output from a Bayesian model. As research into new MCMC methods remains very active, it is likely that future generations of software will incorporate new methods to improve the user experience.
Abstract Background Helicobacter pylori infection is a major cause of non-cardia gastric cancer (NCGC), but uncertainty remains about the associations between sero-positivity to different H. pylori antigens and risk of NCGC and cardia gastric cancer (CGC) in different populations. Methods A case-cohort study in China included ∼500 each of incident NCGC and CGC cases and ∼2000 subcohort participants. Sero-positivity to 12 H. pylori antigens was measured in baseline plasma samples using a multiplex assay. Hazard ratios (HRs) of NCGC and CGC for each marker were estimated using Cox regression. These were further meta-analysed with studies using same assay. Results In the subcohort, sero-positivity for 12 H. pylori antigens varied from 11.4% (HpaA) to 70.8% (CagA). Overall, 10 antigens showed significant associations with risk of NCGC (adjusted HRs: 1.33 to 4.15), and four antigens with CGC (HRs: 1.50 to 2.34). After simultaneous adjustment for other antigens, positive associations remained significant for NCGC (CagA, HP1564, HP0305) and CGC (CagA, HP1564, HyuA). Compared with CagA sero-positive only individuals, those who were positive for all three antigens had an adjusted HR of 5.59 (95% CI 4.68–6.66) for NCGC and 2.17 (95% CI 1.54–3.05) for CGC. In the meta-analysis of NCGC, the pooled relative risk for CagA was 2.96 (95% CI 2.58–3.41) [Europeans: 5.32 (95% CI 4.05–6.99); Asians: 2.41 (95% CI 2.05–2.83); Pheterogeneity<0.0001]. Similar pronounced population differences were also evident for GroEL, HP1564, HcpC and HP0305. In meta-analyses of CGC, two antigens (CagA, HP1564) were significantly associated with a higher risk in Asians but not Europeans. Conclusions Sero-positivity to several H. pylori antigens was significantly associated with an increased risk of NCGC and CGC, with varying effects between Asian and European populations.
In nutritional epidemiology, self-reported assessments of dietary exposure are prone to measurement errors, which is responsible for bias in the association between dietary factors and risk of disease. In this study, self-reported dietary assessments were complemented by biomarkers of dietary intake. Dietary and serum measurements of folate and vitamin-B6 from two nested case-control studies within the European Prospective Investigation into Cancer and Nutrition (EPIC) study were integrated in a Bayesian model to explore the measurement error structure of the data, and relate dietary exposures to risk of site-specific cancer. A Bayesian hierarchical model was developed, which included: 1) an exposure model, to define the distribution of unknown true exposure (X); 2) a measurement model, to relate observed assessments, in turn, dietary questionnaires (Q), 24-hour recalls (R) and biomarkers (M) to X measurements; 3) a disease model, to estimate exposures/cancer relationships. The marginal posterior distribution of model parameters was obtained from the joint posterior distribution, using Markov Chain Monte Carlo (MCMC) sampling techniques in JAGS. The study included 554 and 882 case/control pairs for kidney and lung cancer, respectively. In the measurement error component, the error correlation between Q measurements of vitamin-B6 and folate was estimated to be equal to 0.82 (95% CI: 0.76, 0.87). After adjustment for age, center, sex, BMI and smoking status, the kidney cancer odds ratios (OR) were 0.55 (0.16, 1.31) and 1.07 (0.33, 3.44) for one standard deviation increase of vitamin-B6 and folate, respectively. For lung cancer ORs were 0.85 (0.27, 2.42) for vitamin-B6 and 0.55 (0.14, 1.39) for folate. Bayesian models offer powerful solutions to handle complex data structures. After accounting for the role of measurement error, folate and vitamin-B6 were not associated to the risk of kidney and lung cancer.
Objectives To systematically assess the sero-prevalence and associated factors of major infectious pathogens in China, where there are high incidence rates of certain infection-related cancers. Design Cross-sectional study. Setting 10 (5 urban, 5 rural) geographically diverse areas in China. Participants A subcohort of 2000 participants from the China Kadoorie Biobank. Primary measures Sero-prevalence of 19 pathogens using a custom-designed multiplex serology panel and associated factors. Results Of the 19 pathogens investigated, the mean number of sero-positive pathogens was 9.4 (SD 1.7), with 24.4% of participants being sero-positive for >10 pathogens. For individual pathogens, the sero-prevalence varied, being for example, 0.05% for HIV, 6.4% for human papillomavirus (HPV)-16, 53.5% for Helicobacter pylori ( H. pylori ) and 99.8% for Epstein-Barr virus . The sero-prevalence of human herpesviruses (HHV)-6, HHV-7 and HPV-16 was higher in women than men. Several pathogens showed a decreasing trend in sero-prevalence by birth cohort, including hepatitis B virus (HBV) (51.6% vs 38.7% in those born <1940 vs >1970), HPV-16 (11.4% vs 5.4%), HHV-2 (15.1% vs 8.1%), Chlamydia trachomatis (65.6% vs 28.8%) and Toxoplasma gondii (22.0% vs 9.0%). Across the 10 study areas, sero-prevalence varied twofold to fourfold for HBV (22.5% to 60.7%), HPV-16 (3.4% to 10.9%), H. pylori (16.2% to 71.1%) and C. trachomatis (32.5% to 66.5%). Participants with chronic liver diseases had >7-fold higher sero-positivity for HBV (OR=7.51; 95% CI 2.55 to 22.13). Conclusions Among Chinese adults, previous and current infections with certain pathogens were common and varied by area, sex and birth cohort. These infections may contribute to the burden of certain cancers and other non-communicable chronic diseases.
BACKGROUND:Just Another Gibbs Sampling (JAGS) is a convenient tool to draw posterior samples using Markov Chain Monte Carlo for Bayesian modeling. However, the built-in function dinterval() for censored data misspecifies the default computation of deviance function, which limits likelihood-based Bayesian model comparison.RESULTS:To establish an automatic approach to specifying the correct deviance function in JAGS, we propose a simple and generic alternative modeling strategy for the analysis of censored outcomes. The two illustrative examples demonstrate that the alternative strategy not only properly draws posterior samples in JAGS, but also automatically delivers the correct deviance for model assessment. In the survival data application, our proposed method provides the correct value of mean deviance based on the exact likelihood function. In the drug safety data application, the deviance information criterion and penalized expected deviance for seven Bayesian models of censored data are simultaneously computed by our proposed approach and compared to examine the model performance.CONCLUSIONS:We propose an effective strategy to model censored data in the Bayesian modeling framework in JAGS with the correct deviance specification, which can simplify the calculation of popular Kullback-Leibler based measures for model selection. The proposed approach applies to a broad spectrum of censored data types, such as survival data, and facilitates different censored Bayesian model structures.
Meta-analysis provides important insights for evidence-based medicine by synthesizing evidence from multiple studies which address the same research question. Within the Bayesian framework, meta-analysis is frequently expressed by a Bayesian normal-normal hierarchical model (NNHM). Recently, several publications have discussed the choice of the prior distribution for the between-study heterogeneity in the Bayesian NNHM and used several "vague" priors. However, no approach exists to quantify the informativeness of such priors, and thus, we develop a principled reference analysis framework for the Bayesian NNHM acting at the posterior level. The posterior reference analysis (post-RA) is based on two posterior benchmarks: one induced by the improper reference prior, which is minimally informative for the data, and the other induced by a highly anticonservative proper prior. This approach applies the Hellinger distance to quantify the informativeness of a heterogeneity prior of interest by comparing the corresponding marginal posteriors with both posterior benchmarks. The post-RA is implemented in the freely accessible R package ra4bayesmeta and is applied to two medical case studies. Our findings show that anticonservative heterogeneity priors produce platykurtic posteriors compared with the reference posterior, and they produce shorter 95% credible intervals (CrI) and optimistic inference compared with the reference prior. Conservative heterogeneity priors produce leptokurtic posteriors, longer 95% CrI and cautious inference. The novel post-RA framework could support numerous Bayesian meta-analyses in many research fields, as it determines how informative a heterogeneity prior is for the actual data as compared with the minimally informative reference prior.
Background Helicobacter pylori infection is a major cause of non-cardia gastric cancer (NCGC), but its causal role in cardia gastric cancer (CGC) is unclean Moreover, the reported magnitude of association with NCGC varies considerably, leading to uncertainty about population-based H pylori screening and eradication strategies in high-risk settings, particularly in China, where approximately half of all global gastric cancer cases occur Our aim was to assess the associations of H pylori infection, both overall and for individual infection biomarkers, with the risks of NCGC and CGC in Chinese adults. Methods A case-cohort study was done in adults from the prospective China Kadoorie Biobank study, aged 30-79 years from ten areas in China (Qingdao, Haikou, Harbin, Suzhou, Liuzhou, Henan, Sichuan, Hunan, Gansu, and Zhejiang), and included 500 incident NCGC cases, 437 incident CGC cases, and 500 subcohort participants who were cancerfree and alive within the first two years since enrolment in 2004-08. H pylori biomarkers were measured in stored baseline plasma samples using a sensitive immunoblot assay (HelicoBlot 2.1), with adapted criteria to define H pylori seropositivity. Cox regression was used to estimate adjusted hazard ratios (HRs) for NCGC and CGC associated with H pylori infection. These values were used to estimate the number of gastric cancer cases attributable to H pylori infection in China. Findings Of the 512 715 adults enrolled in the China Kadoorie Biobank between June, 2004, and July, 2008, 500 incident NCGC cases, 437 incident CGC cases, and 500 subcohort participants were selected for analysis. The seroprevalence of H pylori was 94.4% (95% CI 92.4-96.4) in NGCG, 92.2% (89.7-94.7) in CGC, and 75.6% (71.8-79.4) in subcohort participants. H pylori infection was associated with adjusted HRs of 5.94 (95% CI 3.25-10.86) for NCGC and 3.06 (1.54-6.10) for CGC. Among the seven individual infection biomarkers, cytotoxin-associated antigen had the highest HRs for both NCGC (HR 4.41, 95% CI 2.60-7.50) and CGC (2.94, 1.53-5.68). In this population, 78.5% of NCGC and 62.1% of CGC cases could be attributable to H pylori infection. H pylori infection accounted for an estimated 339 955 cases of gastric cancer in China in 2018. Interpretation Among Chinese adults, H pylori infection is common and is the cause of large numbers of gastric cancer cases. Population-based mass screening and the eradication of H pylori should be considered to reduce the burden of gastric cancer in high-risk settings. Copyright (C) 2021 The Author(s). Published by Elsevier Ltd.
CD4-based multi-state back-calculation methods are key for monitoring the HIV epidemic, providing estimates of HIV incidence and diagnosis rates by disentangling their inter-related contribution to the observed surveillance data. This paper, extends existing approaches to age-specific settings, permitting the joint estimation of age- and time-specific incidence and diagnosis rates and the derivation of other epidemiological quantities of interest. This allows the identification of specific age-groups at higher risk of infection, which is crucial in directing public health interventions. We investigate, through simulation studies, the suitability of various bivariate splines for the non-parametric modelling of the latent age- and time-specific incidence and illustrate our method on routinely collected data from the HIV epidemic among gay and bisexual men in England and Wales.