Epigenome-wide mediation analysis aims to identify high-dimensional DNA methylation at cytosine–phosphate–guanine (CpG) sites that mediate the causal effect of linking smoking with Crohn’s disease (CD) outcome. Studies have shown that smoking has significant detrimental effects on the course of CD. So we assessed whether DNA methylation mediates the association between smoking and CD. Among 103 CD cases and 174 controls, we estimated whether the effects of smoking on CD are mediated through DNA methylation CpG sites, which we referred to as causal mediation effect. Based on the causal diagram, we first implemented sure independence screening (SIS) to reduce the pool of potential mediator CpGs from a very large to a moderate number; then, we implemented variable selection with de-sparsifying the LASSO regression. Finally, we carried out a comprehensive mediation analysis and conducted sensitivity analysis, which was adjusted for potential confounders of age, sex, and blood cell type proportions to estimate the mediation effects. Smoking was significantly associated with CD under odds ratio (OR) of 2.319 (95% CI: 1.603, 3.485, p < 0.001) after adjustment for confounders. Ninety-nine mediator CpGs were selected from SIS, and then, seven candidate CpGs were obtained by de-sparsifying the LASSO regression. Four of these CpGs showed statistical significance, and the average causal mediation effects (ACME) were attenuated from 0.066 to 0.126. Notably, three significant mediator CpGs had absolute sensitivity parameters of 0.40, indicating that these mediation effects were robust even when the assumptions were slightly violated. Genes (BCL3 and FKBP5) harboring these four CpGs were related to CD. These findings suggest that changes in methylation are involved in the mechanism by which smoking increases risk of CD.
目的 通过统计模拟和实例数据分析,探索当存在不可观测的混杂因素时,Logis-tic回归分析模型中调整工具变量(instrumental variable,IV)对估计因果效应的影响.方法 设定变量均服从二项分布,在Logistic回归分析模型中依次使用不同的参数进行统计模拟,以因果效应估计值的偏倚和标准误作为评价指标;实例数据分析是基于山东省多家医院健康体检中心的体检随访数据,以高血压为目标结局,构建纵向观察队列,筛选单核苷酸多态性(single nucleo-tide polymorphism,SNP)位点rs12149832作为IV,在Logistic回归分析模型中,采用不同策略(纳入/不纳入rs12149832协变量)来分析BMI与患高血压风险之间的关系.结果 统计模拟结果显示在以Logistic回归分析模型估计暴露与结局间的效应时,协变量集中纳入IV会增大效应估计的偏倚和标准误,但增大程度较小;实例分析中,高血压队列共纳入1240名女性,基线年龄为(37.7±10.5)岁,BMI为(22.1±3.1)kg/m2.纳入IV的模型所得的效应估计值为0.225(P<0.001),略小于不包含IV的回归模型所得的效应估计值(0.228,P<0.001),基本验证了关于纳入IV进行调整的统计模拟结果.结论 观察性流行病学研究中,Logistic回归分析模型误纳入IV对效应估计值的偏倚和标准误均有影响.
目的 探索DNA甲基化是否为BMI和胰岛素-妊娠期糖尿病(insulin-treated gestation-al diabetes mellitus,I-GDM)之间的中介变量,并估计其中介效应.方法 研究资料来自基因表达数据库(gene expression omnibus,GEO),检索号为GSE88929,由产科医生收集44例I-GDM病例和64例对照,共纳入212991个胞嘧啶-磷酸-鸟嘌呤双核苷酸(cytosine-phosphate-guanine pairs of nucleotides,CpG)位点.采用因果推断检验(causal inference test,CIT)筛选出潜在的中介CpG位点.调整孕龄、胎儿性别和胎龄三个混杂因素,进一步通过CpG位点估计BMI对I-GDM的中介效应.结果 CIT过程第一步结果表明BMI与I-GDM有关联(OR=1.057,95% CI:1.014~1.105,P=0.010);第二步调整BMI,采用伪发现率(false discorery rate,FDR)方法进行多重检验校正,共筛选出与I-GDM相关联的6348个CpG位点纳入下一步分析;第三步调整I-GDM后,确定529个CpG位点分别与BMI有关联;第四步分别调整6个CpG位点后,BMI与I-GDM相互独立.因此,CIT检验出6个CpG位点(cg00542041、cg08589721、cg25775742、cg15819225、cg26824326、cg15110463)作为BMI与I-GDM的中介变量.进一步采用中介分析因果推断模型,证实上述6个CpG存在中介效应.结论 本研究发现6个DNA甲基化CpG位点在BMI和I-GDM之间发挥了中介作用,均可能作为I-GDM发病机制中新的生物标记物,为研究BMI和I-GDM之间复杂的生物学机制提供了参考依据.
OBJECTIVES:In observational studies, epidemiologists often attempt to estimate the total effect of an exposure on an outcome of interest. However, when the underlying diagram is unknown and limited knowledge is available, dissecting bias performances is essential to estimating the total effect of an exposure on an outcome when mistakenly adjusting for mediators under logistic regression. Through simulation, we focused on six causal diagrams concerning different roles of mediators. Sensitivity analysis was conducted to assess the bias performances of varying across exposure-mediator effects and mediator-outcome effects when adjusting for the mediator.SETTING:Based on the causal relationships in the real world, we compared the biases of varying across the effects of exposure-mediator with those of varying across the effects of mediator-outcome when adjusting for the mediator. The magnitude of the bias was defined by the difference between the estimated effect (using logistic regression) and the total effect of the exposure on the outcome.RESULTS:In four scenarios (a single mediator, two series mediators, two independent parallel mediators or two correlated parallel mediators), the biases of varying across the effects of exposure-mediator were greater than those of varying across the effects of mediator-outcome when adjusting for the mediator. In contrast, in two other scenarios (a single mediator or two independent parallel mediators in the presence of unobserved confounders), the biases of varying across the effects of exposure-mediator were less than those of varying across the effects of mediator-outcome when adjusting for the mediator.CONCLUSIONS:The biases were more sensitive to the variation of effects of exposure-mediator than the effects of mediator-outcome when adjusting for the mediator in the absence of unobserved confounders, while the biases were more sensitive to the variation of effects of mediator-outcome than those of exposure-mediator in the presence of an unobserved confounder.
Objective To establish a prediction model to estimate risks of cataract among health management population aged above 50 years.Methods Based on the Shandong Multi-center Longitudinal Cohort for Health Management,a prediction model for cataract was constructed using Cox's proportional hazards regression model.The predictability was evaluated with the area under the receiver operating characteristic (ROC) curve (AUC).The stability was tested with ten-fold cross-validation.Results During the follow-up period,there were 1010 new cataract cases,and the incidence density was 24.76‰.The risk factors included in prediction model were age,sex,smoking habit,hyperviscosemia,tympanic diseases,ametropia,diabetes,total cholesterol and systolic blood pressure (SBP).The AUC of the prediction model was 0.712 (95% CI:0.693-0.732).The ten-fold cross-validation showed that the AUC was 0.714.Conclusion The prediction model of cataract has high predictability and reliability.It can provide scientific basis for identifying high-risk groups of cataract.
BACKGROUND:Confounders can produce spurious associations between exposure and outcome in observational studies. For majority of epidemiologists, adjusting for confounders using logistic regression model is their habitual method, though it has some problems in accuracy and precision. It is, therefore, important to highlight the problems of logistic regression and search the alternative method.METHODS:Four causal diagram models were defined to summarize confounding equivalence. Both theoretical proofs and simulation studies were performed to verify whether conditioning on different confounding equivalence sets had the same bias-reducing potential and then to select the optimum adjusting strategy, in which logistic regression model and inverse probability weighting based marginal structural model (IPW-based-MSM) were compared. The "do-calculus" was used to calculate the true causal effect of exposure on outcome, then the bias and standard error were used to evaluate the performances of different strategies.RESULTS:Adjusting for different sets of confounding equivalence, as judged by identical Markov boundaries, produced different bias-reducing potential in the logistic regression model. For the sets satisfied G-admissibility, adjusting for the set including all the confounders reduced the equivalent bias to the one containing the parent nodes of the outcome, while the bias after adjusting for the parent nodes of exposure was not equivalent to them. In addition, all causal effect estimations through logistic regression were biased, although the estimation after adjusting for the parent nodes of exposure was nearest to the true causal effect. However, conditioning on different confounding equivalence sets had the same bias-reducing potential under IPW-based-MSM. Compared with logistic regression, the IPW-based-MSM could obtain unbiased causal effect estimation when the adjusted confounders satisfied G-admissibility and the optimal strategy was to adjust for the parent nodes of outcome, which obtained the highest precision.CONCLUSIONS:All adjustment strategies through logistic regression were biased for causal effect estimation, while IPW-based-MSM could always obtain unbiased estimation when the adjusted set satisfied G-admissibility. Thus, IPW-based-MSM was recommended to adjust for confounders set.
Objective To construct prediction models to estimate the risks of developing type 2 diabetes mellitus (T2DM) in 3 years among the health management population in mainland China.Methods Non-diabetic people aged 20 to 75 years at the baseline were chosen from Shandong Multi-center Longitudinal Cohort for Health Management to compose our cohort.Cox's proportional hazards regression model was adopted to build T2DM prediction model.The area under the receiver operating characteristic (ROC) curve (AUC) was used to evaluate the predictability of the model.Ten-fold cross-validation was adopted to test the stability of the model.Results During the follow-up of 3.68 ± 2.8 years,1,624 cases of new-onset diabetes occurred.The incidence density of male and female was 15.00‰ and 10.83‰,respectively.The risk factors for the male model included age,body mass index (BMI),fasting plasma glucose (FPG),triglyceride,alanine aminotransferase (ALT),and white blood cell (WBC) count.The risk factors for the female model included age,FPG,triglyceride,high density lipoprotein cholesterol (HDL-C),and ALT.The AUC of the male model and female model was 0.795 (95% CI:0.764-0.827) and 0.707 (95% CI:0.654-0.759),respectively.Conclusion The male and female prediction models we constructed have high predictability and reliability among the health management population.
Objective To introduce the theory of sub-distribution hazard model and its applications in health risk assessment.Methods Given that competing risks are commonly encountered in health risk assessment,we have introduced the sub-distribution hazard model,and further evaluate itsefficiency and application based on the Shandong Multi-center cohort of hemorrhagic cerebral apoplexy.Results Under the framework of competing risk,the sub-distribution hazard model took the competing endpoint into the construction of the risk set,other than treating the competing endpoint simply as the censoring.Thus,it can have better performance than the traditional Cox model.Based on the Shandong Multi-center cohort of hemorrhagic cerebral apoplexy,it showed a good practicability in risk assessment of stroke death.Conclusion The sub-distribution hazard model can directly link the covariates with the cumulative incidence function,and can efficiently deal with the competing risk problem.
Objective To introduce the theory of cause-specific hazard model and its applications in health risk assessment.Methods Given that competing risks are commonly encountered in health risk assessment,we have introduced the cause-specific hazard model from the perspective of modeling principle and parameter estimator,and further evaluate the efficiency and application based on the Shandong Multi-center cohort of hemorrhagic cerebral apoplexy.Results Under the framework of competing risk,the cause-specific hazard model constructed the cox-type survival model for each cause-specific endpoint.The parameter estimator can be obtained from the partial likelihood and thus have good properties,this can make sure that the absolute risk is accurate.Based on the Shandong Multi-center cohort of hemorrhagic cerebral apoplexy,it showed a good practicability in risk assessment of stroke death.Conclusion When competing risk can not be ignored,the cause-specific hazard model are preferred in health risk assessment.
In observational studies, matched case-control designs are routinely conducted to improve study precision. How to select covariates for match or adjustment, however, is still a great challenge for estimating causal effect between the exposure E and outcome D.