We discuss maximum likelihood and estimating equations methods for combining results from multiple studies in pooling projects and data consortia using a meta-analysis model, when the multivariate estimates with their covariance matrices are available. The estimates to be combined are typically regression slopes, often from relative risk models in biomedical and epidemiologic applications. We generalize the existing univariate meta-analysis model and investigate the efficiency advantages of the multivariate methods, relative to the univariate ones. We generalize a popular univariate test for between-studies homogeneity to a multivariate test. The methods are applied to a pooled analysis of type of carotenoids in relation to lung cancer incidence from seven prospective studies. In these data, the expected gain in efficiency was evident, sometimes to a large extent. Finally, we study the finite sample properties of the estimators and compare the multivariate ones to their univariate counterparts.
With the growing number of epidemiologic publications on the relation between dietary factors and cancer risk, pooled analyses that summarize results from multiple studies are becoming more common. Here, the authors describe the methods being used to summarize data on diet-cancer associations within the ongoing Pooling Project of Prospective Studies of Diet and Cancer, begun in 1991. In the Pooling Project, the primary data from prospective cohort studies meeting prespecified inclusion criteria are analyzed using standardized criteria for modeling of exposure, confounding, and outcome variables. In addition to evaluating main exposure-disease associations, analyses are also conducted to evaluate whether exposure-disease associations are modified by other dietary and nondietary factors or vary among population subgroups or particular cancer subtypes. Study-specific relative risks are calculated using the Cox proportional hazards model and then pooled using a random- or mixed-effects model. The study-specific estimates are weighted by the inverse of their variances in forming summary estimates. Most of the methods used in the Pooling Project may be adapted for examining associations with dietary and nondietary factors in pooled analyses of case-control studies or case-control and cohort studies combined.
Stephanie A. Smith-Warner, Donna Spiegelman, John Ritz, Demetrius Albanes, W. Lawrence Beeson, Leslie Bernstein, Franco Berrino, Piet A. van den Brandt, Julie E. Buring, Eunyoung Cho, Graham A. Colditz, Aaron R. Folsom, Jo L. Freudenheim, Edward Giovannucci, R. Alexandra Goldbohm, Saxon Graham, Lisa Harnack, Pamela L. HornRoss, Vittorio Krogh, Michael F. Leitzmann, Marjorie L. McCullough, Anthony B. Miller, Carmen Rodriguez, Thomas E. Rohan, Arthur Schatzkin, Roy Shore, Mikko Virtanen, Walter C. Willett, Alicja Wolk, Anne Zeleniuch-Jacquotte, Shumin M. Zhang, and David J. Hunter
Background: Although smoking is the primary cause of lung cancer, much is unknown about lung cancer etiology, including risk determinants for nonsmokers and modifying factors for smokers.
Background: Epidemiologic studies have generally reported positive associations between alcohol consumption and risk for colorectal cancer. However, findings related to specific alcoholic beverages or different anatomic sites in the large bowel have been inconsistent.Objective: To examine the relationship of total alcohol intake and intake from specific beverages to the incidence of colorectal cancer and to evaluate whether other potential risk factors modify the association.Design: Pooled analysis of primary data from 8 cohort studies in 5 countries.Setting: North America and Europe.Participants: 489 979 women and men with no history of cancer other than nonmelanoma skin cancer at baseline.Measurements: Alcohol intake was assessed in each study at baseline by using a validated food-frequency questionnaire.Results: During a maximum of 6 to 16 years of follow-up across the studies, 4687 cases of colorectal cancer were documented. In categorical analyses, increased risk for colorectal cancer was limited to persons with an alcohol intake of 30 g/d or greater (approximately greater than or equal to2 drinks/d), a consumption level reported by 4% of women and 13% of men. Compared with nondrinkers, the pooled multivariate relative risks were 1.16 (95% Cl, 0.99 to 1.36) for persons who consumed 30 to less than 45 g/d and 1.41 (Cl, 1.16 to 1.72) for those who consumed 45 g/d or greater. No significant heterogeneity by study or sex was observed. The association was evident for cancer of the proximal colon, distal colon, and rectum. No clear difference in relative risks was found among specific alcoholic beverages.Limitations: The study included only one measure of alcohol consumption at baseline and could not investigate lifetime alcohol consumption, alcohol consumption at younger ages, or changes in alcohol consumption during follow-up. It also could not examine drinking patterns or duration of alcohol use.Conclusions: A single determination of alcohol intake correlated with a modest relative elevation in colorectal cancer rate, mainly at the highest levels of alcohol intake.
Certain statistical models specify a conditional mean function, given a random effect and covariates of interest. On the other hand, one may instead model a marginal mean only in terms of the covariates. We discuss some common situations where conditional and marginal means coincide. In a Gaussian linear mixed effects model we have equivalent interpretations of the conditional and marginal regression parameter estimates. Similar results exist for more general link functions. In this paper we give a short overview of some models, where conditional and marginal results are equivalent and we illustrate this with some examples. When the conditional mean is additive in a random effect on the log scale, it is seen that the marginal mean equals the conditional mean plus a constant, such that slope parameters have the same interpretation in both formulations. No further distributional assumptions are needed in either of these cases. With a logit link and a double exponential random effect, a closed form marginal link function is derived from the conditional model. When a logit or probit link is used with a normal random effect, the marginal mean parameters become attenuated by a factor which depends on parameters of the distribution of the covariates. In a conditional Weibull proportional hazards model with a positive stable frailty, the marginal hazards are again Weibull but with slope parameters attenuated towards zero.
BACKGROUND Epidemiologic studies have suggested a lower risk of coronary heart disease (CHD) at higher intakes of fruit, vegetables, and whole grain. Whether this association is due to antioxidant vitamins or some other factors remains unclear. OBJECTIVE We studied the relation between the intake of antioxidant vitamins and CHD risk. DESIGN A cohort study pooling 9 prospective studies that included information on intakes of vitamin E, carotenoids, and vitamin C and that met specific criteria was carried out. During a 10-y follow-up, 4647 major incident CHD events occurred in 293 172 subjects who were free of CHD at baseline. RESULTS Dietary intake of antioxidant vitamins was only weakly related to a reduced CHD risk after adjustment for potential nondietary and dietary confounding factors. Compared with subjects in the lowest dietary intake quintiles for vitamins E and C, those in the highest intake quintiles had relative risks of CHD incidence of 0.84 (95% CI: 0.71, 1.00; P=0.17) and 1.23 (1.04, 1.45; P=0.07), respectively, and the relative risks for subjects in the highest intake quintiles for the various carotenoids varied from 0.90 to 0.99. Subjects with higher supplemental vitamin C intake had a lower CHD incidence. Compared with subjects who did not take supplemental vitamin C, those who took >700 mg supplemental vitamin C/d had a relative risk of CHD incidence of 0.75 (0.60, 0.93; P for trend <0.001). Supplemental vitamin E intake was not significantly related to reduced CHD risk. CONCLUSIONS The results suggest a reduced incidence of major CHD events at high supplemental vitamin C intakes. The risk reductions at high vitamin E or carotenoid intakes appear small.
Background: Epidemiologic studies have suggested a lower risk of coronary heart disease (CHD) at higher intakes of fruit, vegetables, and whole grain. Whether this association is due to antioxidant vitamins or some other factors remains unclear. Objective: We studied the relation between the intake of antioxidant vitamins and CHD risk. Design: A cohort study pooling 9 prospective studies that included information on intakes of vitamin E, carotenoids, and vitamin C and that met specific criteria was carried out. During a 10-y follow-up, 4647 major incident CHD events occurred in 293 172 subjects who were free of CHD at baseline. Results: Dietary intake of antioxidant vitamins was only weakly related to a reduced CHD risk after adjustment for potential nondietary and dietary confounding factors. Compared with subjects in the lowest dietary intake quintiles for vitamins E and C, those in the highest intake quintiles had relative risks of CHD incidence of 0.84 (95% CI: 0.71, 1.00; P 0.17) and 1.23 (1.04, 1.45; P 0.07), respectively, and the relative risks for subjects in the highest intake quintiles for the various carotenoids varied from 0.90 to 0.99. Subjects with higher supplemental vitamin C intake had a lower CHD incidence. Compared with subjects who did not take supplemental vitamin C, those who took 700 mg supplemental vitamin C/d had a relative risk of CHD incidence of 0.75 (0.60, 0.93; P for trend 0.001). Supplemental vitamin E intake was not significantly related to reduced CHD risk. Conclusions: The results suggest a reduced incidence of major CHD events at high supplemental vitamin C intakes. The risk reductions at high vitamin E or carotenoid intakes appear small. Am J Clin Nutr 2004;80:1508–20.
Lung cancer rates are highest in countries with the greatest fat intakes. In several case-control studies, positive associations have been observed between lung cancer and intakes of total and saturated fat, particularly among nonsmokers. We analyzed the association between fat and cholesterol intakes and lung cancer risk in eight prospective cohort studies that met predefined criteria. Among the 280,419 female and 149,862 male participants who were followed for up to 6-16 years, 3,188 lung cancer cases were documented. Using the Cox proportional hazards model, we calculated study-specific relative risks that were adjusted for smoking history and other potential risk factors. Pooled relative risks were computed using a random effects model. Fat intake was not associated with lung cancer risk. For an increment of 5% of energy from fat, the pooled multivariate relative risks were 1.01 [95% confidence interval (CI), 0.98-1.05] for total, 1.03 (95% CI, 0.96-1.11) for saturated, 1.01 (95% CI, 0.93-1.10) for monounsaturated, and 0.99 (95% CI, 0.90-1.10) for polyunsaturated fat. No associations were observed between intakes of total or specific types of fat and lung cancer risk among never, past, or current smokers. Dietary cholesterol was not associated with lung cancer incidence [for a 100-mg/day increment, the pooled multivariate relative risk was 1.01 (95% CI, 0.97-1.05)]. There was no statistically significant heterogeneity among studies or by sex. These data do not support an important relation between fat or cholesterol intakes and lung cancer risk. The means to prevent this important disease remains avoidance of smoking.
Consider the use of Generalized Estimating Equations (GEEs) for estimation of regression parameters from longitudinal series. Let Y-it be the outcome for the i(th) series at time t and let X(i) = (X(il),...,X(ini))' be covariate vectors associated with n(i) observation times. We investigate bias of GEE estimates for population-average (PA) and conditional parameters under model misspecification which takes the form of omission of past history from the model for Y-it\X(i). We provide exact bias results for the identity link, a bias approximation for nonlinear links and simulation results. Bias for either parameter can be positive or negative and depends on the size of the series and the strength of association between observations at times t and t - 1.