Risk models for disease incidence can be useful for allocating resources for disease prevention if risk assessment is not too expensive. Assume there is a preventive intervention that should be given to everyone, but preventive resources are limited. We optimize risk-based prevention strategies and investigate robustness to modeling assumptions. The optimal strategy defines the proportion of the population to be given risk assessment and who should be offered intervention. The optimal strategy depends on the ratio of available resources to resources needed to intervene on everyone, and on the ratio of the costs of risk assessment to intervention. Risk assessment is not recommended if it is too expensive. Preventive efficiency decreases with decreasing compliance to risk assessment or intervention. Risk measurement error has little effect nor does misspecification of the risk distribution. Ignoring population substructure has small effects on optimal prevention strategy but can lead to modest over- or under-spending. We give conditions under which ignoring population substructure has no effect on optimal strategy. Thus, a simple one-population model offers robust guidance on prevention strategy but requires data on available resources, costs of risk assessment and intervention, population risk distribution, and probabilities of acceptance of risk assessment and intervention.
Background: There is no model to estimate absolute invasive breast cancer risk for Hispanic women.Methods: The San Francisco Bay Area Breast Cancer Study (SFBCS) provided data on Hispanic breast cancer case patients (533 US-born, 553 foreign-born) and control participants (464 US-born, 947 foreign-born). These data yielded estimates of relative risk (RR) and attributable risk (AR) separately for US-born and foreign-born women. Nativity-specific absolute risks were estimated by combining RR and AR information with nativity-specific invasive breast cancer incidence and competing mortality rates from the California Cancer Registry and Surveillance, Epidemiology, and End Results program to develop the Hispanic risk model (HRM). In independent data, we assessed model calibration through observed/expected (O/E) ratios, and we estimated discriminatory accuracy with the area under the receiver operating characteristic curve (AUC) statistic.Results: The US-born HRM included age at first full-term pregnancy, biopsy for benign breast disease, and family history of breast cancer; the foreign-born HRM also included age at menarche. The HRM estimated lower risks than the National Cancer Institute's Breast Cancer Risk Assessment Tool (BCRAT) for US-born Hispanic women, but higher risks in foreign-born women. In independent data from the Women's Health Initiative, the HRM was well calibrated for US-born women (observed/expected [O/E] ratio - 1.07, 95% confidence interval [CI] - 0.81 to 1.40), but seemed to overestimate risk in foreign-born women (O/E ratio - 0.66, 95% CI - 0.41 to 1.07). The AUC was 0.564 (95% CI = 0.485 to 0.644) for US-born and 0.625 (95% CI = 0.487 to 0.764) for foreign-born women.Conclusions: The HRM is the first absolute risk model that is based entirely on data specific to Hispanic women by nativity. Further studies in Hispanic women are warranted to evaluate its validity.
Among 2258 Helicobacter pylori-seropositive subjects randomly assigned to receive one-time H. pylori treatment with amoxicillin-omeprazole or its placebo, we evaluated the 15-year effect of treatment on gastric cancer incidence and mortality in subgroups defined by age, baseline gastric histopathology, and post-treatment infection status. We used conditional logistic and Cox regressions for covariable adjustments in incidence and mortality analyses, respectively. Treatment was associated with a statistically significant decrease in gastric cancer incidence (odds ratio = 0.36; 95% confidence interval [CI] = 0.17 to 0.79) and mortality (hazard ratio = 0.26; 95% CI = 0.09 to 0.79) at ages 55 years and older and a statistically significant decrease in incidence among those with intestinal metaplasia or dysplasia at baseline (odds ratio = 0.56; 95% CI = 0.34 to 0.91). Treatment benefits for incidence and mortality among those with and without post-treatment infection were similar. Thus H. pylori treatment can benefit older members and those with advanced baseline histopathology, and benefits are present even with post-treatment infection, suggesting treatment can benefit an entire population, not just the young or those with mild histopathology.
Ruth Pfeiffer and colleagues describe models to calculate absolute risks for breast, endometrial, and ovarian cancers for white, non-Hispanic women over 50 years old using easily obtainable risk factors. Please see later in the article for the Editors' Summary
Some molecular analyses require microgram quantities of DNA, yet many epidemiologic studies preserve only the buffy coat. In Frederick, Maryland, in 2010, we estimated DNA yields from 5 mL of whole blood and from equivalent amounts of all-cell-pellet (ACP) fraction, buffy coat, and residual blood cells from fresh blood (n = 10 volunteers) and from both fresh and frozen blood (n = 10). We extracted DNA with the QIAamp DNA Blood Midi Kit (Qiagen Sciences, Germantown, Maryland) for silica spin column capture and measured double-stranded DNA. Yields from frozen blood fractions were not statistically significantly different from those obtained from fresh fractions. ACP fractions yielded 80.6% (95% confidence interval: 66, 97) of the yield of frozen whole blood and 99.3% (95% confidence interval: 86, 100) of the yield of fresh blood. Frozen buffy coat and residual blood cells each yielded only half as much DNA as frozen ACP, and the yields were more variable. Assuming that DNA yield and quality from frozen ACP are stable, we recommend freezing plasma and ACP. Not only does ACP yield twice as much DNA as buffy coat but it is easier to process, and its yield is less variable from person to person. Long-term stability studies are needed. If one wishes to separate buffy coat before freezing, one should also save the residual blood cell fraction, which contains just as much DNA.
In the Shandong Intervention Trial, 2 weeks of antibiotic treatment for Helicobacter pylori reduced the prevalence of precancerous gastric lesions, whereas 7.3 years of oral supplementation with garlic extract and oil (garlic treatment) or vitamin C, vitamin E, and selenium (vitamin treatment) did not. Here we report 14.7-year follow-up for gastric cancer incidence and cause-specific mortality among 3365 randomly assigned subjects in this masked factorial placebo-controlled trial. Conditional logistic regression was used to estimate the odds of gastric cancer incidence, and the Cox proportional hazards model was used to estimate the relative hazard of cause-specific mortality. All statistical tests were two-sided. Gastric cancer was diagnosed in 3.0% of subjects who received H pylori treatment and in 4.6% of those who received placebo (odds ratio = 0.61, 95% confidence interval = 0.38 to 0.96, P = .032). Gastric cancer deaths occurred among 1.5% of subjects assigned H pylori treatment and among 2.1% of those assigned placebo (hazard ratio [HR] of death = 0.67, 95% CI = 0.36 to 1.28). Garlic and vitamin treatments were associated with non-statistically significant reductions in gastric cancer incidence and mortality. Vitamin treatment was associated with statistically significantly fewer deaths from gastric or esophageal cancer, a secondary endpoint (HR = 0.51, 95% CI = 0.30 to 0.87; P = .014).
BACKGROUNDThe Breast Cancer Risk Assessment Tool (BCRAT) of the National Cancer Institute is widely used for estimating absolute risk of invasive breast cancer. However, the absolute risk estimates for Asian and Pacific Islander American (APA) women are based on data from white women. We developed a model for projecting absolute invasive breast cancer risk in APA women and compared its projections to those from BCRAT.METHODSData from 589 women with breast cancer (case patients) and 952 women without breast cancer (control subjects) in the Asian American Breast Cancer Study were used to compute relative and attributable risks based on the age at menarche, number of affected mothers, sisters, and daughters, and number of previous benign biopsies. Absolute risks were obtained by combining this information with ethnicity-specific data from the National Cancer Institute's Surveillance, Epidemiology, and End Results (SEER) program and with US ethnicity-specific mortality data to create the Asian American Breast Cancer Study model (AABCS model). Independent data from APA women in the Women's Health Initiative (WHI) were used to check the calibration and discriminatory accuracy of the AABCS model.RESULTSThe AABCS model estimated absolute risk separately for Chinese, Japanese, Filipino, Hawaiian, Other Pacific Islander, and Other Asian women. Relative and attributable risks for APA women were comparable to those in BCRAT, but the AABCS model usually estimated lower-risk projections than BCRAT in Chinese and Filipino, but not in Hawaiian women, and not in every age and ethnic subgroup. The AABCS model underestimated absolute risk by 17% (95% confidence interval = 1% to 38%) in independent data from WHI, but APA women in the WHI had incidence rates approximately 18% higher than those estimated from the SEER program.CONCLUSIONSThe AABCS model was calibrated to ethnicity-specific incidence rates from the SEER program for projecting absolute invasive breast cancer risk and is preferable to BCRAT for counseling APA women.
Background Although modifiable risk factors have been included in previous models that estimate or project breast cancer risk, there remains a need to estimate the effects of changes in modifiable risk factors on the absolute risk of breast cancer.Methods Using data from a case-control study of women in Italy (2569 case patients and 2588 control subjects studied from June 1, 1991, to April 1, 1994) and incidence and mortality data from the Florence Registries, we developed a model to predict the absolute risk of breast cancer that included five non-modifiable risk factors (reproductive characteristics, education, occupational activity, family history, and biopsy history) and three modifiable risk factors (alcohol consumption, leisure physical activity, and body mass index). The model was validated using independent data, and the percent risk reduction was calculated in high-risk subgroups identified by use of the Lorenz curve.Results The model was reasonably well calibrated (ratio of expected to observed cancers = 1.10, 95% confidence interval [CI] = 0.96 to 1.26), but the discriminatory accuracy was modest. The absolute risk reduction from exposure modifications was nearly proportional to the risk before modifying the risk factors and increased with age and risk projection time span. Mean 20-year reductions in absolute risk among women aged 65 years were 1.6% (95% CI = 0.9% to 2.3%) in the entire population, 3.2% (95% CI = 1.8% to 4.8%) among women with a positive family history of breast cancer, and 4.1% (95% CI = 2.5% to 6.8%) among women who accounted for the highest 10% of the total population risk, as determined from the Lorenz curve.Conclusions These data give perspective on the potential reductions in absolute breast cancer risk from preventative strategies based on lifestyle changes. Our methods are also useful for calculating sample sizes required for trials to test lifestyle interventions.
PURPOSE:The Gail model combines relative risks (RRs) for five breast cancer risk factors with age-specific breast cancer incidence rates and competing mortality rates from the Surveillance, Epidemiology, and End Results (SEER) program from 1983 to 1987 to predict risk of invasive breast cancer over a given time period. Motivated by changes in breast cancer incidence during the 1990s, we evaluated the model's calibration in two recent cohorts.METHODS:We included white, postmenopausal women from the National Institutes of Health (NIH) -AARP Diet and Health Study (NIH-AARP, 1995 to 2003), and the Prostate, Lung, Colorectal and Ovarian Cancer Screening Trial (PLCO, 1993 to 2006). Calibration was assessed by comparing the number of breast cancers expected from the Gail model with that observed. We then evaluated calibration by using an updated model that combined Gail model RRs with 1995 to 2003 SEER invasive breast cancer incidence rates.RESULTS:Overall, the Gail model significantly underpredicted the number of invasive breast cancers in NIH-AARP, with an expected-to-observed ratio of 0.87 (95% CI, 0.85 to 0.89), and in PLCO, with an expected-to-observed ratio of 0.86 (95% CI, 0.82 to 0.90). The updated model was well-calibrated overall, with an expected-to-observed ratio of 1.03 (95% CI, 1.00 to 1.05) in NIH-AARP and an expected-to-observed ratio of 1.01 (95% CI: 0.97 to 1.06) in PLCO. Of women age 50 to 55 years at baseline, 13% to 14% had a projected Gail model 5-year risk lower than the recommended threshold of 1.66% for use of tamoxifen or raloxifene but >or= 1.66% when using the updated model. The Gail model was well calibrated in PLCO when the prediction period was restricted to 2003 to 2006.CONCLUSION:This study highlights that model calibration is important to ensure the usefulness of risk prediction models for clinical decision making.
PURPOSE Given the high incidence of colorectal cancer (CRC), and the availability of procedures that can detect disease and remove precancerous lesions, there is a need for a model that estimates the probability of developing CRC across various age intervals and risk factor profiles. METHODS The development of separate CRC absolute risk models for men and women included estimating relative risks and attributable risk parameters from population-based case-control data separately for proximal, distal, and rectal cancer and combining these estimates with baseline age-specific cancer hazard rates based on Surveillance, Epidemiology, and End Results (SEER) incidence rates and competing mortality risks. RESULTS For men, the model included a cancer-negative sigmoidoscopy/colonoscopy in the last 10 years, polyp history in the last 10 years, history of CRC in first-degree relatives, aspirin and nonsteroidal anti-inflammatory drug (NSAID) use, cigarette smoking, body mass index (BMI), current leisure-time vigorous activity, and vegetable consumption. For women, the model included sigmoidoscopy/colonoscopy, polyp history, history of CRC in first-degree relatives, aspirin and NSAID use, BMI, leisure-time vigorous activity, vegetable consumption, hormone-replacement therapy (HRT), and estrogen exposure on the basis of menopausal status. For men and women, relative risks differed slightly by tumor site. A validation study in independent data indicates that the models for men and women are well calibrated. CONCLUSION We developed absolute risk prediction models for CRC from population-based data, and a simple questionnaire suitable for self-administration. This model is potentially useful for counseling, for designing research intervention studies, and for other applications.
Combining data from several case-control genome-wide association (GWA) studies can yield greater efficiency for detecting associations of disease with single nucleotide polymorphisms (SNPs) than separate analyses of the component studies. We compared several procedures to combine GWA study data both in terms of the power to detect a disease-associated SNP while controlling the genome-wide significance level, and in terms of the detection probability (DP). The DP is the probability that a particular disease-associated SNP will be among the T most promising SNPs selected on the basis of low p-values. We studied both fixed effects and random effects models in which associations varied across studies. In settings of practical relevance, meta-analytic approaches that focus on a single degree of freedom had higher power and DP than global tests such as summing chi-square test-statistics across studies, Fisher's combination of p-values, and forming a combined list of the best SNPs from within each study.
BACKGROUND The Breast Cancer Risk Assessment Tool of the National Cancer Institute (NCI) is widely used for counseling and determining eligibility for breast cancer prevention trials, although its validity for projecting risk in African American women is uncertain. We developed a model for projecting absolute risk of invasive breast cancer in African American women and compared its projections with those from the Breast Cancer Risk Assessment Tool. METHODS Data from 1607 African American women with invasive breast cancer and 1647 African American control subjects in the Women's Contraceptive and Reproductive Experiences (CARE) Study were used to compute relative and attributable risks that were based on age at menarche, number of affected mother or sisters, and number of previous benign biopsy examinations. Absolute risks were obtained by combining this information with data on invasive breast cancer incidence in African American women from the NCI's Surveillance, Epidemiology and End Results Program and with national mortality data. Eligibility screening data from the Study of Tamoxifen and Raloxifene (STAR) trial were used to determine how the new model would affect eligibility, and independent data from the Women's Health Initiative (WHI) were used to assess how well numbers of invasive breast cancers predicted by the new model agreed with observed cancers. RESULTS Tables and graphs for estimating relative risks and projecting absolute invasive breast cancer risk with confidence intervals were developed for African American women. Relative risks for family history and number of biopsies and attributable risks estimated from the CARE population were lower than those from the Breast Cancer Risk Assessment Tool, as was the discriminatory accuracy (i.e., concordance). Using eligibility screening data from the STAR trial, we estimated that 30.3% of African American women would have had 5-year invasive breast cancer risks of at least 1.66% by use of the CARE model, compared with only 14.5% by use of the Breast Cancer Risk Assessment Tool. The numbers of cancers predicted by the CARE model agreed well with observed numbers of cancers (i.e., it was well calibrated) in data from the WHI, except that it underestimated risk in African American women with breast biopsy examinations. CONCLUSIONS The CARE model usually gave higher risk estimates for African American women than the Breast Cancer Risk Assessment Tool and is recommended for counseling African American women regarding their risk of breast cancer.
Studies to detect genetic association with disease can be family‐based, often using families with multiple affected members, or population based, as in population‐based case‐control studies. If data on both study types are available from the same population, it is useful to combine them to improve power to detect genetic associations. Two aspects of the data need to be accommodated, the sampling scheme and potential residual correlations among family members. We propose two approaches for combining data from a case‐control study and a family study that collected families with multiple cases. In the first approach, we view a family as the sampling unit and specify the joint likelihood for the family members using a two‐level mixed effects model to account for random familial effects and for residual genetic correlations among family members. The ascertainment of the families is accommodated by conditioning on the ascertainment event. The individuals in the case‐control study are treated as families of size one, and their unconditional likelihood is combined with the conditional likelihood for the families. This approach yields subject specific maximum likelihood estimates of covariate effects. In the second approach, we view an individual as the sampling unit. The sampling scheme is accommodated using two‐phase sampling techniques, marginal covariate effects are estimated, and correlations among family members are accounted for in the variance calculations. The models are compared in simulations. Data from a case‐control and a family study from north‐eastern Italy on melanoma and a low‐risk melanoma‐susceptibility gene, MC1R, are used to illustrate the approaches. Genet. Epidemiol . 2008. Published 2008 Wiley‐Liss, Inc.
The effects of a 7.3-y supplementation with garlic and micronutrients and of anti-Helicobacter pylori treatment with amoxicillin (1 g twice daily) and omeprazole (20 mg twice daily) on serum folate, vitamin B-12, homocysteine, and glutathione concentrations were assessed in a rural Chinese population. A randomized, double-blind, placebo-controlled, factorial trial was conducted to compare the ability of 3 treatments to retard the development of precancerous gastric lesions in 3411 subjects. The treatments were: 1) anti-H. pylori treatment with amoxicillin and omeprazole; 2) 7.3-y supplementation with aged garlic and steam-distilled garlic oil; and 3) 7.3-y supplementation with vitamin C, vitamin E, and selenium. All 3 treatments were given in a 2(3) factorial design to subjects seropositive for H. pylori infection; only the garlic supplement and vitamin and selenium supplement were given in a 2(2) factorial design to the other subjects. Thirty-four subjects were randomly selected from each of the 12 treatment strata. Sera were analyzed after 7.3 y to measure effects on folate, vitamin B-12, homocysteine, and glutathione concentrations. Regression analyses adjusted for age, gender, and smoking indicated an increase of 10.2% (95%CI: 2.9-18.1%) in serum folate after garlic supplementation and an increase of 13.4% (95%CI: 5.3-22.2%) in serum glutathione after vitamin and selenium supplementation. The vitamin and selenium supplement did not affect other analytes and the amoxicillin and omeprazole therapy did not affect any of the variables tested. In this rural Chinese population, 7.3 y of garlic supplementation increased the serum folate concentration and the vitamin and selenium supplement increased that of glutathione, but neither affected serum concentrations of vitamin B-12 or homocysteine.
We analyzed data from the Breast Cancer Detection Demonstration Project (BCDDP) to obtain multivariate relative hazard models for breast cancer that included mammographic density (MD) in addition to standard risk factors. Data from the BCDDP were collected from a stratified case–control study in the screening phase (1973–1980) and from follow-up of three subcohorts in the follow-up phase (1980–1995). For both phases, MD measurements were only available for about half the women who developed breast cancer (cases) and a small fraction of noncases. We used a logistic regression model for the stratified case–control study and developed a general pseudo-likelihood approach to accommodate missing covariate data (MD) by adapting the method of Scott and Wild and Breslow and Holubkov. We showed that this method was substantially more efficient than a previously proposed weighted-likelihood method. We assumed piecewise exponential models for the analysis of each subcohort, with the missing covariate (MD) distribution conditional on the observed information modeled with polytomous logistic regression. We developed an EM algorithm for estimation, which allowed for time-varying covariates, incomplete follow-up, and left truncation. We analyzed the three follow-up subcohorts separately and then combined the relative hazard models from the case–control and cohort data. The final model included main effects for MD, weight, age at first live birth, number of previous breast biopsies, and number of sisters or mother with breast cancer and was more discriminating (higher concordance) than the original model of Gail et al., which included standard risk factors but not MD. In a separate work, we combined this relative hazard model with other data to project absolute breast cancer risk.
Purpose Validation of an absolute risk prediction model for colorectal cancer (CRC) by using a large, population-based cohort. Patients and Methods The National Institutes of Health (NIH) –American Association of Retired Persons (AARP) diet and health study, a prospective cohort study, was used to validate the model. Men and women age 50 to 71 years at baseline answered self-administered questionnaires that asked about demographic characteristics, diet, lifestyle, and medical histories. We compared expected numbers of CRC patient cases predicted by the model to the observed numbers of CRC patient cases identified in the NIH-AARP study overall and in subgroups defined by risk factor combinations. The discriminatory power was measured by the area under the receiver-operating characteristic curve (AUC). Results During an average of 6.9 years of follow-up, we identified 2,092 and 832 incident CRC patient cases in men and women, respectively. The overall expected/observed ratio was 0.99 (95% CI, 0.95 to 1.04) in men and 1.05 (95% CI, 0.98 to 1.11) in women. Agreement between the expected and the observed number of cases was good in most risk factor categories, except for in subgroups defined by CRC screening and polyp history. This discrepancy may be caused by differences in the question on screening and polyp history between two studies. The AUC was 0.61 (95% CI, 0.60 to 0.62) for men and 0.61 (95% CI, 0.59 to 0.62) for women, which was similar to other risk prediction models. Conclusion The absolute risk model for CRC was well calibrated in a large prospective cohort study. This prediction model, which estimates an individual's risk of CRC given age and risk factors, may be a useful tool for physicians, researchers, and policy makers.
SummaryLarge two‐stage genome‐wide association studies (GWASs) have been shown to reduce required genotyping with little loss of power, compared to a one‐stage design, provided a substantial fraction of cases and controls, πsample, is included in stage 1. However, a number of recent GWASs have used πsample < 0.2. Moreover, standard power calculations are not applicable because SNPs are selected in stage 1 by ranking their p‐values, rather than comparing each SNP's statistic to a fixed critical value. We define the detection probability (DP) of a two‐stage design as the probability that a given disease‐associated SNP will have a p‐value among the lowest ranks of p‐values at stage 1, and, among those SNPs selected at stage 1, at stage 2. For 8000 cases and 8000 controls available for study and for odds ratios per allele in the range 1.1‐1.3, we show that DP is substantially reduced for designs with πsample≤ 0.25, and that DP cannot be appreciably increased by analyzing the stage 1 and stage 2 data jointly. These results suggest that multistage designs with small first stages (e.g. πsample≤ 0.25) should be avoided, and that additional genotyping in earlier studies with small first stages will yield previously unselected disease‐associated SNPs.
Some case-control genome-wide association studies (CCGWASs) select promising single nucleotide polymorphisms (SNPs) by ranking corresponding p-values, rather than by applying the same p-value threshold to each SNP. For such a study, we define the detection probability (DP) for a specific disease-associated SNP as the probability that the SNP will be "T-selected," namely have one of the top T largest chi-square values (or smallest p-values) for trend tests of association. The corresponding proportion positive (PP) is the fraction of selected SNPs that are true disease-associated SNPs. We study DP and PP analytically and via simulations, both for fixed and for random effects models of genetic risk, that allow for heterogeneity in genetic risk. DP increases with genetic effect size and case-control sample size and decreases with the number of nondisease-associated SNPs, mainly through the ratio of T to N, the total number of SNPs. We show that DP increases very slowly with T, and the increment in DP per unit increase in T declines rapidly with T. DP is also diminished if the number of true disease SNPs exceeds T. For a genetic odds ratio per minor disease allele of 1.2 or less, even a CCGWAS with 1000 cases and 1000 controls requires T to be impractically large to achieve an acceptable DP, leading to PP values so low as to make the study futile and misleading. We further calculate the sample size of the initial CCGWAS that is required to minimize the total cost of a research program that also includes follow-up studies to examine the T-selected SNPs. A large initial CCGWAS is desirable if genetic effects are small or if the cost of a follow-up study is large.