Supplementary Figure 1 from CYP1A1/2 Haplotypes and Lung Cancer and Assessment of Confounding by Population Stratification
<div>Abstract<p>Prior studies of lung cancer and <i>CYP1A1/2</i> in African-American and Latino populations have shown inconsistent results and have not yet investigated the haplotype block structure of <i>CYP1A1/2</i> or addressed potential population stratification. To investigate haplotypes in the <i>CYP1A1/2</i> region and lung cancer in African-Americans and Latinos, we conducted a case-control study (1998–2003). African-Americans (<i>n</i> = 535) and Latinos (<i>n</i> = 412) were frequency matched on age, sex, and self-reported race/ethnicity. We used a custom genotyping panel containing 50 single nucleotide polymorphisms in the <i>CYP1A1/2</i> region and 184 ancestry informative markers selected to have large allele frequency differences between Africans, Europeans, and Amerindians. Latinos exhibited significant haplotype main effects in two blocks even after adjusting for admixture [odds ratio (OR), 2.02; 95% confidence interval (95% CI), 1.28–3.19 and OR, 0.55; 95% CI, 0.36–0.83], but no main effects were found among African-Americans. Adjustment for admixture revealed substantial confounding by population stratification among Latinos but not African-Americans. Among Latinos and African-Americans, interactions between smoking level and haplotypes were not statistically significant. Evidence of population stratification among Latinos underscores the importance of adjusting for admixture in lung cancer association studies, particularly in Latino populations. These results suggest that a variant occurring within the <i>CYP1A2</i> region may be conferring an increased risk of lung cancer in Latinos. [Cancer Res 2009;69(6):2340–8]</p></div>
Supplementary Tables 1-4 from <i>CYP1A1/2</i> Haplotypes and Lung Cancer and Assessment of Confounding by Population Stratification
Background: The observance of nonrandom space-time groupings of childhood cancer has been a concern of health professionals and the general public for decades. Many childhood cancers are suspected to have initiated in utero; therefore, we examined the spatial-temporal randomness of the birthplace of children who later developed cancer. Methods: We performed a space-time cluster analysis using birth addresses of 5,896 cases and 23,369 population-based, age-, sex-, and race/ethnicity-matched controls in California from 1997 to 2007, evaluating 20 types of childhood cancer and three a priori designated subgroups of childhood acute lymphoblastic leukemia (ALL). We analyzed data using a newly designed semiparametric analysis program, ClustR, and a common algorithm, SaTScan. Results: We observed evidence for nonrandom space-time clustering for ALL diagnosed at 2-6 years of age in the South San Francisco Bay Area (ClustR P = 0.04, SaTScan P = 0.07), and malignant gonadal germ cell tumors in a region of Los Angeles (ClustR P = 0.03, SaTScan P = 0.06). ClustR did not identify evidence of clustering for other childhood cancers, although SaTScan suggested some clustering for Hodgkin lymphoma (P = 0.09), astrocytoma (P = 0.06), and retinoblastoma (P = 0.06). Conclusions: Our study provides evidence that childhood ALL diagnosed at 2-6 years and malignant gonadal germ cell tumors sporadically occurs in nonrandom space-time clusters. Further research is warranted to identify epidemiologic features that may inform the underlying etiology.
Background: Until recently, large individual-level longitudinal data were unavailable to investigate clusters of disease, driving a need for suitable statistical tools. We introduce a robust, efficient, intuitive R package, ClustR, for space-time cluster analysis of individual-level data. Methods: We developed ClustR and evaluated the tool using a simulated dataset mirroring the population of California with constructed clusters. We assessed Cluster's performance under various conditions and compared it with another space-time clustering algorithm: SaTScan. Results: ClustR mostly exhibited high sensitivity for urban clusters and low sensitivity for rural clusters. Specificity was generally high. Compared with SaTScan, ClustR ran faster and demonstrated similar sensitivity, but had lower specificity. Select cluster types were detected better by ClustR than SaTScan and vice versa. Conclusion: ClustR is a user-friendly, publicly available tool designed to perform efficient cluster analysis on individual-level data, filling a gap among current tools. ClustR and SaTScan exhibited different strengths and may be useful in conjunction.
A series of puzzles are often described in terms of statistical issues and methods. Some examples are presented.
This statistical tool is a useful and interesting approach to describing the relationships within tabled data and is illustrated by a description of the relationship between smoking exposure and levels of socioeconomic status data. The measure capitalizes on the observed maximum values to identify association.
A poem entirely devoted to the description of statistics by a Nobel prize winning poet.
The term confounding is defined and discussed as a general statistical concept and extensively illustrated by an analysis of black/white differences in infant perinatal mortality with particular emphasis on stratification and weighted average summaries.
A discussion and complete description of the odds ratio in the context of epidemiologic data.
Binary data is often an important type of data. The description of twin pairs presents an extensive example. Also introduced are the statistical consequences of random pairs illustrated by the analysis of the occurrence of twins with birth defects.
A complete discussion of a faulty analysis that produced the conclusion that major risk factors in the frequency of coronary disease range from 87% to 90%.
A description of elementary probability techniques and applications useful to understanding the statistical tools that follow later in the book.
Human disease and mortality rates are important statistical measures. The analysis and inferences based on calculated rates are defined and their estimation is described.
Odds are a frequently used and popular metric in the medical and epidemiologic literature and are often generated by logistic regression analysis.
Often the incidence of a specific virus is estimated from collected data and a process of pooling these data provides an efficient and easily applied shortcut.
A description of a simple and robust statistical technique to estimate a straight line to summarize the relationship between paired observations.
Like the normal distribution, the log-normal distribution’s properties are useful and often revealing. A short description presents an analysis of log-normal data collected to asses cancer risk associated with a specific pesticide.
A number of approaches are possible. A specific analysis describes the analysis of tabulated data using the concept of additivity. Several examples illustrate the description.
The mean value is perhaps the most fundamental statistic and the chapter describes its properties and features that make it an important and necessary analytic tool.