The objective of this book is to describe procedures for analyzing genome-wide association studies (GWAS). Some of the material is unpublished and contains commentary and unpublished research; other chapters (Chapters 4 through 7) have been published in other journals. Each previously published chapter investigates a different genomics model, but all focus on identifying the strengths and limitations of various statistical procedures that have been applied to different GWAS scenarios.
In this report, we address a scenario that uses synthetic genotype case-control data that is influenced by environmental factors in a genome-wide association study (GWAS) context. The precise way the environmental influence contributes to a given phenotype is typically unknown. Therefore, our study evaluates how to approach a GWAS that may have an environmental component. Specifically, we assess different statistical models in the context of a GWAS to make association predictions when the form of the environmental influence is questionable. We used a simulation approach to generate synthetic data corresponding to a variety of possible environmental-genetic models, including a “main effects only” model as well as a “main effects with interactions” model. Our method takes into account the strength of the association between phenotype and both genotype and environmental factors, but we focus on low-risk genetic and environmental risks that necessitate using large sample sizes (N = 10,000 and 200,000) to predict associations with high levels of confidence. We also simulated different Mendelian gene models, and we analyzed how the collection of factors influences statistical power in the context of a GWAS. Using simulated data provides a “truth set” of known outcomes such that the association-affecting factors can be unambiguously determined. We also test different statistical methods to determine their performance properties. Our results suggest that the chances of predicting an association in a GWAS is reduced if an environmental effect is present and the statistical model does not adjust for that effect. This is especially true if the environmental effect and genetic marker do not have an interaction effect. The functional form of the statistical model also matters. The more accurately the form of the environmental influence is portrayed by the statistical model, the more accurate the prediction will be. Finally, even with very large samples sizes, association predictions involving recessive markers with low risk can be poor.
The objective of this book is to describe procedures for analyzing genome-wide association studies (GWAS). Some of the material is unpublished and contains commentary and unpublished research; other chapters (Chapters 4 through 7) have been published in other journals. Each previously published chapter investigates a different genomics model, but all focus on identifying the strengths and limitations of various statistical procedures that have been applied to different GWAS scenarios.
301 Received August 22, 2016; Accepted November 1, 2016; Epub November 14, 2016; https://doi.org/10.14573/altex.1608221 systems for accurate and cost-effective prediction of human health outcomes. Development of increasingly sophisticated in vitro tissue models has accelerated in recent years, driven by the recognition that conventional two-dimensional (2D) cell culture formats do not adequately recapitulate the 3D arrangements of cells and extracellular matrix of tissues and
Nearly one thousand human genome wide association studies (GWAS) have examined over 210 diseases and traits and found over 1,200 SNP associations. With improved genotyping technologies and the growing number of available markers, case-control Genome Wide Association Studies (GWAS) have become a key tool for investigating complex diseases. This study assesses the influence of genotype and diagnosis errors present in GWAS by analyzing a synthetic gene dataset incorporating factors known to influence association measurement. Monte Carlo methods were used to generate the synthetic gene data, which incorporated factors including gene inheritance, relative risk levels, disease penetrance, genotype distribution, sample size, as well as the two error factors that are the focus of this study. The resulting dataset provides a truth set for assessing statistical method performance and association sensitivity. While previously understood, these results quantify and document the extent of the relationship between genotype and diagnosis error measures and statistical power loss. Our results also demonstrate that for low risk non-recessive loci, sample sizes in the range of 1,000 - 2,000 cases will achieve 80% power thresholds for error type I error levels of 10-8 even with realistic genotype and phenotype error assumptions. Nevertheless, compensating for power loss due to the presence of genotype and diagnosis errors by increasing sample size should not be underestimated. Our estimates indicate that sample size increase requirements are in the range of 20% to 40%, depending on the gene inheritance model assumed.
Maternal nutrition influences a child’s birthweight, which affects the child’s growth and subsequent survival. However, the broad consequences of maternal undernutrition and the outcomes of interventions to improve maternal nutrition take years to manifest. To examine the long-term health outcomes of low birthweight infants in response to a maternal nutritional supplementation intervention without this obstacle, we developed the Forecasting Population Progress (FPOP) microsimulation model. The intervention we assessed was based on the findings of a published clinical trial outcome that reduced the incidence of low birthweight, a known cause of stunting. We implemented the “before intervention” and “after intervention” simulations and generated the difference in outcomes, using a spatially explicit synthetic baseline population of Indonesia generated from a microdata sample of the Indonesian 2010 census. We focused specifically on two provinces—Yogyakarta and Bali—which represent different levels of fertility and mortality but both exhibit significant underweight birth. The baseline scenario represented the current nutritional status of pregnant women in the two Indonesian provinces and projected that implementing a multiple nutrition supplementation intervention would, after 30 years, avert 8 per 1,000 low birthweight births, 3.8 per 1,000 stunted children younger than 5 years of age, and 0.25 infant deaths per 1,000 births. As our model results demonstrate, improvement in maternal nutrition would reduce infant mortality, but an even greater impact could be the reduction in growth stunting.
The objective of this book is to describe procedures for analyzing genome-wide association studies (GWAS). Some of the material is unpublished and contains commentary and unpublished research; other chapters (Chapters 4 through 7) have been published in other journals. Each previously published chapter investigates a different genomics model, but all focus on identifying the strengths and limitations of various statistical procedures that have been applied to different GWAS scenarios.
Translating in vitro biological data into actionable information related to human health holds the potential to improve disease treatment and risk assessment of chemical exposures. While genomics has identified regulatory pathways at the cellular level, translation to the organism level requires a multiscale approach accounting for intra-cellular regulation, inter-cellular interaction, and tissue/organ-level effects. Tissue-level effects can now be probed in vitro thanks to recently developed systems of three-dimensional (3D), multicellular, "organotypic" cell cultures, which mimic functional responses of living tissue. However, there remains a knowledge gap regarding interactions across different biological scales, complicating accurate prediction of health outcomes from molecular/genomic data and tissue responses. Systems biology aims at mathematical modeling of complex, non-linear biological systems. We propose to apply a systems biology approach to achieve a computational representation of tissue-level physiological responses by integrating empirical data derived from organotypic culture systems with computational models of intracellular pathways to better predict human responses. Successful implementation of this integrated approach will provide a powerful tool for faster, more accurate and cost-effective screening of potential toxicants and therapeutics. On September 11, 2015, an interdisciplinary group of scientists, engineers, and clinicians gathered for a workshop in Research Triangle Park, North Carolina, to discuss this ambitious goal. Participants represented laboratory-based and computational modeling approaches to pharmacology and toxicology, as well as the pharmaceutical industry, government, non-profits, and academia. Discussions focused on identifying critical system perturbations to model, the computational tools required, and the experimental approaches best suited to generating key data.
Background: Influenza vaccination is administered throughout the influenza disease season, even as late as March. Given such timing, what is the value of vaccinating the population earlier than currently being practiced?Methods: We used real data on when individuals were vaccinated in Allegheny County, Pennsylvania, and the following 2 models to determine the value of vaccinating individuals earlier (by the end of September, October, and November): Framework for Reconstructing Epidemiological Dynamics (FRED), an agent-based model (ABM), and FluEcon, our influenza economic model that translates cases from the ABM to outcomes and costs [health care and lost productivity costs and quality-adjusted life-years (QALYs)]. We varied the reproductive number (R-0) from 1.2 to 1.6.Results: Applying the current timing of vaccinations averted 223,761 influenza cases, $16.3 million in direct health care costs, $50.0 million in productivity losses, and 804 in QALYs, compared with no vaccination (February peak, R-0 1.2). When the population does not have preexisting immunity and the influenza season peaks in February (R-0 1.2-1.6), moving individuals who currently received the vaccine after September to the end of September could avert an additional 9634-17,794 influenza cases, $0.6-$1.4 million in direct costs, $2.1-$4.0 million in productivity losses, and 35-64 QALYs. Moving the vaccination of just children to September (R-0 1.2-1.6) averted 11,366-1660 influenza cases, $0.6-$0.03 million in direct costs, $2.3-$0.2 million in productivity losses, and 42-8 QALYs. Moving the season peak to December increased these benefits, whereas increasing preexisting immunity reduced these benefits.Conclusion: Even though many people are vaccinated well after September/October, they likely are still vaccinated early enough to provide substantial cost-savings.
Many published influenza models treat each simulation day as a weekday and do not distinguish weekend days. Consequently, the weekend effect on influenza transmission has not been fully explored. To assess whether distinguishing between weekday and weekend transmissions in simulation models of flu-like infectious disease models matters, this study uses an agent-based model of the Chicago, Illinois metropolitan area. Our study assesses whether including weekend effects is offset by increases in weekend contact patterns and if implementing 3-day weekends dampens disease transmission enough to warrant its use as a containment strategy. Results indicate that ignoring weekend behaviors without incorporating increases in community-based non-school contacts (i.e., compensatory behaviors) causes the peak case incidence day to occur 7 days earlier and can reduce the peak attack rate by as much as 60 %. These results are sensitive to the proportion of symptomatic cases that are assumed to remain at home until they recover. The 3-day weekend intervention has interesting possibilities, but the benefits may only be effective for mild epidemics. However, a 3-day weekend for schools would be less detrimental to the educational process than sustained permanent closing because student and teacher contact is maintained throughout the epidemic period. Also, a 4-day school and work week may be more easily accommodated by many types of schools and businesses. On the other hand, an additional day per week of school closure could result in substantial societal costs, with lost productivity and child care costs outweighing the savings of preventing influenza cases.
Objectives: From 2003 to 2013, RTI International served as the data repository for the National Institute of Diabetes, Digestive and Kidney Diseases (NIDDK). RTI worked closely with two sample repository partners to build and maintain the Central Repository (CR) that made data and samples available to approved requestors. In this paper, we recap aspects of establishing the mechanism; detail the challenges and limitations of data and sample sharing, and explore the future of resource sharing in light of the evolving environment of research funding.Design and methods: Effective maintenance required the system to be flexible and dynamic while at the same time compliant with established data standards.Results: Our years serving as the CR for NIDDK have yielded a number of observations about the difficulties of running a repository, an operation that is by definition dependent on many outside parties whose degree of expertise and efficiency have a direct impact on repository functioning.Conclusion: The bio-banking industry will likely continue to become more globally centralized for studying specific genetic diseases and monitoring the health of our environment. The dynamic relationship between emerging technologies and the infrastructure will be needed to support future research that requires the ability of organizations providing support to remain flexible even while following established standards. (C) 2013 The Canadian Society of Clinical Chemists. Published by Elsevier Inc. All rights reserved.
The National Institute of Diabetes and Digestive Disease (NIDDK) Central Data Repository (CDR) is a web-enabled resource available to researchers and the general public. The CDR warehouses clinical data and study documentation from NIDDK funded research, including such landmark studies as The Diabetes Control and Complications Trial (DCCT, 1983-93) and the Epidemiology of Diabetes Interventions and Complications (EDIC, 1994-present) follow-up study which has been ongoing for more than 20 years. The CDR also houses data from over 7 million biospecimens representing 2 million subjects. To help users explore the vast amount of data stored in the NIDDK CDR, we developed a suite of search mechanisms called the public query tools (PQTs). Five individual tools are available to search data from multiple perspectives: study search, basic search, ontology search, variable summary and sample by condition. PQT enables users to search for information across studies. Users can search for data such as number of subjects, types of biospecimens and disease outcome variables without prior knowledge of the individual studies. This suite of tools will increase the use and maximize the value of the NIDDK data and biospecimen repositories as important resources for the research community. Database URL:https://www.niddkrepository.org/niddk/home.do
Forecasting Populations (FPOP) is a microsimulation model (MSM) that is the demographic core of an extensible modeling framework. The framework, with FPOP at its core, enables the geospatial projection of a population under purely demographic processes or under the additional influence of exogenous factors such as disease, policy changes and prevention programs, or environmental stressors. Empirically-derived transition probabilities of life events such as birth, death, marriage, divorce and migration, captured in lookup table format, drive the simulation. These transition probabilities can be modified dynamically by external user-defined functions or other external MSMs. The use of MSM structures and methodologies enables FPOP to portray the impact of heterogeneity in the geospatial dimension (e.g., distribution of environmental factors or distribution of intervention programs), as well as the social dimension (e.g., household or social network correlates), on the projections. POP is designed and structured
BACKGROUND:Mathematical and computational models provide valuable tools that help public health planners to evaluate competing health interventions, especially for novel circumstances that cannot be examined through observational or controlled studies, such as pandemic influenza. The spread of diseases like influenza depends on the mixing patterns within the population, and these mixing patterns depend in part on local factors including the spatial distribution and age structure of the population, the distribution of size and composition of households, employment status and commuting patterns of adults, and the size and age structure of schools. Finally, public health planners must take into account the health behavior patterns of the population, patterns that often vary according to socioeconomic factors such as race, household income, and education levels.RESULTS:FRED (a Framework for Reconstructing Epidemic Dynamics) is a freely available open-source agent-based modeling system based closely on models used in previously published studies of pandemic influenza. This version of FRED uses open-access census-based synthetic populations that capture the demographic and geographic heterogeneities of the population, including realistic household, school, and workplace social networks. FRED epidemic models are currently available for every state and county in the United States, and for selected international locations.CONCLUSIONS:State and county public health planners can use FRED to explore the effects of possible influenza epidemics in specific geographic regions of interest and to help evaluate the effect of interventions such as vaccination programs and school closure policies. FRED is available under a free open source license in order to contribute to the development of better modeling tools and to encourage open discussion of modeling tools being used to evaluate public health policies. We also welcome participation by other researchers in the further development of FRED.