There are bioethical, institutional, economic, legal, and cultural obstacles to creating the robust-precompetitive-data resource that will be required to advance the vision of “precision medicine,” the ability to use molecular data to target therapies to patients for whom they offer the most benefit at the least risk. Creation of such an “information commons” was the central recommendation of the 2011 report Toward Precision Medicine issued by a committee of the National Research Council of the USA (Committee on a Framework for Development of a New Taxonomy of Disease; National Research Council. Toward precision medicine: building a knowledge network for biomedical research and a new taxonomy of disease. 2011). In this commentary, I review the rationale for creating an information commons and the obstacles to doing so; then, I endorse a path forward based on the dynamic consent of research subjects interacting with researchers through trusted mediators. I assert that the advantages of the proposed system overwhelm alternative ways of handling data on the phenotypes, genotypes, and environmental exposures of individual humans; hence, I argue that its creation should be the central policy objective of early efforts to make precision medicine a reality.
>Let me start with a few words about the goals and intended audience for this introduction to republication of two reports of the U.S.National Research Council(NRC)–Mapping and Sequencing the Human Genome(1988)and Toward Precision Medicine(2011).I do not intend to summarize what the
Claims data alone are typically lacking in key clinical information capturing disease severity, specificity of diagnosis, patient reported outcomes, and many other factors. Attempts to estimate treatment effects with claims data alone, with missing key control variables, have a greater likelihood to produce biased result. This study uses claims and lab test results data from a large national health plan, in an attempt to circumvent or mitigate problem of biases in treatment effect estimators attributed to missing key clinical variables. A linked dataset of claims and laboratory results is used to estimate treatment effect on resource utilization for a Hepatitis C (HCV) sample (N=2031). To empirically assess the bias from omitting the clinical information in the laboratory results (APRI scores), treatment effects are estimated using claims alone. Laboratory results are then attached to the claims through statistical linkages of records from clinical records and claims records, simulating the statistical linkage of medical datasets in situations where direct patient-level linkages are not possible. Estimation of treatment effects were calculated through the bootstrap method by running negative binomial count regressions 500 times in samples with replacement size of 500. The results from the dataset with the linked laboratory results showed bias reduction of 78.8% and 23.8% in HCV-related ambulatory visits and All-cause ambulatory visits, respectively. Statistical linkages do not appear to have any impact on the treatment effect estimators in the ER visit outcomes which are confirmed by the lack of correlations in APRI scores and outcomes. When confronted with missing key clinical variables, record linkage has the effect of reducing bias in treatment effects. The degree of this bias reduction will be a function of the strength of the correlations among the missing variable, the treatment variable, and the outcome variable.
Fulfilling the promise of the genetic revolution requires the analysis of large datasets containing information from thousands to millions of participants. However, sharing human genomic data requires protecting subjects from potential harm. Current models rely on de-identification techniques in which privacy versus data utility becomes a zero-sum game. Instead, we propose the use of trust-enabling techniques to create a solution in which researchers and participants both win. To do so we introduce three principles that facilitate trust in genetic research and outline one possible framework built upon those principles. Our hope is that such trust-centric frameworks provide a sustainable solution that reconciles genetic privacy with data sharing and facilitates genetic research.
The role of the E3 ubiquitin ligase murine double minute 2 (Mdm2) in regulating the stability of the p53 tumor suppressor is well documented. By contrast, relatively little is known about p53-independent activities of Mdm2 and the role of Mdm2 in cellular differentiation. Here we report a novel role for Mdm2 in the initiation of adipocyte differentiation that is independent of its ability to regulate p53. We show that Mdm2 is required for cAMP-mediated induction of CCAAT/enhancer-binding protein δ (C/EBP δ ) expression by facilitating recruitment of the cAMP regulatory element-binding protein (CREB) coactivator, CREB-regulated transcription coactivator (Crtc2)/TORC2, to the c/ebpδ promoter. Our findings reveal an unexpected role for Mdm2 in the regulation of CREB-dependent transactivation during the initiation of adipogenesis. As Mdm2 is able to promote adipogenesis in the myoblast cell line C2C12, it is conceivable that Mdm2 acts as a switch in cell fate determination.
A VME-based data acquisition system for beam-loss monitors has been developed and is in use in the Tevatron and Main Injector accelerators at the Fermilab complex. The need for enhanced beam-loss protection when the Tevatron is operating in collider-mode was the main driving force for the new design. Prior to the implementation of the present system, the beam-loss monitor system was disabled during collider operation and protection of the Tevatron magnets relied on the quench protection system. The new Beam-Loss Monitor system allows appropriate abort logic and thresholds to be set over the full set of collider operating conditions. The system also records a history of beam-loss data prior to a beam-abort event for post-abort analysis. Installation of the Main Injector system occurred in the fall of 2006 and the Tevatron system in the summer of 2007. Both systems were fully operation by the summer of 2008. In this paper we report on the overall system design, provide a description of its normal operation, and show a number of examples of its use in both the Main Injector and Tevatron.
It is well known that average levels of population structure are higher on the X chromosome compared to autosomes in humans. However, there have been surprisingly few analyses on the spatial distribution of population structure along the X chromosome. With publicly available data from the HapMap Project and Perlegen Sciences, we show a strikingly punctuated pattern of X chromosome population structure. Specifically, 87% of X-linked HapMap SNPs within the top 1% of F(ST) values cluster into five distinct loci. The largest of these regions spans 5.4 Mb and contains 66% of the most highly differentiated HapMap SNPs on the X chromosome. We demonstrate that the extreme clustering of highly differentiated SNPs on the X chromosome is not an artifact of ascertainment bias, nor is it specific to the populations genotyped in the HapMap Project. Rather, additional analyses and resequencing data suggest that these five regions have been substrates of recent and strong adaptive evolution. Finally, we discuss the implications that patterns of X-linked population structure have on the evolutionary history of African populations.
Using next-generation sequencing technology alone, we have successfully generated and assembled a draft sequence of the giant panda genome. The assembled contigs (2.25 gigabases (Gb)) cover approximately 94% of the whole genome, and the remaining gaps (0.05 Gb) seem to contain carnivore-specific repeats and tandem repeats. Comparisons with the dog and human showed that the panda genome has a lower divergence rate. The assessment of panda genes potentially underlying some of its unique traits indicated that its bamboo diet might be more dependent on its gut microbiome than its own genetic composition. We also identified more than 2.7 million heterozygous single nucleotide polymorphisms in the diploid genome. Our data and analyses provide a foundation for promoting mammalian genetic research, and demonstrate the feasibility for using next-generation sequencing technologies for accurate, cost-effective and rapid de novo assembly of large eukaryotic genomes.
Background Methylotrophy describes the ability of organisms to grow on reduced organic compounds without carbon-carbon bonds. The genomes of two pink-pigmented facultative methylotrophic bacteria of the Alpha-proteobacterial genus Methylobacterium, the reference species Methylobacterium extorquens strain AM1 and the dichloromethane-degrading strain DM4, were compared. Methodology/Principal Findings The 6.88 Mb genome of strain AM1 comprises a 5.51 Mb chromosome, a 1.26 Mb megaplasmid and three plasmids, while the 6.12 Mb genome of strain DM4 features a 5.94 Mb chromosome and two plasmids. The chromosomes are highly syntenic and share a large majority of genes, while plasmids are mostly strain-specific, with the exception of a 130 kb region of the strain AM1 megaplasmid which is syntenic to a chromosomal region of strain DM4. Both genomes contain large sets of insertion elements, many of them strain-specific, suggesting an important potential for genomic plasticity. Most of the genomic determinants associated with methylotrophy are nearly identical, with two exceptions that illustrate the metabolic and genomic versatility of Methylobacterium. A 126 kb dichloromethane utilization (dcm) gene cluster is essential for the ability of strain DM4 to use DCM as the sole carbon and energy source for growth and is unique to strain DM4. The methylamine utilization (mau) gene cluster is only found in strain AM1, indicating that strain DM4 employs an alternative system for growth with methylamine. The dcm and mau clusters represent two of the chromosomal genomic islands (AM1: 28; DM4: 17) that were defined. The mau cluster is flanked by mobile elements, but the dcm cluster disrupts a gene annotated as chelatase and for which we propose the name “island integration determinant” (iid). Conclusion/Significance These two genome sequences provide a platform for intra- and interspecies genomic comparisons in the genus Methylobacterium, and for investigations of the adaptive mechanisms which allow bacterial lineages to acquire methylotrophic lifestyles.
The application of new technology to sequence the genome of an individual yields few biological insights. Nonetheless, the feat heralds an era of 'personal genomics' based on cheap sequencing.
Large-insert genome analysis (LIGAN) is a broadly applicable, high-throughput technology designed to characterize genome-scale structural variation. Fosmid paired-end sequences and DNA fingerprints from a query genome are compared to a reference sequence using the Genomic Variation Analysis (GenVal) suite of software tools to pinpoint locations of insertions, deletions, and rearrangements. Fosmids spanning regions that contain new structural variants can then be sequenced. Clonal pairs of Pseudomonas aeruginosa isolates from four cystic fibrosis patients were used to validate the LIGAN technology. Approximately 1.5 Mb of inserted sequences were identified, including 743 kb containing 615 ORFs that are absent from published P. aeruginosa genomes. Six rearrangement breakpoints and 220 kb of deleted sequences were also identified. Our study expands the "genome universe" of P. aeruginosa and validates a technology that complements emerging, short-read sequencing methods that are better suited to characterizing single-nucleotide polymorphisms than structural variation.
Genetic variation among individual humans occurs on many different scales, ranging from gross alterations in the human karyotype to single nucleotide changes. Here we explore variation on an intermediate scale--particularly insertions, deletions and inversions affecting from a few thousand to a few million base pairs. We employed a clone-based method to interrogate this intermediate structural variation in eight individuals of diverse geographic ancestry. Our analysis provides a comprehensive overview of the normal pattern of structural variation present in these genomes, refining the location of 1,695 structural variants. We find that 50% were seen in more than one individual and that nearly half lay outside regions of the genome previously described as structurally variant. We discover 525 new insertion sequences that are not present in the human reference genome and show that many of these are variable in copy number between individuals. Complete sequencing of 261 structural variants reveals considerable locus complexity and provides insights into the different mutational processes that have shaped the human genome. These data provide the first high-resolution sequence map of human structural variation--a standard for genotyping platforms and a prelude to future individual genome sequencing projects.
Methods relying on dense arrays of synthetic oligodeoxynucleotides to target specific subsets of the human genome may enable routine resequencing of all human exons or multi-megabase-pair chromosomal regions.
The contribution of large-scale and intermediate-size structural variation (ISV) to human genetic disease and disease susceptibility is only beginning to be understood. The development of high-throughput genotyping technologies is one of the most critical aspects for future studies of linkage disequilibrium (LD) and disease association. Using a simple PCR-based method designed to assay the junctions of the breakpoints, we genotyped seven simple insertion and deletion polymorphisms ranging in size from 6.3 to 24.7 kb among 90 CEPH individuals. We then extended this analysis to a larger collection of samples (n=460) by application of an oligonucleotide extension-ligation genotyping assay. The analysis showed a high level of concordance ( approximately 99%) when compared with PCR/sequence-validated genotypes. Using the available HapMap data, we observed significant LD (r2=0.74-0.95) between each ISV and flanking single nucleotide polymorphisms, but this observation is likely to hold only for similar simple insertion/deletion events. The approach we describe may be used to characterize a large number of individuals in a cost-effective manner once the sequence organization of ISVs is known.
Traditional methods for identifying minor histocompatibility antigens (mHags) are technically challenging and biased against discovery of mHags not expressed in the peripheral blood. In this work, we propose a rapid, unbiased, genetic approach for identification minor antigens resulting from disparities in coding non-synonymous SNPs (“C SNPS”). This approach is capable of testing for responses to candidate minor antigens expressed in virtually any tissue, including those expressed exclusively in tissues targeted by GVHD. The first step in our approach begins with comparison of donor and recipient C SNP genotypes generated using C SNP microarrays. These arrays interrogate approximately 80% of human C SNPs predicted to occur in greater than 5% of the population. Comparison of C SNP genotypes directly identifies protein-altering alleles present in the recipient but not the donor (hereafter referred to as “recipient-restricted” alleles), thereby identifying a transplant-specific set of candidate minor antigens. The second step utilizes conventional HLA-class I epitope prediction performed on all linear peptides that include amino-acid residues defined by a recipient-restricted allele. This two step filtering process identifies a small, “testable” number candidate minor peptide epitopes for an individual expressed HLA-class I allele. Candidate epitopes can then be synthesized, pooled and tested at diagnosis of GVHD using a commercial Granzyme B ELISPOT assay. As proof-of-principal for a direct genotyping approach, we have analyzed C SNP genotypes, performed epitope prediction and generated T-cell lines specific for candidate minor antigens using DNA and PBL from a pair of disease-free HLA-identical siblings. Analysis of 10,000 C SNPs shows that approximately 2,000 C SNP alleles are restricted to one sibling within the pair. BIMAS epitope prediction of short unique peptide sequences determined by each sibling-restricted allele identifies approximately 100 candidate minor epitopes predicted to bind HLA-A*0201. These candidate minor epitopes were ranked using expression microarrays performed on EBV-transformed LCL derived from each sibling, with candidates derived from highly expressed genes ranked above those from genes with lower expression levels. A pool of 12 candidate minor epitopes that were both unique to sibling “A” and derived from genes highly expressed in LCL were synthesized and used to generate CD8+ T-cell lines from sibling “B”. Stimulation utilized autologous (sibling “B”-derived) mature dendritic cells loaded with candidate minor epitopes. After several rounds of in-vitro stimulation, each T-cell line was tested for responses to EBV-LCL from sibling “B” and sibling “A” using a Granzyme B ELISPOT kit. Five out of sixteen lines responded to LCL from sibling A while no line responded to autologous LCL. Thus we show that this approach frequently generates CD8+ T-cell lines specific for sibling-derived target cells, suggesting that this approach efficiently identifies genuine, novel, endogenously processed and presented minor epitopes. Deconvolution of the peptide pool suggests that at least two out of the twelve candidate minor epitopes are naturally processed and presented.
The reference sequence for each human chromosome provides the framework for understanding genome function, variation and evolution. Here we report the finished sequence and biological annotation of human chromosome 1. Chromosome 1 is gene-dense, with 3,141 genes and 991 pseudogenes, and many coding sequences overlap. Rearrangements and mutations of chromosome 1 are prevalent in cancer and many other diseases. Patterns of sequence variation reveal signals of recent selection in specific genes that may contribute to human fitness, and also in regions where no function is evident. Fine-scale recombination occurs in hotspots of varying intensity along the sequence, and is enriched near genes. These and other studies of human biology and disease encoded within chromosome 1 are made possible with the highly accurate annotated sequence, as part of the completed set of chromosome sequences that comprise the reference human genome.