To characterize somatic alterations in colorectal carcinoma, we conducted a genome-scale analysis of 276 samples, analysing exome sequence, DNA copy number, promoter methylation and messenger RNA and microRNA expression. A subset of these samples (97) underwent low-depth-of-coverage whole-genome sequencing. In total, 16% of colorectal carcinomas were found to be hypermutated: three-quarters of these had the expected high microsatellite instability, usually with hypermethylation and MLH1 silencing, and one-quarter had somatic mismatch-repair gene and polymerase e (POLE) mutations. Excluding the hypermutated cancers, colon and rectum cancers were found to have considerably similar patterns of genomic alteration. Twenty-four genes were significantly mutated, and in addition to the expected APC, TP53, SMAD4, PIK3CA and KRAS mutations, we found frequent mutations in ARID1A, SOX9 and FAM123B. Recurrent copy-number alterations include potentially drug-targetable amplifications of ERBB2 and newly discovered amplification of IGF2. Recurrent chromosomal translocations include the fusion of NAV2 and WNT pathway member TCF7L1. Integrative analyses suggest new markers for aggressive colorectal carcinoma and an important role for MYC-directed transcriptional activation and repression.
We present a bacterial genome computational analysis pipeline, called GenVar. The pipeline, based on the program GeneWise, is designed to analyze an annotated genome and automatically identify missed gene calls and sequence variants such as genes with disrupted reading frames (split genes) and those with insertions and deletions (indels). For a given genome to be analyzed, GenVar relies on a database containing closely related genomes (such as other species or strains) as well as a few additional reference genomes. GenVar also helps identify gene disruptions probably caused by sequencing errors. We exemplify GenVar's capabilities by presenting results from the analysis of four Brucella genomes. Brucella is an important human pathogen and zoonotic agent. The analysis revealed hundreds of missed gene calls, new split genes and indels, several of which are species specific and hence provide valuable clues to the understanding of the genome basis of Brucella pathogenicity and host specificity.
The PathoSystems Resource Integration Center (PATRIC) is one of eight Bioinformatics Resource Centers (BRCs) funded by the National Institute of Allergy and Infection Diseases (NIAID) to create a data and analysis resource for selected NIAID priority pathogens, specifically proteobacteria of the genera Brucella, Rickettsia and Coxiella, and corona-, calici- and lyssaviruses and viruses associated with hepatitis A and E. The goal of the project is to provide a comprehensive bioinformatics resource for these pathogens, including consistently annotated genome, proteome and metabolic pathway data to facilitate research into counter-measures, including drugs, vaccines and diagnostics. The project's curation strategy has three prongs: 'breadth first' beginning with whole-genome and proteome curation using standardized protocols, a 'targeted' approach addressing the specific needs of researchers and an integrative strategy to leverage high-throughput experimental data (e.g. microarrays, proteomics) and literature. The PATRIC infrastructure consists of a relational database, analytical pipelines and a website which supports browsing, querying, data visualization and the ability to download raw and curated data in standard formats. At present, the site warehouses complete sequences for 17 bacterial and 332 viral genomes. The PATRIC website (https://patric.vbi.vt.edu) will continually grow with the addition of data, analysis and functionality over the course of the project.
This paper presents the eleventh update of the human obesity gene map, which incorporates published results up to the end of October 2004. Evidence from single-gene mutation obesity cases, Mendelian disorders exhibiting obesity as a clinical feature, transgenic and knockout murine models relevant to obesity, quantitative trait loci (QTLs) from animal cross-breeding experiments, association studies with candidate genes, and linkages from genome scans is reviewed. As of October 2004, 173 human obesity cases due to single-gene mutations in 10 different genes have been reported, and 49 loci related to Mendelian syndromes relevant to human obesity have been mapped to a genomic region, and causal genes or strong candidates have been identified for most of these syndromes. There are 166 genes which, when mutated or expressed as transgenes in the mouse, result in phenotypes that affect body weight and adiposity. The number of QTLs reported from animal models currently reaches 221. The number of human obesity QTLs derived from genome scans continues to grow, and we have now 204 QTLs for obesity-related phenotypes from 50 genome-wide scans. A total of 38 genomic regions harbor QTLs replicated among two to four studies. The number of studies reporting associations between DNA sequence variation in specific genes and obesity phenotypes has also increased considerably with 358 findings of positive associations with 113 candidate genes. Among them, 18 genes are supported by at least five positive studies. The obesity gene map shows putative loci on all chromosomes except Y. Overall, > 600 genes, markers, and chromosomal regions have been associated or linked with human obesity phenotypes. The electronic version of the map with links to useful publications and genornic and other relevant sites can be found at http:// obesitygene.pbrc.edu.
Method for converting a group of important parameters that specify a device in a first set of import data, the method comprising in a set of output parameters that specify the device in a first database: Receiving the set of import parameters from an import file that contains the first set of import data in a variety of import records, each import record contains a variety of import values, where at each import value an import parameters from the group of import parameters corresponds; Receiving the set of output parameters from the first database; and Generating a first implementation of the set of import parameters in the set of output parameters.
This is the ninth update of the human obesity gene map, incorporating published results through October 2002 and continuing the previous format. Evidence from single-gene mutation obesity cases, Mendelian disorders exhibiting obesity as a clinical feature, quantitative trait loci (QTLs) from human genome-wide scans and various animal crossbreeding.. experiments, and association and linkage studies with candidate genes and other markers is reviewed. For the first time, transgenic and knockout murine models exhibiting obesity as a phenotype are incorporated (N = 38). As of October 2002, 33 Mendelian syndromes relevant to human obesity have been mapped to a genomic region, and the causal genes or strong candidates have been identified for 23 of these syndromes. QTLs reported from animal models currently number 168; there are 68 human QTLs for obesity phenotypes from genome-wide scans. Additionally, significant linkage peaks with candidate genes have been identified in targeted studies. Seven genomic regions harbor QTLs replicated among two to five studies. Attempts to relate DNA sequence variation in specific genes to obesity phenotypes continue to grow, with 222 studies reporting positive associations with 71 candidate genes. Fifteen such candidate genes are supported by at least five positive studies. The obesity gene map shows putative loci on all chromosomes except Y. More than 300 genes, markers, and chromosomal regions have been associated or linked with human obesity phenotypes. The electronic version of the map with links to useful sites can be found at http://obesitygene.pbre.edu.
Physical exercise produces several adaptive changes in skeletal muscle. However, the molecular mechanisms of these effects are poorly understood. We performed serial analysis of gene expression (SAGE) to quantify the global gene expression profile in sedentary and endurance-trained muscle. A total of 10869 SAGE tags was sequenced and represented 4727 genes. The genes most expressed in muscle are mainly involved in contraction and energy metabolism. Thirty-three genes were differentially expressed between endurance athletes and sedentary individuals. Four genes such as myosin binding protein C fast-type, glycogen phosphorylase, and pyruvate kinase were expressed less in endurance athletes, whereas eight genes coding for expressed sequence tag similar to (EST) crystallin alpha B, EST myosin light chain 2, EST surfactant pulmonary-associated protein A1, EST thrombospondin, EST fructose-bisphosphate aldolase A, EST cytochrome oxidase 1, NADH dehydrogenase 3, and G8 protein were up-regulated. Most of the upregulated tags corresponded to novel genes. On the other hand, different isoforms of fructose-bisphosphate aldolase A were also differentially expressed. The current study underlying the most highly expressed genes allows a better understanding of global muscle characteristics in normal and endurance-trained individuals. Moreover, the current data suggest novel candidate genes that may be responsible for enhanced endurance performance.
The prevalence of mutations within and in the flanking regions of the gene encoding the melanocortin 4 receptor was investigated in severely obese and normal-weight subjects from the Swedish Obese Subjects study, the Health, Risk Factors, Exercise Training, and Genetics (HERITAGE) Family study, and a Memphis cohort. A total of 433 white and 95 black subjects (94% females) were screened for mutations by direct sequencing. Three previously described missense variants and nine novel (three missense, six silent) variants were detected. None of them showed significant association with obesity or related phenotypes. In addition, two novel deletions were found in two heterozygous obese women: a -65_-64delTG mutation within the 5' noncoding region and a 171delC frameshift mutation predicted to result in a truncated nonfunctional receptor. No pathogenic mutations were found among obese blacks or nonobese controls. Furthermore, none of the null mutations found in other populations was present in this sample. In conclusion, our results do not support the prevailing notion that sequence variation in the melanocortin 4 receptor gene is a frequent cause of human obesity.
Ghrelin and preproghrelin sequences were determined in 96 unrelated female subjects with severe obesity (mean body mass index (BMI) 42.3 +/- 3.4 kg/m(2)) and in 96 non-obese female controls (mean BMI 23.0 +/- 1.4 (kg/m2) of the Swedish Obese Subjects cohort. A mutation at amino acid position 51 (Arg51Gln) of the preproghrelin sequence that corresponds to the last amino acid in mature ghrelin product was identified in six (all heterozygotes) obese subjects (6.3%) but not among controls (p < 0.05). The self-reported weight at 20, 30, and 40 years of age tended to be 7.5, 4.7 and 6.4 kg lower, respectively, among obese Gln allele carriers versus obese non-carriers. In addition, a mutation at codon 72 of the preproghrelin gene (Leu72Met) was detected in 15 obese (12 hetero- and 3 homozygotes) and 12 control (all heterozygotes) subjects. This mutation outside the coding region of the mature ghrelin product tended to be associated with lower age of self-reported onset of obesity (15.6 +/- 7.9 vs. 20.5 +/- 10.5 years; p = 0.09). In addition to these two mutations in coding regions, a G274A base change in a non-coding region between exons one and two was found only in two obese individuals. The Arg51Gln amino acid substitution may alter the cleavage site of endoproteases and the length of the mature ghrelin product. The functional significance of the Leu72Met mutation and a G274A base change remains to be determined. In conclusion, the data provide evidence that a low frequency sequence variation in the ghrelin gene could play a role in the etiology of obesity.
We have developed a computer program, GeneParser, which identifies and determines the fine structure of protein genes in genomic DNA sequences. The program scores all subintervals in a sequence for content statistics indicative of introns and exons, and for sites that identify their boundaries. This information is weighted by a neural network to approximate the log-likelihood that each subinterval exactly represents an intron or exon (first, internal or last). A dynamic programming algorithm is then applied to this data to find the combination of introns and exons that maximizes the likelihood function. Using this method, we can rapidly generate ranked suboptimal solutions, each of which is the optimum solution containing a given intron-exon junction. We have tested the system on a large collection of human genes. On sequences not used in training, we achieved a correlation coefficient for exon nucleotide prediction of 0.89. For a subset of G + C-rich genes, a correlation coefficient of 0.94 was achieved. We have also quantified the robustness of the method to substitution and frame-shift errors and show how the system can be optimized for performance on sequences with known levels of sequencing errors.
High-affinity RNA ligands were generated against intact 30S ribosomes, S1-depleted 30S ribosomes, and purified ribosomal protein S1. Sequence analysis indicated two classes of ligand: unstructured RNAs containing a Shine-Dalgarno sequence and structured RNAs containing a pseudoknot. The Shine-Dalgarno-containing Ligands were generated against S1-depleted 30S ribosomes but, surprisingly, not against intact 30S ribosomes or ribosomal protein S1. In contrast, pseudoknot-containing ligands were generated against intact ribosomes as well as purified S1 protein. The two classes of ligand exhibited specificity for their respective targets, as well as conserved sequence and secondary structure reminiscent of naturally occurring, cis-acting mRNA elements.
Dynamic programming (DP) is applied to the problem of precisely identifying internal exons and introns in genomic DNA sequences. The program GeneParser first scores the sequence of interest for splice sites and for these intron- and exon-specific content measures: codon usage, local compositional complexity, 6-tuple frequency, length distribution and periodic asymmetry. This information is then organized for interpretation by DP. GeneParser employs the DP algorithm to enforce the constraints that introns and exons must be adjacent and non-overlapping and finds the highest scoring combination of introns and exons subject to these constraints. Weights for the various classification procedures are determined by training a simple feed-forward neural network to maximize the number of correct predictions. In a pilot study, the system has been trained on a set of 56 human gene fragments containing 150 internal exons in a total of 158,691 bps of genomic sequence. When tested against the training data, GeneParser precisely identifies 75% of the exons and correctly predicts 86% of coding nucleotides as coding while only 13% of non-exon bps were predicted to be coding. This corresponds to a correlation coefficient for exon prediction of 0.85. Because of the simplicity of the network weighting scheme, generalization performance is nearly as good as with the training set.
The Escherichia coli D-galactose and D-glucose receptor, an aqueous periplasmic receptor that triggers sugar sensing and transport, Possesses a single Ca2+ binding site similar in structure and specificity to the EF-hand class of sites found in eukaryotic Ca2+ signaling proteins including calmodulin and its homologues. A universal feature of these sites is the use of a pentagonal bipyramidal array of seven oxygens to coordinate bound Ca2+. Here we investigate the mechanisms used by this coordinating array to control ion specificity. To vary the cavity size and charge of the array, we have replaced axial glutamine 142 in the prokaryotic site with asparagine, glutamate, and aspartate. The ion selectivities of the resulting engineered sites have been quantitated by measuring dissociation constants for a series of spherical metal ions, differing in increments of radius and charge, from groups Ia, IIa, and IIIa and the lanthanides. Dramatic specificity changes are observed: sites containing an engineered smaller side chain (Asn or Asp) bind the largest cations up to 50-fold more tightly than the native site; and sites containing an engineered negative side chain (Glu or Asp) exhibit preferences for trivalent over divalent cations up to 1900-fold higher than the native site. The results indicate that the cavity size and negative charge of the coordination array play key roles in selective Ca2+ binding and that the array can be engineered to preferentially bind other cations.
The molecular mechanisms by which protein Ca(II) sites selectively bind Ca(II) even in the presence of high concentrations of other metals, particularly Na(I), K(I), and Mg(II), have not been fully described. The single Ca(II) site of the Escherichia coli receptor for D-galactose and D-glucose (GGR) is structurally related to the eukaryotic EF-hand Ca(II) sites and is ideally suited as a model for understanding the structural and electrostatic basis of Ca(II) specificity. Metal binding to the bacterial site was monitored by a Tb(III) phosphorescence assay: Ca(II) in the site was replaced with Tb(III), which was then selectively excited by energy transfer from protein tryptophans. Photons emitted from the bound Tb(III) enabled specific detection of this substrate; for other metals binding was detected by competitive displacement of Tb(III). Representative spherical metal ions from groups IA, IIA, and IIIA and the lanthanides were chosen to study the effects of metal ion size and charge on the affinity of metal binding. A dissociation constant was measured for each metal, yielding a range of KD's spanning over 6 orders of magnitude. Monovalent metal ions of group IA exhibited very low affinities. Divalent group IIA metal ions exhibited affinities related to their size, with optimal binding at an effective ionic radius between those of Mg(II) (0.81 A) and Ca(II) (1.06 A). Trivalent metal ions of group IIIA and the lanthanides also exhibited size-dependent affinities, with an optimal effective ionic radius between those of Sc(III) (0.81 A) and Yb(III) (0.925 A). The results indicate that the GGR site selects metal ions on the basis of both charge and size.(ABSTRACT TRUNCATED AT 250 WORDS)
e Cyberinfrastructure Group (CIG) develops and uses methods, infrastructure, and resources to enable scientifi c discoveries in infectious disease research by applying the principles of cyberinfrastructure to integrate data, computational infrastructure, and people (Atkins, 2003). CIG has developed many public resources for curated, diverse molecular and literature data from various infectious disease systems, and implemented the processes, systems, and databases required to support them. It also conducts research applying its methods, infrastructure and data to make new discoveries of its own. CIG participates in education and outreach activities, resulting in scientifi c discoveries and publications, and an outreach program involving development of project-centric cyberinfrastructure courses for the educators from high schools and undergraduate institutions as well as graduates and postgraduates. In the reporting period, key accomplishments include publication of the Brucella abortus S19 genome, deployment of a pipeline to improve genome annotations, phylogenomic analysis of ten rickettsial genomes, publication of an Alphaproteobacteria phylogenetic tree, development of a program to generate oligonucleotide sequences from whole genome sequences, continuing progress in numerous collaborative research projects, and development of an online self-guided bioinformatics tutorial.