646Aneka is an Application Platform-as-a-Service (Aneka PaaS) for Cloud Computing. It acts as a framework for building customized applications and deploying them on either public or private Clouds. One of the key features of Aneka is its support for provisioning resources on different public Cloud providers such as Amazon EC2, Windows Azure and GoGrid. In this chapter, we will present the Aneka platform and its integration with one of the public Cloud infrastructures, Windows Azure, which enables the usage of Windows Azure Compute Service as a resource provider of Aneka PaaS. The integration of the two platforms allows users to leverage the power of Windows Azure Platform for Aneka Cloud Computing, employing a large number of compute instances to run their applications in parallel. Furthermore, customers of the 647Windows Azure platform can benefit from the integration with Aneka …
An important level at which the expression of programmed cell death (PCD) genes is regulated is alternative splicing. Our previous work identified an intronic splicing regulatory element in caspase-2 (casp-2) gene. This 100-nucleotide intronic element, In100, consists of an upstream region containing a decoy 3' splice site and a downstream region containing binding sites for splicing repressor PTB. Based on the signal of In100 element in casp-2, we have detected the In100-like sequences as a family of sequence elements associated with alternative splicing in the human genome by using computational and experimental approaches. A survey of human genome reveals the presence of more than four thousand In100-like elements in 2757 genes. These In100-like elements tend to locate more frequent in intronic regions than exonic regions. EST analyses indicate that the presence of In100- like elements correlates with the skipping of their immediate upstream exons, with 526 genes showing exon skipping in such a manner. In addition, In100-like elements are found in several human caspase genes near exons encoding the caspase active domain. RT-PCR experiments show that these caspase genes indeed undergo alternative splicing in a pattern predicted to affect their functional activity. Together, these results suggest that the In100-like elements represent a family of intronic signals for alternative splicing in the human genome.
Motivation : mRNA sequences and expressed sequence tags represent some of the most abundant experimental data for identifying genes and alternatively spliced products in metazoans. These transcript sequences are frequently studied by aligning them to a genomic sequence template. For existing programs, error-prone, polymorphic and cross-species data, as well as non-canonical splice sites, still present significant barriers to producing accurate, complete alignments. Results: We took a novel approach to spliced alignment that meaningfully combined information from sequence similarity with that obtained from PSSM splice site models. Scoring systems were chosen to maximize their power of discrimination, and dynamic programming (DP) was employed to guarantee optimal solutions would be found. The resultant program, EXALIN, performed better than other popular tools tested under a wide range of conditions that included detection of micro-exons and human--mouse cross-species comparisons. For improved speed with only a marginal decrease in splice site prediction accuracy, EXALIN could perform limited DP guided by a result from BLASTN. Availability: The source code, binaries, scripts, scoring matrices and splice site models for human, mouse, rice and Caenorhabditis elegans utilized in this study are posted at http://blast.wustl.edu/exalin . The software (scripts, source code and binaries) is copyrighted but free for all to use. Contact: [email protected] Supplementary information: http://blast.wustl.edu/exalin/exalin-supplement.pdf
Transcription factors (TFs) are essential regulators of gene expression, and mutated TF genes have been shown to cause numerous human genetic diseases. Yet to date, no single, comprehensive database of human TFs exists. In this work, we describe the collection of an essentially complete set of TF genes from one depiction of the human ORFeome, and the design of a microarray to interrogate their expression. Taking 1468 known TFs from TRANSFAC, InterPro, and FlyBase, we used this seed set to search the ScriptSure human transcriptome database for additional genes. ScriptSure's genome-anchored transcript clusters allowed us to work with a nonredundant high-quality representation of the human transcriptome. We used a high-stringency similarity search by using BLASTN, and a protein motif search of the human ORFeome by using hidden Markov models of DNA-binding domains known to occur exclusively or primarily in TFs. Four hundred ninety-four additional TF genes were identified in the overlap between the two searches, bringing our estimate of the total number of human TFs to 1962. Zinc finger genes are by far the most abundant family (762 members), followed by homeobox (199 members) and basic helix-loop-helix genes (117 members). We designed a microarray of 50-mer oligonucleotide probes targeted to a unique region of the coding sequence of each gene. We have successfully used this microarray to interrogate TF gene expression in species as diverse as chickens and mice, as well as in humans.
Human chromosome 7 has historically received prominent attention in the human genetics community, primarily related to the search for the cystic fibrosis gene and the frequent cytogenetic changes associated with various forms of cancer. Here we present more than 153 million base pairs representing 99.4% of the euchromatic sequence of chromosome 7, the first metacentric chromosome completed so far. The sequence has excellent concordance with previously established physical and genetic maps, and it exhibits an unusual amount of segmentally duplicated sequence (8.2%), with marked differences between the two arms. Our initial analyses have identified 1,150 protein-coding genes, 605 of which have been confirmed by complementary DNA sequences, and an additional 941 pseudogenes. Of genes confirmed by transcript sequences, some are polymorphic for mutations that disrupt the reading frame.
Since 1995, the WU-BLAST programs (http://blast.wustl.edu) have provided a fast, flexible and reliable method for similarity searching of biological sequence databases. The software is in use at many locales and web sites. The European Bioinformatics Institute's WU-Blast2 (http://www.ebi.ac.uk/blast2/) server has been providing free access to these search services since 1997 and today supports many features that both enhance the usability and expand on the scope of the software.
The expressed sequence tag (EST) collection in dbEST provides an extensive resource for detecting alternative splicing on a genomic scale. Using genomically aligned ESTs, a computational tool (TAP) was used to identify alternative splice patterns for 6400 known human genes from the RefSeq database. With sufficient EST coverage, one or more alternatively spliced forms could be detected for nearly all genes examined. To identify high (>95%) confidence observations of alternative splicing, splice variants were clustered on the basis of having mutually exclusive structures, and sample statistics were then applied. Through this selection, alternative splices expected at a frequency of >5% within their respective clusters were seen for only 17%-28% of genes. Although intron retention events (potentially unspliced messages) had been seen for 36% of the genes overall, the same statistical selection yielded reliable cases of intron retention for <5% of genes. For high-confidence alternative splices in the human ESTs, we also noted significantly higher rates both of cross-species conservation in mouse ESTs and of validation in the GenBank mRNA collection. We suggest quantitative analytical approaches such as these can aid in selecting useful targets for further experimental characterization and in so doing may help elucidate the mechanisms and biological implications of alternative splicing.
In 1990, the United States Human Genome Project was initiated as a fifteen-year endeavor to sequence the approximately three billion bases making up the human genome (Vaughan, 1996). As of December 31, 2001, the public sequencing efforts have sequenced a total of 2.01 billion finished bases representing 63.0% of the human genome (http://www.ncbi.nlm.nih.gov/genome/seq/page.cgi?F=HsProgress.shtml&&ORG=Hs) to a Bermuda quality error rate of 1/10000 (Smith and Carrano, 1996). In addition, 1.11 billion bases representing 34.8% of the human genome has been sequenced to a rough-draft level. Efforts such as UCSC's GoldenPath (Kent and Haussler, 2001) and NCBI's contig assembly (Jang et al., 1999) attempt to assemble the human genome by incorporating both finished and rough-draft sequence. The availability of the human genome data allows us to ask questions concerning the maintenance of specific regions of the human genome. We consider two hypotheses for maintenance of high G+C regions: the presence of specific repetitive elements and compositional mutation biases. Our results rule out the possibility of the G+C content of repetitive elements determining regions of high and low G+C regions in the human genome. We determine that there is a compositional bias for mutation rates. However, these biases are not responsible for the maintenance of high G+C regions. In addition, we show that regions of the human under less selective pressure will mutate towards a higher A+T composition, regardless of the surrounding G+C composition. We also analyze sequence organization and show that previous studies of isochore regions (Bernardi, 1993) cannot be generalized within the human genome. In addition, we propose a method to assemble only those parts of the human genome that are finished into larger contigs. Analysis of the contigs can lead to the mining of meaningful biological data that can give insights into genetic variation and evolution. I suggest a method to help aid in single nucleotide polymorphism (SNP) detection, which can help to determine differences within a population. I also discuss a dynamic-programming based approach to sequence assembly validation and detection of large-scale polymorphisms within a population that is made possible through the availability of large human sequence contigs.
A fundamental problem in the human genome project is uncovering the correct assembly of the human genome. Many studies, including transcriptional analysis, SNP detection and characterization, gene finding and EST clustering, use genome assemblies as templates so it is important to determine the consistency among the various whole genome assemblies. A comparison of the order and orientation of the GenBank entries used to construct the NCBI and UCSC Goldenpath assemblies was made. In addition, a sequence level comparison was performed using MULTI, an efficient database search tool developed to make whole genome comparisons possible. The resulting comparisons show significant discrepancies in the sequence as well as in the order and orientation of GenBank entries used in constructing the NCBI and UCSC assemblies.
Single nucleotide polymorphisms (SNPs) are valuable genetic markers of human disease 1 , 2 , 3 . They also comprise the highest potential density marker set available for mapping experimentally derived mutations in model organisms such as Caenorhabditis elegans . To facilitate the positional cloning of mutations we have identified polymorphisms in CB4856, an isolate from a Hawaiian island that shows a uniformly high density of polymorphisms compared with the reference Bristol N2 strain. Based on 5.4 Mbp of aligned sequences, we predicted 6,222 polymorphisms. Furthermore, 3,457 of these markers modify restriction enzyme recognition sites ('snip-SNPs') and are therefore easily detected as RFLPs. Of these, 493 were experimentally confirmed by restriction digest to produce a snip-SNP map of the worm genome. A mapping strategy using snip-SNPs and bulked segregant analysis 4 (BSA) is outlined. CB4856 is crossed into a mutant strain, and exclusion of CB4856 alleles of a subset of snip-SNPs in mutant progeny is assesed with BSA. The proximity of a linked marker to the mutation is estimated by the relative proportion of each form of the biallelic marker in populations of wildtype and mutant genomes. The usefulness of this approach is illustrated by the rapid mapping of the dyf-5 gene.
transfection, cells were collected and processed for CAT or luciferase activity using standard techniques 14 . GST pull downs and immunoprecipitationsGST±Rb (wild type and mutants 15 ) and other GST fusion proteins were expressed and puri®ed from Escherichia coli XA90 (ref.16).GST fusion proteins that were immobilized on glutathione-sepharose (Pharmacia), or H3-derived peptides 3 bound to Sulfolink beads (Pierce), were incubated with extract in IPH buffer 16 .Complexes were washed four times in IPH buffer before processing for methylase assays or western blotting.Antibodies against HA (12CA5, Roche), Gal4±DBD (DNA-binding domain; sc-510, Santa Cruz), SUV39H1 (M.Cleary), Rb (G3-245; XZ55, PharMingen) or HP1 (ref.3) were used.For immunoprecipitations antibodies were incubated with HeLa nuclear extract (Cell Culture Center) or U2OS nuclear extract in IPH buffer at 4 8C (ref.17).After 2 h a 50:50 mixture of protein A/G-sepharose beads (Pharmacia) was added.To avoid the possibility that DNA mediates the interaction between SUV39H1 and Rb, the immunoprecipitations were probed for the presence of histones with negative results. Histone methylase assays and protein sequencingPrecipitations from pull downs or immunoprecipitations were incubated with 20 mg histones (Sigma) and 1 ml [3H-Me]-S-adenosyl methionine (NEN, 80 Ci mmol -1 ) in PBS at 30 8C for 1 h.Assays were analysed by SDS±PAGE followed by western blotting and autoradiography or were spotted onto P-81 cationic exchange paper (Whatman), washed in carbonate buffer and quanti®ed by scintillation counting 3 .For amino-terminal sequencing, radiolabelled H3 was blotted to polyvinylidene ¯uoride and sequenced by Edman degradation (Protein Sequencing Facility, University of Cambridge).We counted fractions for the presence of tritium. RNA puri®cation and RT-PCR analysisTotal RNA (0.5 mg) was isolated from W12 (wild type) and D3 (SUV39H1 and SUV39H2 double knockout; D.O. and T.J., unpublished observations) female mouse cells, and was used for quantitative RT-PCR, following the Qiagen One Step protocol, for 20, 25 and 30 PCR cycles. Antibody generationRabbits were immunized with a H3 N-terminal lysine-methylated peptide corresponding to amino acids 1±16.Immunoreactive serum was applied to a H3 Lys-9-methylated peptide column to af®nity purify speci®c antibodies, as the antiserum crossreacted with H3 methylated at Lys 4. Chromatin immunoprecipitationChromatin immunoprecipitations were performed using HeLa cells and MEF cells essentially as described 18,19 .Immunoprecipitates were analysed for the presence of cyclin E or Cdc25C promoter fragments by PCR using primers speci®c for single nucleosomes.PCR reactions were repeated exhaustively using varying cycle numbers and different amounts of templates to ensure that results were within the linear range of the PCR.
With the availability of a nearly complete sequence of the human genome, aligning expressed sequence tags (EST) to the genomic sequence has become a practical and powerful strategy for gene prediction. Elucidating gene structure is a complex problem requiring the identification of splice junctions, gene boundaries, and alternative splicing variants. We have developed a software tool, Transcript Assembly Program (TAP), to delineate gene structures using genomically aligned EST sequences. TAP assembles the joint gene structure of the entire genomic region from individual splice junction pairs, using a novel algorithm that uses the EST-encoded connectivity and redundancy information to sort out the complex alternative splicing patterns. A method called polyadenylation site scan (PASS) has been developed to detect poly-A sites in the genome. TAP uses these predictions to identify gene boundaries by segmenting the joint gene structure at polyadenylated terminal exons. Reconstructing 1007 known transcripts, TAP scored a sensitivity (Sn) of 60% and a specificity (Sp) of 92% at the exon level. The gene boundary identification process was found to be accurate 78% of the time. also reports alternative splicing patterns in EST alignments. An analysis of alternative splicing in 1124 genic regions suggested that more than half of human genes undergo alternative splicing. Surprisingly, we saw an absolute majority of the detected alternative splicing events affect the coding region. Furthermore, the evolutionary conservation of alternative splicing between human and mouse was analyzed using an EST-based approach. (See http://stl.wustl.edu/~zkan/TAP/)
UNLABELLED:Identifying and masking repetitive elements is usually the first step when analyzing vertebrate genomic sequence. Current repeat identification software is sensitive but slow, creating a costly bottleneck in large-scale analyses. We have developed MaskerAid, a software enhancement to RepeatMasker that increased the speed of masking more than 30-fold at the most sensitive setting.AVAILABILITY:On request from the authors (see http://sapiens.wustl.edu/MaskerAid).CONTACT:maskeraid@watson.wustl.edu
UNLABELLEDWe have developed a program, MPBLAST, that increases the throughput of batch BLASTN searches by multiplexing (concatenating) query sequences and thereby reducing the number of actual database searches performed. Throughput was observed to increase in reciprocal proportion to the component sequence length. For sequencing read-sized queries of 500 bp, an order of magnitude speed-up was seen.AVAILABILITYFree (see http://blast.wustl.edu)CONTACT[ikorf, gish]@watson.wustl.edu
Untranslated regions (UTR) play important roles in the posttranscriptional regulation of mRNA processing. There is a wealth of UTR-related information to be mined from the rapidly accumulating EST collections. A computational tool, UTR-extender, has been developed to infer UTR sequences from genomically aligned ESTs. It can completely and accurately reconstruct 72% of the 3' UTRs and 15% of the 5' UTRs when tested using 908 functionally cloned transcripts. In addition, it predicts extensions for 11% of the 5' UTRs and 28% of the 3' UTRs. These extension regions are validated by examining splicing frequencies and conservation levels. We also developed a method called polyadenylation site scan (PASS) to precisely map polyadenylation sites in human genomic sequences. A PASS analysis of 908 genic regions estimates that 40-50% of human genes undergo alternative polyadenylation. Using EST redundancy to assess expression levels, we also find that genes with short 3' UTRs tend to be highly expressed.
Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits1. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence6,7 as a template on which to layer often unmapped, fragmentary sequence data8,9,10,11 and to use base quality values12 to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.
Ian Korf合作论文数University of California, Davis3