Although new and emerging next-generation sequencing (NGS) technologies have reduced sequencing costs significantly, much work remains to implement them for de novo sequencing of complex and highly repetitive genomes such as the tetraploid genome of Upland cotton (Gossypium hirsutum L.). Herein we report the results from implementing a novel, hybrid Sanger/454-based BAC-pool sequencing strategy using minimum tiling path (MTP) BACs from Ctg-3301 and Ctg-465, two large genomic segments in A12 and D12 homoeologous chromosomes (Ctg). To enable generation of longer contig sequences in assembly, we implemented a hybrid assembly method to process ~35x data from 454 technology and 2.8-3x data from Sanger method. Hybrid assemblies offered higher sequence coverage and better sequence assemblies. Homology studies revealed the presence of retrotransposon regions like Copia and Gypsy elements in these contigs and also helped in identifying new genomic SSRs. Unigenes were anchored to the sequences in Ctg-3301 and Ctg-465 to support the physical map. Gene density, gene structure and protein sequence information derived from protein prediction programs were used to obtain the functional annotation of these genes. Comparative analysis of both contigs with Arabidopsis genome exhibited synteny and microcollinearity with a conserved gene order in both genomes. This study provides insight about use of MTP-based BAC-pool sequencing approach for sequencing complex polyploid genomes with limited constraints in generating better sequence assemblies to build reference scaffold sequences. Combining the utilities of MTP-based BAC-pool sequencing with current longer and short read NGS technologies in multiplexed format would provide a new direction to cost-effectively and precisely sequence complex plant genomes.
Legumes (Fabaceae or Leguminosae) are unique among cultivated plants for their ability to carry out endosymbiotic nitrogen fixation with rhizobial bacteria, a process that takes place in a specialized structure known as the nodule. Legumes belong to one of the two main groups of eurosids, the Fabidae, which includes most species capable of endosymbiotic nitrogen fixation. Legumes comprise several evolutionary lineages derived from a common ancestor 60 million years ago (Myr ago). Papilionoids are the largest clade, dating nearly to the origin of legumes and containing most cultivated species. Medicago truncatula is a long-established model for the study of legume biology. Here we describe the draft sequence of the M. truncatula euchromatin based on a recently completed BAC assembly supplemented with Illumina shotgun sequence, together capturing ∼94% of all M. truncatula genes. A whole-genome duplication (WGD) approximately 58 Myr ago had a major role in shaping the M. truncatula genome and thereby contributed to the evolution of endosymbiotic nitrogen fixation. Subsequent to the WGD, the M. truncatula genome experienced higher levels of rearrangement than two other sequenced legumes, Glycine max and Lotus japonicus. M. truncatula is a close relative of alfalfa (Medicago sativa), a widely cultivated crop with limited genomics tools and complex autotetraploid genetics. As such, the M. truncatula genome sequence provides significant opportunities to expand alfalfa's genomic toolbox.
Background: Follicular lymphoma (FL) is a form of non-Hodgkin's lymphoma (NHL) that arises from germinal center (GC) B-cells. Despite the significant advances in immunotherapy, FL is still not curable. Beyond transcriptional profiling and genomics datasets, there currently is no epigenome-scale dataset or integrative biology approach that can adequately model this disease and therefore identify novel mechanisms and targets for successful prevention and treatment of FL.Methodology/Principal Findings: We performed methylation-enriched genome-wide bisulfite sequencing of FL cells and normal CD19(+) B-cells using 454 sequencing technology. The methylated DNA fragments were enriched with methyl-binding proteins, treated with bisulfite, and sequenced using the Roche-454 GS FLX sequencer. The total number of bases covered in the human genome was 18.2 and 49.3 million including 726,003 and 1.3 million CpGs in FL and CD19(+) B-cells, respectively. 11,971 and 7,882 methylated regions of interest (MRIs) were identified respectively. The genome-wide distribution of these MRIs displayed significant differences between FL and normal B-cells. A reverse trend in the distribution of MRIs between the promoter and the gene body was observed in FL and CD19+ B-cells. The MRIs identified in FL cells also correlated well with transcriptomic data and ChIP-on-Chip analyses of genome-wide histone modifications such as tri-methyl-H3K27, and tri-methyl-H3K4, indicating a concerted epigenetic alteration in FL cells.Conclusions/Significance: This study is the first to provide a large scale and comprehensive analysis of the DNA methylation sequence composition and distribution in the FL epigenome. These integrated approaches have led to the discovery of novel and frequent targets of aberrant epigenetic alterations. The genome-wide bisulfite sequencing approach developed here can be a useful tool for profiling DNA methylation in clinical samples.
Background Sugarcane ( Saccharum spp.) has become an increasingly important crop for its leading role in biofuel production. The high sugar content species S. officinarum is an octoploid without known diploid or tetraploid progenitors. Commercial sugarcane cultivars are hybrids between S. officinarum and wild species S. spontaneum with ploidy at ~12×. The complex autopolyploid sugarcane genome has not been characterized at the DNA sequence level. Results The microsynteny between sugarcane and sorghum was assessed by comparing 454 pyrosequences of 20 sugarcane bacterial artificial chromosomes (BACs) with sorghum sequences. These 20 BACs were selected by hybridization of 1961 single copy sorghum overgo probes to the sugarcane BAC library with one sugarcane BAC corresponding to each of the 20 sorghum chromosome arms. The genic regions of the sugarcane BACs shared an average of 95.2% sequence identity with sorghum, and the sorghum genome was used as a template to order sequence contigs covering 78.2% of the 20 BAC sequences. About 53.1% of the sugarcane BAC sequences are aligned with sorghum sequence. The unaligned regions contain non-coding and repetitive sequences. Within the aligned sequences, 209 genes were annotated in sugarcane and 202 in sorghum. Seventeen genes appeared to be sugarcane-specific and all validated by sugarcane ESTs, while 12 appeared sorghum-specific but only one validated by sorghum ESTs. Twelve of the 17 sugarcane-specific genes have no match in the non-redundant protein database in GenBank, perhaps encoding proteins for sugarcane-specific processes. The sorghum orthologous regions appeared to have expanded relative to sugarcane, mostly by the increase of retrotransposons. Conclusions The sugarcane and sorghum genomes are mostly collinear in the genic regions, and the sorghum genome can be used as a template for assembling much of the genic DNA of the autopolyploid sugarcane genome. The comparable gene density between sugarcane BACs and corresponding sorghum sequences defied the notion that polyploidy species might have faster pace of gene loss due to the redundancy of multiple alleles at each locus.
Pre-mRNA 5' spliced-leader (SL) trans-splicing occurs in some metazoan groups but not in others. Genome-wide characterization of the trans-spliced mRNA subpopulation has not yet been reported for any metazoan. We carried out a high-throughput analysis of the SL trans-spliced mRNA population of the ascidian tunicate Ciona intestinalis by 454 Life Sciences (Roche) pyrosequencing of SL-PCR-amplified random-primed reverse transcripts of tailbud embryo RNA. We obtained approximately 250,000 high-quality reads corresponding to 8790 genes, approximately 58% of the Ciona total gene number. The great depth of this data revealed new aspects of trans-splicing, including the existence of a significant class of "infrequently trans-spliced" genes, accounting for approximately 28% of represented genes, that generate largely non-trans-spliced mRNAs, but also produce trans-spliced mRNAs, in part through alternative promoter use. Thus, the conventional qualitative dichotomy of trans-spliced versus non-trans-spliced genes should be supplanted by a more accurate quantitative view recognizing frequently and infrequently trans-spliced gene categories. Our data include reads representing approximately 80% of Ciona frequently trans-spliced genes. Our analysis also revealed significant use of closely spaced alternative trans-splice acceptor sites which further underscores the mechanistic similarity of cis- and trans-splicing and indicates that the prevalence of +/-3-nt alternative splicing events at tandem acceptor sites, NAGNAG, is driven by spliceosomal mechanisms, and not nonsense-mediated decay, or selection at the protein level. The breadth of gene representation data enabled us to find new correlations between trans-splicing status and gene function, namely the overrepresentation in the frequently trans-spliced gene class of genes associated with plasma/endomembrane system, Ca(2+) homeostasis, and actin cytoskeleton.
Background MicroRNAs (miRNAs) are small ~22-nt regulatory RNAs that can silence target genes, by blocking their protein production or degrading the mRNAs. Pig is an important animal in the agriculture industry because of its utility in the meat production. Besides, pig has tremendous biomedical importance as a model organism because of its closer proximity to humans than the mouse model. Several hundreds of miRNAs have been identified from mammals, humans, mice and rats, but little is known about the miRNA component in the pig genome. Here, we adopted an experimental approach to identify conserved and unique miRNAs and characterize their expression patterns in diverse tissues of pig. Results By sequencing a small RNA library generated using pooled RNA from the pig heart, liver and thymus; we identified a total of 120 conserved miRNA homologs in pig. Expression analysis of conserved miRNAs in 14 different tissue types revealed heart-specific expression of miR-499 and miR-208 and liver-specific expression of miR-122. Additionally, miR-1 and miR-133 in the heart, miR-181a and miR-142-3p in the thymus, miR-194 in the liver, and miR-143 in the stomach showed the highest levels of expression. miR-22, miR-26b, miR-29c and miR-30c showed ubiquitous expression in diverse tissues. The expression patterns of pig-specific miRNAs also varied among the tissues examined. Conclusion Identification of 120 miRNAs and determination of the spatial expression patterns of a sub-set of these in the pig is a valuable resource for molecular biologists, breeders, and biomedical investigators interested in post-transcriptional gene regulation in pig and in related mammals, including humans.
MicroRNAs (miRNAs) and small-interfering RNAs (siRNAs) have emerged as important regulators of gene expression in higher eukaryotes. Recent studies indicate that genomes in higher plants encode lineage-specific and species-specific miRNAs in addition to the well-conserved miRNAs. Leguminous plants are grown throughout the world for food and forage production. To date the lack of genomic sequence data has prevented systematic examination of small RNAs in leguminous plants. Medicago truncatula, a diploid plant with a near-completely sequenced genome has recently emerged as an important model legume.We sequenced a small RNA library generated from M. truncatula to identify not only conserved miRNAs but also novel small RNAs, if any.Eight novel small RNAs were identified, of which four (miR1507, miR2118, miR2119 and miR2199) are annotated as legume-specific miRNAs because these are conserved in related legumes. Three novel transcripts encoding TIR-NBS-LRR proteins are validated as targets for one of the novel miRNA, miR2118. Small RNA sequence analysis coupled with the small RNA blot analysis, confirmed the expression of around 20 conserved miRNA families in M. truncatula. Fifteen transcripts have been validated as targets for conserved miRNAs. We also characterized Tas3-siRNA biogenesis in M. truncatula and validated three auxin response factor (ARF) transcripts that are targeted by tasiRNAs.These findings indicate that M. truncatula and possibly other related legumes have complex mechanisms of gene regulation involving specific and common small RNAs operating post-transcriptionally.
ABSTRACT We used an expressed sequence tag and 454 pyrosequencing approach to initiate a study of the genome of the screwworm, Cochliomyia hominivorax (Coquerel) (Diptera: Calliphoridae). Two normalized cDNA libraries were constructed from RNA isolated from embryos and second instar larvae from the Panama 95 strain. Approximately 5,400 clones from each library were sequenced from both the 5′ and 3′ directions using the Sanger method. In addition, double-stranded cDNA was prepared from random-primed polyA RNA purified from embryos, second-instar larvae, adult males, and adult females. These four cDNA samples were used for 454 pyrosequencing that produced ≈300,000 independent sequences. Sequences were assembled into a database of assembled contigs and singletons and used to search public protein databases and annotate the sequences. The full database consists of 6,076 contigs and 58,221 singletons assembled from both the traditional expressed sequence tag (EST) and 454 sequences. Annotation of the data led to the identification of several gene coding regions with possible roles in sex determination in the screwworm. This database will facilitate the design of microarray and other experiments to study screwworm gene expression on a larger scale than previously possible.
With the introduction of massively parallel, microminiature-based instrumentation for DNA sequencing, robust, reproducible, optimized methods are needed to prepare the target DNA for analysis using these high-throughput approaches because the cost per instrument run is orders of magnitude more than for typical Sanger dideoxynucleotide sequencing on fluorescence-based capillary systems. The methods provided by the manufacturer for genome sequencing using the 454/Roche GS-20 and GS-FLX instruments are robust. However, in an effort to streamline them for automation, we have incorporated several novel changes and deleted several extraneous steps. As a result of modifying these sample preparation protocols, the number of manual manipulations has also been minimized, and the overall yields have been improved for both shotgun and mixed shotgun/paired-end libraries.
The horn fly, Haematobia irritans L., is an obligate blood-feeding parasite of cattle, and control of this pest is a continuing problem because the fly is becoming resistant to pesticides. Dominant conditional lethal gene systems are being studied as population control technologies against agricultural pests. One of the components of these systems is a female-specific gene promoter that drives expression of a lethality-inducing gene. To identify candidate genes to supply this promoter, microarrays were designed from a horn fly expressed sequence tag (EST) database and probed to identify female-specific and larval-specific gene expression. Analysis of dye swap experiments found 432 and 417 transcripts whose expression levels were higher or lower in adult female flies, respectively, compared with adult male flies. Additionally, 419 and 871 transcripts were identified whose expression levels were higher or lower in first-instar larvae compared with adult flies, respectively. Three transcripts were expressed more highly in adult females flies compared with adult males and also higher in the first-instar larval lifestage compared with adult flies. One of these transcripts, a putative nanos ortholog, has a high female-to-male expression ratio, a moderate expression level in first-instar larvae, and has been well characterized in Drosophila. melanogaster (Meigen). In conclusion, we used microarray technology, verified by reverse transcriptase-polymerase chain reaction and massively parallel pyrosequencing, to study life stage- and sex-specific gene expression in the horn fly and identified three gene candidates for detailed evaluation as a gene promoter source for the development of a female-specific conditional lethality system.
Enzymatic transformation of humic acids (HA), fulvic acids (FA) and indole was examined using naphthalene 1,2-dioxygenase (NDO). NDO was used as a model for dioxygenase enzymes found in various microbial species. Indole was used as a model substrate for NDO-catalyzed reactions resulting in condensation products. Although NDO is not classified as a soil enzyme, all HA and FA tested were susceptible to NDO-induced transformation. The extent of NDO-specific NADH oxidation in solutions containing HA and FA paralleled the percent aromaticity of the HA and FA. Furthermore, the UV–Vis absorptive properties of NDO-treated HA and FA were altered in a manner suggesting condensation reactions similar to the formation of indigo from indole. Condensation reactions were enhanced in NDO-treated mixtures containing indole and an FA. NDO retained activity for 2 weeks under ambient conditions, and retained some enzymatic activity for 9 days based on detection of specific metabolites by HPLC, suggesting prolonged extracellular activity. Humic substances have not previously been known to be substrates for dioxygenases; even more significant was that dioxygenase enzymes can facilitate condensation reactions between indole-like functional groups well-known to be present in HA and FA. These results illustrate how dioxygenases can be potential humic-modifying enzymes when released into the environment upon microbial death and concurrent cell lysis which could alter the bioavailability of organic contaminants associated with dissolved organic matter through specific modulation of enzyme activity involving substrate competition.
Background: The draft genome sequence of the ascidian Ciona intestinalis, along with associated gene models, has been a valuable research resource. However, recently accumulated expressed sequence tag (EST)/cDNA data have revealed numerous inconsistencies with the gene models due in part to intrinsic limitations in gene prediction programs and in part to the fragmented nature of the assembly.Results: We have prepared a less-fragmented assembly on the basis of scaffold-joining guided by paired-end EST and bacterial artificial chromosome (BAC) sequences, and BAC chromosomal in situ hybridization data. The new assembly (115.2 Mb) is similar in length to the initial assembly (116.7 Mb) but contains 1,272 (approximately 50%) fewer scaffolds. The largest scaffold in the new assembly incorporates 95 initial-assembly scaffolds. In conjunction with the new assembly, we have prepared a greatly improved global gene model set strictly correlated with the extensive currently available EST data. The total gene number (15,254) is similar to that of the initial set (15,582), but the new set includes 3,330 models at genomic sites where none were present in the initial set, and 1,779 models that represent fusions of multiple previously incomplete models. In approximately half, 5'-ends were precisely mapped using 5'-full-length ESTs, an important refinement even in otherwise unchanged models.Conclusion: Using these new resources, we identify a population of non-canonical (non-GT-AG) introns and also find that approximately 20% of Ciona genes reside in operons and that operons contain a high proportion of single-exon genes. Thus, the present dataset provides an opportunity to analyze the Ciona genome much more precisely than ever.
BACKGROUND:The Human Microbiome Project (HMP) is one of the U.S. National Institutes of Health Roadmap for Medical Research. Primary interests of the HMP include the distinctiveness of different gut microbiomes, the factors influencing microbiome diversity, and the functional redundancies of the members of human microbiotas. In this present work, we contribute to these interests by characterizing two extinct human microbiotas.METHODOLOGY/PRINCIPAL FINDINGS:We examine two paleofecal samples originating from cave deposits in Durango Mexico and dating to approximately 1300 years ago. Contamination control is a serious issue in ancient DNA research; we use a novel approach to control contamination. After we determined that each sample originated from a different human, we generated 45 thousand shotgun DNA sequencing reads. The phylotyping and functional analysis of these reads reveals a signature consistent with the modern gut ecology. Interestingly, inter-individual variability for phenotypes but not functional pathways was observed. The two ancient samples have more similar functional profiles to each other than to a recently published profile for modern humans. This similarity could not be explained by a chance sampling of the databases.CONCLUSIONS/SIGNIFICANCE:We conduct a phylotyping and functional analysis of ancient human microbiomes, while providing novel methods to control for DNA contamination and novel hypotheses about past microbiome biogeography. We postulate that natural selection has more of an influence on microbiome functional profiles than it does on the species represented in the microbial ecology. We propose that human microbiomes were more geographically structured during pre-Columbian times than today.
The ascomycete Epichloë festucae is a model endophyte that 1) switches between mutualistic and antagonistic states, 2) is seed transmissible, 3) has a sexual state amenable to genetic analysis, and 4) is rich in bioprotective alkaloids. This fungus grows systemically and intercellularly throughout the life of its host plant. On each reproductive tiller the fungus either infects benignly and transmits clonally in seeds, or produces its sexual state (stroma) and chokes inflorescence development. The E. festucae genome was estimated at 29 Mb in six chromosomes. The genome sequence was assembled from cloned insert end reads (4.2 x coverage) and preassembled pyrosequencing reads (454-sequencing: 20 x raw, 1.7 x assembled), giving 3967 supercontigs, of which 1004 were larger than 2 kb and covered 92% of the genome. Gene prediction with FGENESH identified ~10,000 putative genes. We also sequenced 25,000 ESTs from each of two normalised libraries — one of choked inflorescences, the other of benignly infected inflorescences — yielding 5077 E. festucae unigenes, annotated by BLAST and InterPro. Sequence data and annotations are stored in a database for visualisation and inspection with the GBrowse browser. The genomic sequences can be queried by BLAST at http://www.genome.ou.edu/blast/ ef_blastall.html. Keywords: bioinformatics, DNA sequence, Epichloë festucae, expressed sequence tags, Festucae pratensis, fungal genomics, Lolium pratense