Genome replication by the baculovirus DNA polymerase often generates errors in mononucleotide repeat (MNR) sequences due to replication slippage. This results in the inactivation of genes that affects different stages of the cell infection cycle. Here we mapped these MNRs in the 59 baculovirus genomes. We found that the MNR frequencies of baculovirus genomes are different and not correlated with the genome sizes. Although the average A/T content of baculoviruses is 58.67%, the A/T MNR frequency is significantly higher than that of the G/C MNRs. Furthermore, the A7/T7 MNRs are the most frequent of those we studied. Finally, MNR frequencies in different classes of baculovirus genes, such as immediate early genes, show differences between baculovirus genomes, suggesting that the distribution and frequency of different MNRs are unique to each baculovirus species or strain. Therefore, the results of this study can help select appropriate baculoviruses for the development of biological insecticides.
SUMMARY To address the impending need for exploring rapidly increased transcriptomics data generated for non-model organisms, we developed CBrowse, an AJAX-based web browser for visualizing and analyzing transcriptome assemblies and contigs. Designed in a standard three-tier architecture with a data pre-processing pipeline, CBrowse is essentially a Rich Internet Application that offers many seamlessly integrated web interfaces and allows users to navigate, sort, filter, search and visualize data smoothly. The pre-processing pipeline takes the contig sequence file in FASTA format and its relevant SAM/BAM file as the input; detects putative polymorphisms, simple sequence repeats and sequencing errors in contigs and generates image, JSON and database-compatible CSV text files that are directly utilized by different web interfaces. CBowse is a generic visualization and analysis tool that facilitates close examination of assembly quality, genetic polymorphisms, sequence repeats and/or sequencing errors in transcriptome sequencing projects. AVAILABILITY CBrowse is distributed under the GNU General Public License, available at http://bioinfolab.muohio.edu/CBrowse/ CONTACT liangc@muohio.edu or liangc.mu@gmail.com; glji@xmu.edu.cn SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.
Background: The peanut (Arachis hypogaea) is an important crop cultivated worldwide for oil production and food sources. Its complex genetic architecture (e.g., the large and tetraploid genome possibly due to unique cross of wild diploid relatives and subsequent chromosome duplication: 2n = 4x = 40, AABB, 2800 Mb) presents a major challenge for its genome sequencing and makes it a less-studied crop. Without a doubt, transcriptome sequencing is the most effective way to harness the genome structure and gene expression dynamics of this non-model species that has a limited genomic resource.Description: With the development of next generation sequencing technologies such as 454 pyro-sequencing and Illumina sequencing by synthesis, the transcriptomics data of peanut is rapidly accumulated in both the public databases and private sectors. Integrating 187,636 Sanger reads (103,685,419 bases), 1,165,168 Roche 454 reads (333,862,593 bases) and 57,135,995 Illumina reads (4,073,740,115 bases), we generated the first release of our peanut transcriptome assembly that contains 32,619 contigs. We provided EC, KEGG and GO functional annotations to these contigs and detected SSRs, SNPs and other genetic polymorphisms for each contig. Based on both open-source and our in-house tools, PeanutDB presents many seamlessly integrated web interfaces that allow users to search, filter, navigate and visualize easily the whole transcript assembly, its annotations and detected polymorphisms and simple sequence repeats. For each contig, sequence alignment is presented in both bird's-eye view and nucleotide level resolution, with colorfully highlighted regions of mismatches, indels and repeats that facilitate close examination of assembly quality, genetic polymorphisms, sequence repeats and/or sequencing errors.Conclusion: As a public genomic database that integrates peanut transcriptome data from different sources, PeanutDB (http://bioinfolab.muohio.edu/txid3818v1) provides the Peanut research community with an easy-to-use web portal that will definitely facilitate genomics research and molecular breeding in this less-studied crop.