Phylogenetic reconstruction of herpesvirus evolution is generally founded on amino acid sequence comparisons of specific proteins. These are relevant to the evolution of the specific gene (or set of genes), but the resulting phylogeny may vary depending on the particular sequence chosen for analysis (or comparison). In the first part of this report, we compare 13 herpesvirus genomes by using a new multidimensional methodology based on distance measures and partial orderings of dinucleotide relative abundances. The sequences were analyzed with respect to (i) genomic compositional extremes; (ii) total distances within and between genomes; (iii) partial orderings among genomes relative to a set of sequence standards; (iv) concordance correlations of genome distances; and (v) consistency with the alpha-, beta-, gammaherpesvirus classification. Distance assessments within individual herpesvirus genomes show each to be quite homogeneous relative to the comparisons between genomes. The gammaherpesviruses, Epstein-Barr virus (EBV), herpesvirus saimiri, and bovine herpesvirus 4 are both diverse and separate from other herpesvirus classes, whereas alpha- and betaherpesviruses overlap. The analysis revealed that the most central genome (closest to a consensus herpesvirus genome and most individual herpesvirus sequences of different classes) is that of human herpesvirus 6, suggesting that this genome is closest to a progenitor herpesvirus. The shorter DNA distances among alphaherpesviruses supports the hypothesis that the alpha class is of relatively recent ancestry. In our collection, equine herpesvirus 1 (EHV1) stands out as the most central alphaherpesvirus, suggesting it may approximate an ancestral alphaherpesvirus. Among all herpesviruses, the EBV genome is closest to human sequences. In the DNA partial orderings, the chicken sequence collection is invariably as close as or closer to all herpesvirus sequences than the human sequence collection is, which may imply that the chicken (or other avian species) is a more natural or more ancient host of herpesviruses. In the second part of this report, evolutionary relationships among the 13 herpesvirus genomes are evaluated on the basis of recent methods of amino acid alignment applied to four essential protein sequences. In this analysis, the alignment of the two betaherpesviruses (human cytomegalovirus versus human herpesvirus 6) showed lower scores compared with alignments within alphaherpesviruses (i.e., among EHV1, herpes simplex virus type 1, varicella-zoster virus, pseudorabies virus type 1 and Marek's disease virus) and within gammaherpesviruses (EBV versus herpesvirus saimiri).(ABSTRACT TRUNCATED AT 400 WORDS)
The recent sequencing of two relatively long (approximately 100 kb) contigs of E.coli presents unique opportunities for investigating heterogeneity and genomic organization of the E.coli chromosome. We have evaluated a number of common and contrasting sequence features in the two new contigs with comparisons to all available E.coli sequences (> 1.6 Mb). Our analyses include assessments of: (i) counts and distributions of restriction sites, special oligonucleotides (e.g., Chi sites, Dam and Dcm methylase targets), and other marker arrays; (ii) significant distant and close direct and inverted repeat sequences; (iii) sequence similarities between the long contigs and other E.coli sequences; (iv) characterization and identification of rare and frequent oligonucleotides; (v) compositional biases in short oligonucleotides; and (vi) position-dependent fluctuations in sequence composition. The two contigs reveal a number of distinctive features, including: a cluster of five repeat/dyad elements with very regular spacings resembling a transcription attenuator in one of the contigs; REP elements, ERICs, and other long repeats; distinction of the Chi sequence as the most frequent oligonucleotide; regions of clustering, overdispersion, and regularity of certain restriction sites and short palindromes; and comparative domains of inhomogeneities in the two long contigs. These and other features are discussed in relation to the organization of the E.coli chromosome.
A global analysis of the 230-kilobase-pair (kbp) human cytomegalovirus genome revealed three regions that were very rich in repeated sequences. The region with the highest content of inverted and direct repeats lies between 92,100 and 93,500 bp, upstream of the gene encoding the single-stranded DNA binding protein. Cloned restriction fragments containing this region were able to replicate when trans-acting factors were provided by virus infection in a transient replication assay. With this assay, the region between 92,210 and 93,715 bp on the viral genome was defined as the minimal replication origin, oriLyt. The sequence composition and repeats within oriLyt were used to divide the region into two domains that may be important in origin function. Sequences flanking either the left or right side of the minimal oriLyt contributed to efficient replication; however, these sequences were not essential for origin function. Thus, the region of the viral genome with the most striking concentration of direct and inverted repeats corresponds to the oriLyt of human cytomegalovirus.
Epstein-Barr virus (EBV) has two different modes of existence: latent and productive. There are eight known genes expressed during latency (and hardly at all during the productive phase) and about 70 other ("productive") genes. It is shown that the EBV genes known to be expressed during latency display codon usage strikingly different from that of genes that are expressed during lytic growth. In particular, the percentage of S3 (G or C in codon site 3) is persistently lower (about 20%) in all latent genes than in nonlatent genes. Moreover, S3 is lower in each multicodon amino acid form. Also, the percentage of S in silent codon sites 1 of leucine and arginine is lower in latent than in nonlatent genes. The largest absolute differences in amino acid usage between latent and nonlatent genes emphasize codon types SSN and WWN (W means nucleotide A or T and N is any nucleotide). Two principal explanations to account for the EBV latent versus productive gene codon disparity are proposed. Latent genes have codon usage substantially different from that of host cell genes to minimize the deleterious consequences to the host of viral gene expression during latency. (Productive genes are not so constrained.) It is also proposed that the latency genes of EBV were acquired recently by the viral genome. Evidence and arguments for these proposals are presented.