The severity of Helicobacter pylori-related disease is correlated with a pathogenicity island (the Cag region of about 26 genes) whose presence is associated with the up-regulation of an IL-8 cytokine inflammatory response in gastric epithelial cells. Statistical analysis of the Cag gene sequences calculated from the complete genome of strain 26695 revealed several unusual features. The Cag7 sequence (1,927 aa) has two repeat regions. Repeat region I runs 317 aa in a form of AAA proximal to the protein N terminal; repeat region II extends 907 aa in the middle of the protein sequence consisting of 74 contiguous segments composed from selections among six consensus sequences and includes 58 regularly distributed cysteine residues with consecutive cysteines mostly 12, 18, or 24 aa apart. This "regular" cysteine arrangement may provide a scaffolding of linker elements stabilized by disulfide bridges. When Cag7 homologues from different strains are compared, differences were found almost exclusively in the repeat regions, resulting from deletion and/or insertion of repeating units. These observations suggest that the anomalous repetitive structure of the sequence plays an important role in the conformation of Cag7 gene product and potentially in the function of the pathogenicity island. Other facets of the Cag7 sequence show significant charge clusters, high multiplet count, and extremes of amino acid usage.
ABSTRACT: We present new methods for calculating codon bias of a group of genes or an individual gene relative to a standard gene class. This method is suitable for identifying alien (e.g., horizontally transferred) and highly expressed genes. In yeast and several bacterial genomes, highly expressed genes typically include ribosomal protein genes, elongation factors, chaperonins (heat shock proteins), and a subset of genes involved in glycolysis generally essential in exponential growth. Highly expressed genes of the Synechocystis genome feature several photosystem II genes, and highly expressed genes in several methanogens (Methanococcus jannaschii, M. thermoautotrophicum) are essential for methanogenesis. Alien genes mostly consist of ORFs of unknown function, transposases, prophage genes, and restriction/modification enzymes. Notably, nuclear ribosomal proteins of yeast are highly expressed, whereas mitochondrial ribosomal protein genes appear to be alien genes. Alien genes often occur in clusters, suggesting in these cases that transfer events entail several genes.
Several bacterial genomes exhibit preference for G over C on the DNA leading strand extending from the origin of replication to the ter-region in the genomes of Escherichia coli, Mycoplasma genitalium, Bacillus subtilis, and marginally in Haemophilus influenzae, Mycoplasma pneumoniae, and Helicobacter pylori. Strand compositional asymmetry is not observed in the cyanobacterium Synechocystis sp. genome nor in the archaeal genomes of Methanococcus jannaschii, Methanobacterium thermoautotrophicum, and Archaeoglobus fulgidus. A strong strand compositional asymmetry is observed in beta-type but not alpha- or gamma-type human herpesviruses featuring G > C downstream of oriL and C > G upstream of oriL. Dinucleotide relative abundances (i.e., dinucleotide representations normalized by the component nucleotide frequencies) are consonant with respect to the leading and lagging strands. Strand compositional asymmetry may reflect on differences in replication synthesis of the leading versus lagging strand, on differences between template and coding strand associated with transcription-coupled repair mechanisms, on differences in gene density between the two strands, on differences in residue and codon biases in relation to gene function, expression level, or operon organization, or on differences in single or context-dependent base mutational rates. The absence of strand asymmetry in the archaeal genomes may reflect the presence of multiple origins of replication.
Early biochemical experiments measuring nearest neighbor frequencies established that the set of dinucleotide relative abundance values (dinucleotide biases) is a remarkably stable property of the DNA of an organism. Analyses of currently available genomic sequence data have extended these earlier results, showing that the dinucleotide biases evaluated for successive 50 kb segments of a genome are significantly more similar to each other than to those of sequences from more distant organisms. From this perspective, the set of dinucleotide biases constitutes a ‘genomic signature’ that can discriminate sequences from different organisms. The dinucleotide biases appear to reflect species-specific properties of DNA stacking energies, modification, replication, and repair mechanisms. The genomic signature is useful for detecting pathogenicity islands in bacterial genomes.
Several human neurological disorders are associated with proteins containing abnormally long runs of glutamine residues. Strikingly, most of these proteins contain two or more additional long runs of amino acids other than glutamine. We screened the current human, mouse, Drosophila, yeast, and Escherichia coli protein sequence data bases and identified all proteins containing multiple long homopeptides. This search found multiple long homopeptides in about 12% of Drosophila proteins but in only about 1.7% of human, mouse, and yeast proteins and none among E. coli proteins. Most of these sequences show other unusual sequence features, including multiple charge clusters and excessive counts of homopeptides of length > or = two amino acid residues. Intriguingly, a large majority of the identified Drosophila proteins are essential developmental proteins and, in particular, most play a role in central nervous system development. Almost half of the human and mouse proteins identified are homeotic homologs. The role of long homopeptides in fine-tuning protein conformation for multiple functional activities is discussed. The relative contributions of strand slippage and of dynamic mutation are also addressed. Several new experiments are proposed.
Synonymous codon usage is biased and the bias seems to be different in different organisms. Factors with proposed roles in causing codon bias include degree and timing of gene expression, codon – anticodon inter actions, transcription and translation rate and fidelity, codon context, and global and local G + C content. We offer a new perspective and new methods for elucidating codon choices applied especially to the human genome. We present data supporting the thesis that codon choices for human genes are largely a consequence of two factors: (1) amino acid constraints, (2) maintaining DNA structures dependent on base-step con-formational tendencies consistent with the organism's genome signature that is determined by genome-wide processes of DNA modification, replication and repair. The related codon signature defined as the dinucleotide relative abundances at the distinct codon positions {1, 2}, {2, 3}, and {3, 4} (4=1 of the next codon) accommodates both the global genome signature and amino acid constraints. In human genes, codon positions {2, 3} and {3, 4} containing the silent site have similar codon signatures reflecting DNA symmetry. Strong CG and TA dinucleotide underrepresen tation is observed at all codon positions as well as in non-coding regions. Estimates of synonymous codon usage based on codon signatures are in excellent agreement with the actual codon usage in human and general vertebrate genes. These properties are largely independent of the isochore compartment (G + C content), gene size, and transcriptional and translational constraints. We hypothesize that major influences on codon usage in human genes result from residue preferences and diresidue associations in proteins coupled to biases on the DNA level, related to replication and repair processes and/or DNA structural requirements.
Statistically significant charge clusters (basic, acidic, or of mixed charge) in tertiary protein structures are identified by new methods from a large representative collection of protein structures. About 10% of protein structures show at least one charge cluster, mostly of mixed type involving about equally anionic and cationic residues. Positive charge clusters are very rare. Negative (or histidine-acidic) charge clusters often coordinate calcium, or magnesium or zinc ions [e.g., thermolysin (PDB code: 3tln), mannose-binding protein (2msb), aminopeptidase (1amp)]. Mixed-charge clusters are prominent at interchain contacts where they stabilize quaternary protein formation [e.g., glutathione S-transferase (2gst), catalase (8act), and fructose-1,6-bisphosphate aldolase (1fba)]. They are also involved in protein-protein interaction and in substrate binding. For example, the mixed-charge cluster of aspartate carbamoyl-transferase (8atc) envelops the aspartate carbonyl substrate in a flexible manner (alternating tense and relaxed states) where charge associations can vary from weak to strong. Other proteins with charge clusters include the P450 cytochrome family (BM-3, Terp, Cam), several flavocytochromes, neuraminidase, hemagglutinin, the photosynthetic reaction center, and annexin. In each case in Table 2 we discuss the possible role of the charge clusters with respect to protein structure and function.
We present new methods for identifying and analyzing statistically significant residue clusters that occur in three-dimensional (3D) protein structures. Residue clusters of different kinds occur in many contexts. They often feature the active site (e.g., in substrate binding), the interface between polypeptide units of protein complexes, regions of protein-protein and protein-nucleic acid interactions, or regions of metal ion coordination. The methods are illustrated with 3D clusters centering on four themes. (i) Acidic or histidine-acidic clusters associated with metal ions. (ii) Cysteine clusters including coordination of metals such as zinc or iron-sulfur structures, cysteine knots prominent in growth factors, multiple sets of buried disulfide pairings that putatively nucleate the hydrophobic core, or cysteine clusters of mostly exposed disulfide bridges. (iii) Iron-sulfur proteins and charge clusters. (iv) 3D environments of multiple histidine residues. Study of diverse 3D residue clusters offers a new perspective on protein structure and function. The algorithms can aid in rapid identification of distinctive sites, suggest correlations among protein structures, and serve as a tool in the analysis of new structures.
Genomic similarities and contrasts are investigated in a collection of 23 bacteriophages, including phages with temperate, lytic, and parasitic life histories, with varied sequence organizations and with different hosts and with different morphologies. Comparisons use relative abundances of di-, tri-, and tetranucleotides from entire genomes. We highlight several specific findings. (i) As previously shown for cellular genomes, each viral genome has a distinctive signature of short oligonucleotide abundances that pervade the entire genome and distinguish it from other genomes. (ii) The enteric temperate double-stranded (ds) phages, like enterobacteria, exhibit significantly high relative abundances of GpC = GC and significantly low values of TA, but no such extremes exist in ds lytic phages. (iii) The tetranucleotide CTAG is of statistically low relative abundance in most phages. (iv) The DAM methylase site GATC is of statistically low relative abundance in most phages, but not in P1. This difference may relate to controls on replication (e.g., actions of the host SeqA gene product) and to MutH cleavage potential of the Escherichia coli DAM mismatch repair system. (v) The enteric temperate dsDNA phages form a coherent group: they are relatively close to each other and to their bacteria] hosts in average differences of dinucleotide relative abundance values. By contrast, the lytic dsDNA phages do not form a coherent group. This difference may come about because the temperate phages acquire more sequence characteristics of the host because they use the host replication and repair machinery, whereas the analyzed lytic phages are replicated by their own machinery. (vi) The nonenteric temperate phages with mycoplasmal and mycobacterial hosts are relatively close to their respective hosts and relatively distant from any of the enteric hosts and from the other phages. (vii) The single-stranded RNA phages have dinucleotide relative abundance values closest to those for random sequences, presumably attributable to the mutation rates of RNA phages being much greater than those of DNA phages.
Evolution of Recombination 32 condition as in the haploid model: there must be relatively strong negative epistasis. Final note { The above analysis is consistent with the qualitative conclusions of Feldman et al. (1980) up to a point. They argued that a tightly linked modiier would invade if it increased recombination and " < 0 or if it decreased recombination and " > 0, \but if the linkage is loose enough (when R ? exists), it may be eliminated." Numerically evaluating their stability condition indicates that an R ? between 0 and 1 2 can exist (a switch in stability may occur) for both positive and negative epistasis. In contrast, this reanalysis indicates that an R ? between 0 and 1 2 can only exist when epistasis is negative. Modiier alleles that increase recombination are never favored if there is positive epistasis and are favored under negative epistasis only if the modiier is suuciently tightly linked. (1980) that W ij = W ji , W 2j = W 3j and W j2 = W j3 and deening the marginal tness of an allele j as W j = P 4 i=1 W 1i x i , we nd that the value of the characteristic polynomial evaluated at = 1 equals: At R = 0, this reduces to equation (11) of Feldman et al. (1980) under the assumption that there are no cis-trans eeects on viability, such that an AB=ab genotype has the same tness as an Ab=aB genotype (W 14 = W 23 ; but see Nordborg et al. 1995). At R = 0, as observed by Feldman et al. (1980), invasion of a modiier allele that increases recombination will occur only if ^ D < 0, while invasion of a modiier allele that decreases recombination will occur if ^ D > 0. At R = 1=2, equation (20) becomes Again, a switch in stability will occur between R = 0 and R = 1=2 only if the term in square brackets is negative. To leading order in , this term equals 2 ? 3W 12 + W 12 W 14 + O(). Deening W 12 = 1?s and W 14 = (1?s) 2 +", we regain condition (19). Thus, when mutation is weak relative to selection, a switch in stability occurs in the diploid model under the same Evolution of Recombination 30 that relate ^ x 1 and to …
Early biochemical experiments established that the set of dinucleotide odds ratios or 'general design' is a remarkably stable property of the DNA of an organism, which is essentially the same in protein-coding DNA, bulk genomic DNA, and in different renaturation rate and density gradient fractions of genomic DNA in many organisms. Analysis of currently available genomic sequence data has extended these earlier results, showing that the general designs of disjoint samples of a genome are substantially more similar to each other than to those of sequences from other organisms and that closely related organisms have similar general designs. From this perspective, the set of dinucleotide odds ratio (relative abundance) values constitute a signature of each DNA genome, which can discriminate between sequences from different organisms. Dinucleotide-odds ratio values appear to reflect not only the chemistry of dinucleotide stacking energies and base-step conformational preferences, but also the species-specific properties of DNA modification, replication and repair mechanisms.
RecA protein sequences from 62 eubacterial sources were compared with one another and relative to one archaebacterial RecA-like and a number of eukaryotic RecA-like sequences. Pairwise similarity scores were determined by a novel method based on significant segment pair alignment. The sequences of different species were grouped on the basis of mutually high similarity scores within groups and consistency of score ranges in comparison to other groups. Following this protocol, the gamma-proteobacteria can be subclassified into two major groups, those of mostly vertebrate hosts and those of mostly soil habitat. The alpha-proteobacterial sequences also divide into two distinct groups, whereas classification of the beta-proteobacteria is more complex. The gram-positive bacterial sequences split into three groups of low and three groups of high G+C genome content. However, neither the combined low-G+C-content nor the combined high-G+C-content group nor the aggregate of all gram-positive bacteria form homogeneous groups. The mycoplasma sequences score best with the Bacillus subtilis sequence, consistent with their presumed origin from a gram-positive ancestor. The eukaryotic RAD proteins generally show a single high-scoring segment pair with the proteobacterial RecA sequences around the ATP-binding domain. The bacteriophage T4 UvsX protein aligns best with RecA sequences on two segments disjoint from the ATP-binding domain. The distribution of the most highly conserved regions shared between RecA and noneubacterial RecA-like sequences suggests a mosaic character and evolution of RecA. The discussion considers some questions on the validity and consistency of bacterial classifications derived from RecA sequence comparisons.