The nucleotide sequence of the DNA of bacteriophage λ has been determined using the dideoxy chain termination method in conjunction with random cloning in M13 vectors. Various methods were studied for sequencing specific regions to complete the sequence, but all were much slower than the random approach. The DNA in its circular form contains 48,502 base-pairs. Open reading frames were identified and, where possible, ascribed to genes by comparing with the previously determined genetic map. The reading frames for 46 genes were clearly identified, though in about 20 the position of the protein initiation site could not be rigorously established. Probable positions for the kil, cIII and lom genes are suggested but remain uncertain. There are about 20 other unidentified reading frames that may code for proteins.
We present here the complete 16,338 nucleotide DNA sequence of the bovine mitochondrial genome. This sequence is homologous to that of the human mitochondrial genome (Anderson et al., 1981) and the genes are organized in virtually identical fashion. The bovine mitochondrial protein genes are 63 to 79% homologous to their human counterparts, and most of the nucleotide differences occur in the third positions of codons. The minimum rate of base substitution that accounts for the nucleotide differences in the codon third positions is very high: at least 6 × 10−9 changes per position per year. The bovine and human mitochondrial transfer RNA genes exhibit more interspecies variation than do their cytoplasmic counterparts, with the “TΨC” loop being the most variable part of the molecule. The bovine 12 S and 16 S ribosomal RNA genes, when compared with those from human mitochondrial DNA, show conserved features that are consistent with proposed secondary structure models for the ribosomal RNAs. Unlike the pattern of moderate-to-high homology between the bovine and human mitochondrial DNAs found over most of the genome, the DNA sequence in the bovine D-loop region is only slightly homologous to the corresponding region in the human mitochondrial genome. This region is also quite variable in length, and accounts for the bulk of the size difference between the human and bovine mitochondrial DNAs.
The complete sequence of the 16,569-base pair human mitochondrial genome is presented. The genes for the 12S and 16S rRNAs, 22 tRNAs, cytochrome c oxidase subunits I, II and III, ATPase subunit 6, cytochrome b and eight other predicted protein coding genes have been located. The sequence shows extreme economy in that the genes have none or only a few noncoding bases between them, and in many cases the termination codons are not coded in the DNA but are created post-transcriptionally by polyadenylation of the mRNAs.
An approach to DNA sequencing using chain-terminating inhibitors (Sanger et al., 1977) combined with cloning of small fragments of DNA in a single-stranded DNA bacteriophage is described. Random fragments from restriction enzyme digestion of the DNA are inserted into the EcoRI site of the modified bacteriophage M13mp2 (Gronenborn & Messing, 1978) using a linker oligonucleotide. Individual recombinant plaques are collected, 1-ml cultures grown, and the DNA isolated. A "flankingprimer" from the vector is used to determine a nucleotide sequence in each inserted DNA fragment by the chain-terminating method. This is a relatively rapid and simple method of accumulating sequence data. The 2771-nucleotide sequence of the largest MboI restriction enzyme fragment from human mitochondrial DNA was determined by this method.
Analysis of an almost complete mammalian mitochondrial DNA sequence has identified 23 possible tRNA genes and we speculate here that these are sufficient to translate all the codons of the mitochondrial genetic code. This number is much smaller than the minimum of 31 required by the wobble hypothesis. For each of the eight genetic code boxes with four codons for one amino acid we find a single specific tRNA gene with T in the first (wobble) position of the anticodon. We suggest that these tRNAs with U in the wobble position can recognize all four codons in these genetic code boxes either by a "two out of three" base interaction or by U.N wobble.
The human mitochondrial (mt) genome consists of a closed circular duplex DNA approximately 10 x 106 daltons and has been the most intensely studied animal mt genetic system. The positions of the origin of replication of H strand synthesis (Crews et al. 1979), the 12S and 16S ribosomal RNA genes (Robberson et al. 1972) and 19 tRNA genes (Angerer et al. 1976) have been located on the genetic map shown in Figure 1. A number of discrete products of mitochondrial protein synthesis have been demonstrated and three of them identified as subunits 1, 2 and 3 of the cytochrome oxidase complex (Hare et al. 1980). In comparison with other mito-systems, genes for up to four subunits of the ATPase complex, one of the cytochrome bc1 complex and possibly for a ribosomal protein would be expected to be present (see review by Borst 1977). Both strands are thought to be completely transcribed symmetrically from a point near the origin of the H strand synthesis (Aloni and Attardi 1971; Murphy et al. 1975). These transcripts are then processed to give the rRNAs, the tRNAs and a number of polyadenylated but not capped mRNAs (Attardi et al. 1979). Both the L and H strands have been shown to be coding with the L strand containing the sense sequence of the rRNA genes, most of the tRNA genes and most of the stable polyadenylated mRNAs.
The nucleotide sequence of the region coding for the F protein of bacteriophage φX174 has now been completed, in conjunction with extended peptide sequence analysis of the gene F protein. The protein is 426 amino acids in length, with a molecular weight, calculated from the sequence, of 48,340.
The complete nucleotide sequence of the DNA of bacteriophage φX174 has been determined. The provisional sequence (Sanger et al., 1977a) deduced largely by the plus and minus method, has been completed and confirmed, predominantly using the terminator method (Sanger et al., 1977b). About 30 revisions were found to be necessary in the 5386-nucleotide sequence. The amino acid sequences of the ten proteins for which the DNA codes have also been deduced.
A new method for determining nucleotide sequences in DNA is described. It is similar to the "plus and minus" method [Sanger, F. & Coulson, A. R. (1975) J. Mol. Biol. 94, 441-448] but makes use of the 2',3'-dideoxy and arabinonucleoside analogues of the normal deoxynucleoside triphosphates, which act as specific chain-terminating inhibitors of DNA polymerase. The technique has been applied to the DNA of bacteriophage varphiX174 and is more rapid and more accurate than either the plus or the minus method.
A DNA sequence for the genome of bacteriophage phi X174 of approximately 5,375 nucleotides has been determined using the rapid and simple 'plus and minus' method. The sequence identifies many of the features responsible for the production of the proteins of the nine known genes of the organism, including initiation and termination sites for the proteins and RNAs. Two pairs of genes are coded by the same region of DNA using different reading frames.
Sequence determination of mutants in genes A and B confirms that these constitute another pair of genes in φX174 which are translated from the same DNA sequence using different reading frames.
Application of three different methods to obtain nucleotide sequences from X174 DNA has yielded a continuous sequence of 281 nucleotides and partial sequences totalling another 263 nucleotides from adjoining regions. Simultaneous work on amino acid sequences of X174 coat proteins showed that these nucleotide sequences are part of the F gene, coding for the largest coat protein of the bacteriophage. A total of 136 codons has been identified, of which 56% have thymine as the base in the third position of the coding triplet.
The nucleotide sequence of the coding region of gene G of φX174 and the amino acid sequence of the G-coded “spike” protein of the virion have now been completed. From the 5′ A of the initiating ATG to the 3′ end of the terminator triplet, the gene consists of 528 nucleotides and codes for a protein of 175 amino acids, molecular weight 19,053.
A simple and rapid method for determining nucleotide sequences in single-stranded DNA by primed synthesis with DNA polymerase is described. It depends on the use of Escherichia coli DNA polymerase I and DNA polymerase from bacteriophage T4 under conditions of different limiting nucleoside triphosphates and concurrent fractionation of the products according to size by ionophoresis on acrylamide gels. The method was used to determine two sequences in bacteriophage φX174 DNA using the synthetic decanucleotide A-G-A-A-A-T-A-A-A-A and a restriction enzyme digestion product as primers.
DNA, the chemical component of the gene, plays a central role in biology and contains the whole information for the development of an organism, coded in the form of sequences of the four nucleotide residues. The lecture describes the development and application of some methods that can be employed to deduce sequences in these very large molecules. Special attention has been applied to a rapid simple method in which DNA polymerase is primed with specific oligonucleotide primers, thus making it possible to study small sections of radioactively labelled DNA. The techniques have been applied to the single-stranded DNA of bacteriophage ϕX 174, and two sequences of about 250 nucleotides long have been deduced and related to the amino acid sequences of the proteins for which they code.
Gene G of bacteriophage φX174 DNA has been characterized by three techniques. (1) Transcription by RNA polymerase of appropriate DNA fragments from restriction enzyme digestion. (2) Use of such DNA fragments as primers for DNA polymerase. (3) Amino acid sequencing of the gene product. The combined results give a DNA sequence of 195 nucleotides which codes for the N-terminal 65 amino acids of the protein sequence.
Determination of nucleotide sequences in RNA has been facilitated by the development of small-scale techniques for fractionating 32P-labelled oligonucleotides produced by digestion with ribonucleases, particularly ribonuclease T1. Two-dimensional fractionations on modified papers and thin layers by ionophoresis and chromatography give "fingerprints", which may be used to characterise a given RNA and to isolate pure oligonucleotides as a first step in sequence analysis. Micro methods have been developed for determining the composition and sequence of the isolated oligonucleotides. These techniques have made it possible to determine the complete sequences of RNAs up to about 180 nucleotides long and to deduce a large part of the sequence of an RNA bacteriophage (3300 nucleotides long). Progress in the determination of sequences in DNA has been slower than in RNA due to the large size of the simplest DNA molecules and the lack of suitable deoxyribonucleases for specific degradation. However the small-scale techniques have recently been extended to the study of 32p-labe lied DNA and it has been possible to determine two sequences of about 50 residues from bacteriophage øX. The methods used and other possible approaches to DNA sequencing will be discussed. Another method has been developed in which a specific region in a DNA molecule is sequenced by a copying procedure using DNA polymerase primed by a specific oligonucleotide. In this way a sequence of 50 residues was deduced in bacteriophage f1 DNA.