The mouse is an important model organism in the Human Genome Project and is set to play a pivotal role in the comparative analysis of diverse genomes that will be increasingly important for studies of gene function as we move post-genomics. In mouse genetics, it is recognised that the systematic generation of new mouse mutations along with identification of the underlying genes is an important challenge for gene function studies in the post-genomics era. Large numbers of mouse mutations affecting a plethora of biological pathways and with a diverse range of phenotypes need to be generated and mapped in the mouse genome. Many mouse mutations will be homologues of known human genetic disease loci and uncovering the underlying gene to any mouse mutation will shed light on gene function both in the mouse, human and other species. Sequencing of the mouse genome and its comparison with human genome sequence will not only aid gene identification but will provide a rapid route to uncovering genes underlying any mouse mutation. Furthermore, the sequence comparison of two mammalian genomes can be expected to provide profound new insights into gene regulation and genome evolution.
This article aims to give an introduction to comparative sequencing and analysis, containing a brief section on sequence production, but focusing on the computer-based (in silico) analysis of genomic sequence.
We describe progress in a continuing project aimed at the generation of an overlapping cosmid DNA clone map of the short arm of human chromosome 11. The automated procedures used to prepare DNA samples and the computerized data collection and recording systems are described. We also demonstrate the use of the clones as reagents for the rapid isolation of genomic DNAs containing smaller probed regions. We have isolated approximately 4700 human cosmid DNA clones from mouse/human hybrid cell lines that contain predominantly human chromosomal region 11p. Of the DNA in the cell lines, 60% is derived from this chromosomal region, and the remaining 40% is derived from regions of chromosomes 3, 19, and 20. A total of 4159 clones have been fingerprinted to identify potential overlaps, and we have developed 535 sets (“contigs”). Using random modeling, it is estimated that 65% of 11p must be contained in the analyzed cosmids. The database of clones has been used to identify single or overlapping clones from noncosmid DNA probes. Examples are presented. It is proposed that cosmid reference filters be distributed to requesting laboratories.