Background The genomic integration of a lentiviral vector developed for the treatment of Wiskott-Aldrich syndrome (WAS) was assessed by localizing the vector insertion sites (IS) in a murine model of gene therapy for the disease.Methods Transduced hematopoietic progenitor cells were transplanted into mice or cultured in vitro. The IS were determined in the genomic DNA from blood, the bone marrow of the animals and from cultured cells.Results Sequencing vector-genomic DNA junctions yielded more than 150 IS of which 50-70% were located in transcription units. To obtain additional sequences from the population of cultured cells, we used a vector-tag concatenation technique providing 190 additional IS. Altogether, the profiles confirmed the bias for integration in transcription units. The vector did not congregate as hotspots and did not appear to target specific categories of genes. The diversity of the IS reflected the initial marking of a polyclonal population of cells. However, relatively few vector IS were found in vivo because only 30-40 unique IS were identified in each mouse using this approach. Although four to ten IS were shared by the blood and bone marrow, no common IS was found between mice or between any mouse and the cultured cells.Conclusions Taken as a whole, the pattern of genomic insertion of the WAS lentiviral vector was diverse and similar to that previously described for other HIV-1-derived lentiviral vectors. Testing cells destined for transplantation is unlikely to predict specific IS to be selected in vivo. Copyright (C) 2009 John Wiley & Sons, Ltd.
Retroviral vectors have induced subtle clonal skewing in many gene therapy patients and severe clonal proliferation and leukemia in some of them, emphasizing the need for comprehensive integration site analyses to assess the biosafety and genomic pharmacokinetics of vectors and clonal fate of gene-modified cells in vivo. Integration site analyses such as linear amplification-mediated PCR (LAM-PCR) require a restriction digest generating unevenly small fragments of the genome. Here we show that each restriction motif allows for identification of only a fraction of all genomic integrants, hampering the understanding and prediction of biological consequences after vector insertion. We developed a model to define genomic access to the viral integration site that provides optimal restriction motif combinations and minimizes the percentage of nonaccessible insertion loci. We introduce a new nonrestrictive LAM-PCR approach that has superior capabilities for comprehensive unbiased integration site retrieval in preclinical and clinical samples independent of restriction motifs and amplification inefficiency.
Retroviral vectors are commonly used gene delivery tools in clinical gene therapy providing stable integration and continuous gene expression of the transgene in the treated host cell. However, integration of the reverse transcribed vector DNA into the host genome is, by itself, a mutagenic eventthat may directly contribute to severe adverse events. The latter has dramatically been obbserved in individual cases in several, otherwise successful, gene therapy trials. Thus, a comprehensive analysis of the existing integration site pool in a transduced sample is indispensable to identify potential in vivo selection of affected cell clones and uncontrolled vector-induced cell proliferation. To date, there are several methods available to study the integration site distribution of retroviral vectors or other integrating elements as transposons. Each of these techniques makes use of restriction enzymes to digest the genomic DNA. To reveal particular vector integrations, a recognition motif of the used restriction enzyme has to be located in an appropriate distance to the integration locus in the host genome. Therefore, the genomic distribution of the recognition sequences directly impact the outcome of restriction enzyme dependent integration site analysis. We here report a validated genomic accessibility model which precisely determines the fraction of the human genome that can be analyzed with one reaction set up (i.e. restriction enzyme used). For our modeling, we used the clinically relevant linear amplification mediated PCR (LAM-PCR) as integration site analysis method of choice and the commonly used frequently cutting restriction enzymes (‘four-cutters’). We show that the most frequent four cutter motif (AATT) gives access to 54.5% of all possible integrations in the human genome, whereas the rarest distributed motif (CGCG) only identifies 2.9%. This restriction bias can be minimized by analyzing the same sample with different enzymes. A combination of the 5 most potent four cutter restriction enzymes gives access to 88.7% of the analyzable genome. Furthermore, we established an unbiased, non-restrictive integration site analysis technique based on (nr) LAM-PCR. Direct ligation of a single-stranded DNA sequence to the linear PCR product evades the need for restriction enzymes to recover integration sites. While standard LAM-PCR was done repeatedly with 3 different enzymes to detect integration sites present in lentivirally transduced single cell clones, nrLAM-PCR detected all integrations in these clones in one single reaction setup. This newly developed method comprehensively recovers genomic locations of integrating elements regardless of a restriction enzyme introduced bias. Our data show that the recovery rate of integration sites present in a transduced sample strongly depends on the restriction enzyme(s) used. However, we demonstrate that the genomic accessibility of viral integration sites indeed can be determined and minimized a priori, and that a non restrictive LAM-PCR approach circumvents the existing limitations. Analysis of the clonal inventory by these methods will allow determining the pharmacodynamics of insertional vectors with unprecedented precision, facilitating development and clinical testing of insertional vector systems.
The only natural mechanism of malaria transmission in sub-Saharan Africa is the mosquito, generally Anopheles gambiae. Blocking malaria parasite transmission by stopping the development of Plasmodium in the insect vector would provide a useful alternative to the current methods of malaria control. Toward this end, it is important to understand the molecular basis of the malaria parasite refractory phenotype in An. gambiae mosquito strains. We have selected and sequenced six bacterial artificial chromosome (BAC) clones from the Pen-1 region that is the major quantitative trait locus involved in Plasmodium encapsulation. The sequence and the annotation of five overlapping BAC clones plus one adjacent, but not contiguous clone, totaling 585kb of genomic sequence from the centromeric end of the Pen-1 region of the PEST strain were compared to that of the genome sequence of the same strain produced by the whole genome shotgun technique. This project identified 23 putative mosquito genes plus putative copies of the retrotransposable elements BEL12 and TRANSIBN1_AG in the six BAC clones. Nineteen of the predicted genes are most similar to their Drosophila melanogaster homologs while one is more closely related to vertebrate genes. Comparison of these new BAC sequences plus previously published BAC sequences to the cognate region of the assembled genome sequence identified three retrotransposons present in one sequence version but not the other. One of these elements, Indy, has not been previously described. These observations provide evidence for the recent active transposition of these elements and demonstrate the plasticity of the Anopheles genome. The BAC sequences strongly support the public whole genome shotgun assembly and automatic annotation while also demonstrating the benefit of complementary genome sequences and of human curation. Importantly, the data demonstrate the differences in the genome sequence of an individual mosquito compared to that of a hypothetical, average genome sequence generated by whole genome shotgun assembly.
We performed genome-wide sequence comparisons at the protein coding level between the genome sequences of Drosophila melanogaster and Anopheles gambiae. Such comparisons detect evolutionarily conserved regions (ecores) that can be used for a qualitative and quantitative evaluation of the available annotations of both genomes. They also provide novel candidate features for annotation. The percentage of ecores mapping outside annotations in the A. gambiae genome is about fourfold higher than in D. melanogaster. The A. gambiae genome assembly also contains a high proportion of duplicated ecores, possibly resulting from artefactual sequence duplications in the genome assembly. The occurrence of 4063 ecores in the D. melanogaster genome outside annotations suggests that some genes are not yet or only partially annotated. The present work illustrates the power of comparative genomics approaches towards an exhaustive and accurate establishment of gene models and gene catalogues in insect genomes.
Chromosome 14 is one of five acrocentric chromosomes in the human genome. These chromosomes are characterized by a heterochromatic short arm that contains essentially ribosomal RNA genes, and a euchromatic long arm in which most, if not all, of the protein-coding genes are located. The finished sequence of human chromosome 14 comprises 87,410,661 base pairs, representing 100% of its euchromatic portion, in a single continuous segment covering the entire long arm with no gaps. Two loci of crucial importance for the immune system, as well as more than 60 disease genes, have been localized so far on chromosome 14. We identified 1,050 genes and gene fragments, and 393 pseudogenes. On the basis of comparisons with other vertebrate genomes, we estimate that more than 96% of the chromosome 14 genes have been annotated. From an analysis of the CpG island occurrences, we estimate that 70% of these annotated genes are complete at their 5′ end.
The establishment of an exhaustive inventory of genesis the primary goal of genome sequencing projects. Whenlooking at multicellular genome annotations that areavailable in sequence data banks or on other sites, thelevel of available information is quite variable fromgenome to genome, and the degree of completion that hasbeen reached among the gene inventories of the genomessequenced to date is very difficult to assess. These geneinventories are typically carried out in an automated fashion by the annotation platforms of the major data banks,and rely mainly on two types of predictions: ab initio predictions and those based on sequence comparisons. A genomic DNA sequence can be subjected to direct or indirect comparisons. In direct comparisons the genomicDNA sequence is aligned with sequences of expressionproducts, namely ESTs, cDNAs, or proteins from thesame species. In indirect comparisons the genomic sequence is aligned with genomic or expressed sequencesfrom other organisms...