By analyzing 1,780,295 5'-end sequences of human full-length cDNAs derived from 164 kinds of oligo-cap cDNA libraries, we identified 269,774 independent positions of transcriptional start sites (TSSs) for 14,628 human RefSeq genes. These TSSs were clustered into 30,964 clusters that were separated from each other by more than 500 bp and thus are very likely to constitute mutually distinct alternative promoters. To our surprise, at least 7674 (52%) human RefSeq genes were subject to regulation by putative alternative promoters (PAPs). On average, there were 3.1 PAPs per gene, with the composition of one CpG-island-containing promoter per 2.6 CpG-less promoters. In 17% of the PAP-containing loci, tissue-specific use of the PAPs was observed. The richest tissue sources of the tissue-specific PAPs were testis and brain. It was also intriguing that the PAP-containing promoters were enriched in the genes encoding signal transduction-related proteins and were rarer in the genes encoding extracellular proteins, possibly reflecting the varied functional requirement for and the restricted expression of those categories of genes, respectively. The patterns of the first exons were highly diverse as well. On average, there were 7.7 different splicing types of first exons per locus partly produced by the PAPs, suggesting that a wide variety of transcripts can be achieved by this mechanism. Our findings suggest that use of alternate promoters and consequent alternative use of first exons should play a pivotal role in generating the complexity required for the highly elaborated molecular systems in humans.
Characterization of dedifferentiated chondrocytes (DECs) and mesenchymal stem cells capable of differentiating into chondrocytes is of biological and clinical interest. We isolated DECs and bone marrow stromal cells (BMSCs), H4-1 and H3-4, and demonstrated that the cells started to produce extracellular matrices, such as type II collagen and aggrecan, at an early stage of chondrosphere formation. Furthermore, cDNA sequencing of cDNA libraries constricted by the oligocapping method was performed to analyze difference in mRNA expression profiling between DECs and marrow stromal cells. Upon redifferentiation of DECs, cartilage-related extracellular matrix genes, such as those encoding leucine-rich small proteoglycans, cartilage oligomeric matrix protein, and chitinase 3-like 1 (cartilage glycoprotein-39), were highly expressed. Growth factors such as FGF7 and CTGF were detected at a high frequency in the growth stage of monolayer stromal cultures. By combining the expression profile and flow cytometry, we demonstrated that isolated stromal cells, defined by CD34−, c-kit−, and CD140α− or low, have chondrogenic potential. The newly established human mesenchymal cells with expression profiling provide a powerful model for a study of chondrogenic differentiation and further understanding of cartilage regeneration in the means of redifferentiated DECs and BMSCs.
As a base for human transcriptome and functional genomics, we created the “full-length long Japan” (FLJ) collection of sequenced human cDNAs. We determined the entire sequence of 21,243 selected clones and found that 14,490 cDNAs (10,897 clusters) were unique to the FLJ collection. About half of them (5,416) seemed to be protein-coding. Of those, 1,999 clusters had not been predicted by computational methods. The distribution of GC content of nonpredicted cDNAs had a peak at ∼58% compared with a peak at ∼42%for predicted cDNAs. Thus, there seems to be a slight bias against GC-rich transcripts in current gene prediction procedures. The rest of the cDNAs unique to the FLJ collection (5,481) contained no obvious open reading frames (ORFs) and thus are candidate noncoding RNAs. About one-fourth of them (1,378) showed a clear pattern of splicing. The distribution of GC content of noncoding cDNAs was narrow and had a peak at ∼42%, relatively low compared with that of protein-coding cDNAs.
Gene expression of synoviocytes stimulated with tumor necrosis factor-alpha (TNFalpha) was studied by macroarray analysis to elucidate the cellular response and identify new biological functions of known and unknown genes. 10035 cDNA clones were used to make cDNA macroarrays of representative genes. Synoviocytes expressed large amounts of fibronectin and collagen mRNA. Statistical analysis of the macroarray data revealed 26 genes, including six new genes, which underwent significant alteration of gene expression in response to TNFalpha stimulation. These findings suggest that the synoviocyte response to TNFalpha stimulation forms the basis of development of various aspects of the pathophysiology of rheumatoid arthritis.
Conventionally, an amino acid frame display has generally been used for the extraction of amino acid sequence from a cDNA sequence. In the frame each position of initiation and termination codon is displayed and a segment that starts at an initiation codon and terminates at a termination codon is identied; the obtained segments are identied as possible open reading frames (ORF), and among them, the longest ORF is identied as an amino acid sequence extracted from the cDNA. In the case where a frame shift error exists on a cDNA sequence, an ORF is split and displayed over 2 frames. Further, since the border of the split ORF is not clear, an amino acid sequence is, in general, identied with an error of tens of bases. It has been reported that statistical information included in a DNA sequence, such as coding potential, can be used to identify cloning errors including frame shifts [2]. Dr. Hirosawa showed that the application of a modied GeneMark program for detection of artifacts in cDNA clones. This program serves to provide a warning when any spurious split of protein-coding regions is detected. Though this method is eectiv e for detecting the split of protein-coding regions, it is dicult to detect the strict location of the frame-shift, because of the limitation of the statistical analysis. The most reliable method to identify the frame-shift errors in a DNA sequence is to use similarity information to known amino acid sequences. Methods of comparing a cDNA sequence with amino acid sequences in consideration of the occurrence of frame-shift errors in the DNA sequence have been developed including FASTY [5] and TRANS series developed by our laboratory [3]. Using
The Helix Research Institute (HRI) in Japan is releasing 4356 HUman Novel Transcripts and related information in the newly established HUNT database. The institute is a joint research project principally funded by the Japanese Ministry of International Trade and Industry, and the clones were sequenced in the governmental New Energy and Industrial Technology Development Organization (NEDO) Human cDNA Sequencing Project. The HUNT database contains an extensive amount of annotation from advanced analysis and represents an essential bioinformatics contribution towards understanding of the gene function. The HRI human cDNA clones were obtained from full-length enriched cDNA libraries constructed with the oligo-capping method and have resulted in novel full-length cDNA sequences. A large fraction has little similarity to any proteins of known function and to obtain clues about possible function we have developed original analysis procedures. Any putative function deduced here can be validated or refuted by complementary analysis results. The user can also extract information from specific categories like PROSITE patterns, PFAM domains, PSORT localization, transmembrane helices and clones with GENIUS structure assignments. The HUNT database can be accessed at http://www.hri.co.jp/HUNT.
Annotation and database system of full-length cDNA sequences was developed. As the components of the system, ORF annotation system, functional annotation system based on database search results, mapping annotation system, and integrated retrieval and display system were developed. In the ORF annotation system integrated analyses using conventional tools are performed and useful retrieval interface using motif list are introduced. In the functional annotation system based on database search results, a new method that characterizes a given unknown cDNA was developed by using a profile of similarity level over words appearing in sequence database entries. In the mapping annotation system, we linked by similarity searches full-length cDNA sequences with database DNA sequences that are already mapped on chromosomes. By using these links, full-length cDNAs can be retrieved by the retrieval condition of physical mapping information. Genetic disease information mapped on the physical mapping site can also be displayed by this system. Furthermore, we constructed an integrated database system for these analyzed data, and thus enabled annotation and selection of full-length cDNAs from points of both gene function and mapping information.
With progress of the human genome project, the whole sequences of the human genome will be determined completely in a few years. Functional analysis of genes of human genome is now being accelerated. Full-length cDNA clones are now being collected because they can be used as a starting material for functional analyses of genes. The oligo-capping cDNA library developed by Maruyama and Sugano is an effective source of full-length cDNA clones [1]. The aim of this study is to construct an integrated database of full-length cDNA sequences obtained from oligo-capping cDNA library for the functional analysis of genes. For this purpose we introduced a new annotation method using expression profile information obtained from cDNA sequence databases. The database system developed here has a function to retrieve the tissue specific genes. We also proposed a new method that can statistically compare the frequency distributions of gene expression over tissues. Finally the validity of this method was tested using known tissue specific genes. This study is a part of a project to determine the fulllength cDNA clones and construct a cDNA database system financed by New Energy and Industrial Technology Developmental Organization (NEDO).
Sumio Sugano合作论文数Department of Medical Genome Sciences, Graduate School of Frontier Sciences, The University of Tokyo
4
Kenta Nakai合作论文数Laboratory of Functional Analysis in silico
Human Genome Center
The Institute of Medical Science
The University of Tokyo2