BACKGROUND:Understanding the molecular mechanisms involved in disease is critical for the development of more effective and individualized strategies for prevention and treatment. The amount of disease-related literature, including new genetic information on the molecular mechanisms of disease, is rapidly increasing. Extracting beneficial information from literature can be facilitated by computational methods such as the knowledge-discovery approach. Several methods for mining gene-disease relationships using computational methods have been developed, however, there has been a lack of research evaluating specific disease candidate genes.RESULTS:We present a novel method for gathering and prioritizing specific disease candidate genes. Our approach involved the construction of a set of Medical Subject Headings (MeSH) terms for the effective retrieval of publications related to a disease candidate gene. Information regarding the relationships between genes and publications was obtained from the gene2pubmed database. The set of genes was prioritized using a "weighted literature score" based on the number of publications and weighted by the number of genes occurring in a publication. Using our method for the disease states of pain and Alzheimer's disease, a total of 1101 pain candidate genes and 2810 Alzheimer's disease candidate genes were gathered and prioritized. The precision was 0.30 and the recall was 0.89 in the case study of pain. The precision was 0.04 and the recall was 0.6 in the case study of Alzheimer's disease. The precision-recall curve indicated that the performance of our method was superior to that of other publicly available tools.CONCLUSIONS:Our method, which involved the use of a set of MeSH terms related to disease candidate genes and a novel weighted literature score, improved the accuracy of gathering and prioritizing candidate genes by focusing on a specific disease.
A comprehensive gene-expression analysis during platelet (PLT) production from megakaryocytes may give important information on genes involved in the PLT production process. However, the low abundance of primary megakaryocytes makes the gene expression analysis difficult. Therefore, we employed MEG-01 cells, a human megakaryocytic cell line, and confirmed that the cell line produces PLT-like particles by treatment with phorbol myristate acetate (PMA). After treatment of MEG-01 cells with PMA for 8 or 24 h, comprehensive gene expression analysis was carried out using a microarray and Reverse Transcription-Polymerase Chain Reaction (RTPCR). From the microarray analysis, 141 genes were up-regulated (>2-fold) and 164 genes were down-regulated (<1/2-fold). However, known PLT-related genes were not included in the up- or down-regulated genes. On the other hand, RT-PCR analysis detected increased expression of beta 1-tubulin, CD62P, gpIb alpha and gpIII, which are related to PLT function and megakaryocyte differentiation, following PMA treatment for 24 h. These results indicate that the MEG-01 cell may be an alternative model system to study the process of human PLT production from megakaryocytes. The gene-expression analysis might be a powerful tool for identifying genes related to PLT production, if the experimental conditions are optimized.
Murine megakaryocytes (MKs) are defined by CD41/CD61 expression and acetylcholinesterase (AChE) activity; however, their stages of differentiation in bone marrow (BM) have not been fully elucidated. In murine lineage-negative (Lin(-))/CD45(+) BM cells, we found CD41(+) MKs without AChE activity (AChE(-)) except for CD41(++) MKs with AChE activity (AChE(+)), in which CD61 expression was similar to their CD41 level. Lin(-)/CD41(+)/CD45(+)/AChE(-) MKs could differentiate into AChE(+), with an accompanying increase in CD41/CD61 during in vitro culture. Both proplatelet formation (PPF) and platelet (PLT) production for Lin(-)/CD41(+)/CD45(+)/AChE(-) MKs were observed later than for Lin(-)/CD41(++)/CD45(+)/AChE(+) MKs, whereas MK progenitors were scarcely detected in both subpopulations. GeneChip and semiquantitative polymerase chain reaction analyses revealed that the Lin(-)/CD41(+)/CD45(+)/AChE(-) MKs are assigned at the stage between the progenitor and PPF preparation phases in respect to the many MK/PLT-specific gene expressions, including beta1-tubulin. In normal mice, the number of Lin(-)/CD41(+)/CD45(+)/AChE(-) MKs was 100 times higher than that of AChE(+) MKs in BM. When MK destruction and consequent thrombocytopenia were caused by an antitumor agent, mitomycin-C, Lin(-)/CD41(+)/CD45(+)/AChE(-) MKs led to an increase in AChE(+) MKs and subsequent PLT recovery with interleukin-11 administration. It was concluded that MKs in murine BM at least in part consist of immature Lin(-)/CD41(+)/CD45(+)/AChE(-) MKs and more differentiated Lin(-)/CD41(++)/CD45(+)/AChE(+) MKs. Immature Lin(-)/CD41(+)/CD45(+)/AChE(-) MKs are a major MK population compared with AChE(+) MKs in BM and play an important role in rapid PLT recovery in vivo.
Understanding the coupling specificity between G protein-coupled receptors (GPCRs) and specific classes of G proteins is important for further elucidation of receptor functions within a cell. Increasing information on GPCR sequences and the G protein family would facilitate prediction of the coupling properties of GPCRs. In this study, we describe a novel approach for predicting the coupling specificity between GPCRs and G proteins. This method uses not only GPCR sequences but also the functional knowledge generated by natural language processing, and can achieve 92.2% prediction accuracy by using the C4.5 algorithm. Furthermore, rules related to GPCR-G protein coupling are generated. The combination of sequence analysis and text mining improves the prediction accuracy for GPCR-G protein coupling specificity, and also provides clues for understanding GPCR signaling.
Platelets (PLT) are produced from megakaryocytes (Mks) via proplatelet formation (PPF). However, the molecular mechanisms from Mks to PPF are not clearly elucidated, because the maturational steps of the Mks in bone marrow (BM) are not analyzed in detail. Until now, mouse Mks have been only isolated as acetylcholinesterase (AchE) positive cells and they are understood as well maturated population. In this study, we found the presence of different megakaryocytic subpopulations in BM by flowcytometry. To isolate the Mks, first we depleted lineage marker (CD4, CD8a, CD11b, B220, CD71, CD90, TER119, Gr-1, F4/80, 7/4) positive cells from BM cells of BALB/c mice. The analysis of the expression-pattern of CD41, CD45 and CD61 in the lineage negative (Lin − ) cells showed the presence of two types of megakaryocytic subpopulations. By sorting, they were identified as Lin − CD41 + /45 + /61 + cells (AchE negative) and Lin − CD41 ++ /45 + /61 ++ cells (partially AchE positive), respectively. To assess the maturational stages of the subpopulations, each population was cultured with 10ng/mL of TPO followed by counting of PPF and PLT production. Both PPF and PLT production were observed in Lin − CD41 + /45 + /61 + cells later than those in Lin − CD41 ++ /45 + /61 ++ cells. On the other hand, CFU-Mk was scarcely detected in each subpopulation. The results indicate that both populations are the committed megakaryocytes and Lin − CD41 + /45 + /61 + cells are more immature population than Lin − CD41 ++ /45 + /61 ++ cells. Then to characterize these subpopulations in detail, gene expression profiling was performed against four-megakaryocytic lineage-populations, Lin − CD41 − Thy1 low c-kit + cells as stem/progenitor, Lin − CD41 + /45 + /61 + cells, Lin − CD41 ++ /45 + /61 ++ cells and PLT using GeneChipU74 or RT-PCR. These analyses revealed that many PLT-specific genes including gpIb/IX, P-selectin, thrombin-R and ADP-R were already expressed on Lin − CD41 + /45 + /61 + cells but less than Lin − CD41 ++ /45 + /61 ++ cells. Especially, beta-1 tubulin that is necessary for PPF was only expressed on Lin − CD41 ++ /45 + /61 ++ cells. On the contrary, the expression of c-kit gene was gradually decreasing from stem/progenitor fraction to PLT. In conclusion, we succeeded in the isolation of new subpopulations distinguishable between immature Mks and more matured Mks beginning to prepare PLT. The present finding can contribute to elucidate the molecular mechanisms during terminal maturation.
1. We have confirmed the Diabetes Mellitus OLETF type I ( Dmo1 ) effect on hyperphagia, dyslipidaemia and obesity in the Otsuka Long-Evans Tokushima Fatty (OLETF) strain. The critical interval was narrowed down to 570 kb between D1Got258 to p162CA1 by segregation analyses using congenic lines. 2. Within the critical 570 kb region of the Dmo1 locus, we identified the G-protein-coupled receptor gene GPR10 as the causative gene mutated in the OLETF strain. The ATG translation initiation codon of GPR10 is changed into ATA in this strain and, so, is unavailable for the initiation of translation. 3. The GPR10 protein has a cognate ligand, namely prolactin-releasing peptide (PrRP). Centrally administered PrRP suppressed the food intake of congenic rats that have a Brown Norway derived Dmo1 region (i.e. with wild-type GPR10 ), but did not suppress that of the OLETF strain, indicating that GPR10 is without function and could explain hyperphagia in the OLETF strain. 4. Moreover, when restricted in food volume to the same level consumed by the congenic strain, OLETF rats showed few differences in the parameters of dyslipidaemia and obesity compared with congenic strains. 5. Taken together, these results demonstrate that the mutated GPR10 receptor is responsible for the hyperphagia leading to obesity and dyslipidaemia in the obese diabetic strain rat.
1. Dmo1 (Diabetes Mellitus OLETF type I) is a major quantitative trait locus for dyslipidaemia, obesity and diabetes phenotypes of male Otsuka Long Evans Tokushima Fatty (OLETF) rats. 2. Our congenic lines, produced by transferring Dmo1 chromosomal segments from the non-diabetic Brown Norway (BN) rat into the OLETF strain, have confirmed the strong, wide-range therapeutic effects of Dmo1 on dyslipidaemia, obesity and diabetes in the fourth (BC4) and fifth (BC5) generations of congenic animals. Analysis of a relatively small number of BC5 rats (n = 71) suggested that the critical Dmo1 interval lies within a < 4.9 cM region between D1Rat461 and D1Rat459. 3. To confirm the assignment of the Dmo1 critical interval, we intercrossed BC5 animals to produce a larger study population (BC5:F1 males; n = 406). For the present study, we used bodyweight at 18 weeks of age as an index of obesity; this phenotype is representative of the closely associated dyslipidaemia and hyperglycaemia phenotypes. 4. Interval mapping assigned logarithm of odds (LOD) peaks at the D1Rat90 marker (LOD = 9.11). One LOD support interval lies within the < 1.7 cM region between D1Rat461 and D1Rat459. 5. This large intercross study confirms that Dmo1 is likely localized within the interval.
As a base for human transcriptome and functional genomics, we created the “full-length long Japan” (FLJ) collection of sequenced human cDNAs. We determined the entire sequence of 21,243 selected clones and found that 14,490 cDNAs (10,897 clusters) were unique to the FLJ collection. About half of them (5,416) seemed to be protein-coding. Of those, 1,999 clusters had not been predicted by computational methods. The distribution of GC content of nonpredicted cDNAs had a peak at ∼58% compared with a peak at ∼42%for predicted cDNAs. Thus, there seems to be a slight bias against GC-rich transcripts in current gene prediction procedures. The rest of the cDNAs unique to the FLJ collection (5,481) contained no obvious open reading frames (ORFs) and thus are candidate noncoding RNAs. About one-fourth of them (1,378) showed a clear pattern of splicing. The distribution of GC content of noncoding cDNAs was narrow and had a peak at ∼42%, relatively low compared with that of protein-coding cDNAs.
1. Whole-genome scans have identified Dmo1 as a major quantitative trait locus for dyslipidaemia and obesity in the Otsuka Long Evans Tokushima Fatty (OLETF) rat. 2. We have produced congenic rats for the Dmo1 locus through successive back-cross breeding with diabetic OLETF rats. Marker-assisted speed congenic protocols were applied to efficiently transfer chromosomal segments from non-diabetic Brown Norway (BN) rats into the OLETF background. 3. In the fourth generation of congenic animals, we observed a substantial therapeutic effect of the Dmo1 locus on lipid metabolism, obesity control and plasma glucose homeostasis. 4. We have concluded that Dmo1 primarily affects lipid homeostasis, obesity control and/or glucose homeostasis at fasting and is secondarily involved in glucose homeostasis after loading. 5. The results of the present study show that single-allele correction of a genetic defect of the Dmo1 locus can generate a substantial therapeutic effect, despite the complex polygenic nature of type II diabetic syndromes.
We have isolated more than 12,000 clones containing microsatellite sequences, mainly consisting of (CA)n dinucleotide repeats, using genomic DNA from the BN strain of laboratory rat. Data trimming yielded 9636 non-redundant microsatellite sequences, and we designed oligonucleotide primer pairs to amplify 8189 of these. PCR amplification of genomic DNA from five different rat strains yielded clean amplification products for 7040 of these simple-sequence-length-polymorphism (SSLP) markers; 3019 markers had been mapped previously by radiation hybrid (RH) mapping methods (Nat Genet 22, 27–36, 1998). Here we report the characterization of these newly developed microsatellite markers as well as the release of previously unpublished microsatellite marker information. In addition, we have constructed a genome-wide linkage map of 515 markers, 204 of which are derived from our new collection, by genotyping 48 F2 progeny of (OLETFxBN)F2 crosses. This map spans 1830.9 cM, with an average spacing of 3.56 cM. Together with our ongoing project of preparing a whole-genome radiation hybrid map for the rat, this dense linkage map should provide a valuable resource for genetic studies in this model species.
Biological processes are controlled by direct and specific molecular interactions, involving DNA, RNA and proteins. The integration of the structure and function of them with the knowledge of macromolecular interactions and networks is an important step towards the construction of a unified and physiological view of the organism. Most of the data resources on biological functions such as expressed patterns and interactions are still only in the biological literature. The rapid growth of these collections makes it difficult for human beings to access the required information in a convenient and effective manner. The problem is that most of the documents are written in a natural language that computers can not deal with easily, and it is time-consuming for human beings to extract the knowledge from them. Therefore, the biologists have started asking for an intelligent information extraction system be developed method to save time and labor [1, 2, 3]. Here, we propose a system for the information extraction of protein-protein interactions from scientific text. We think that the system presented here could support biological researchers in various situations, e.g. constructing a database on protein interaction.
A whole-genome radiation hybrid (RH) panel was used to construct a high-resolution map of the rat genome based on microsatellite and gene markers. These include 3,019 new microsatellite markers described here for the first time and 1,714 microsatellite markers with known genetic locations, allowing comparison and integration of maps from different sources. A robust RH framework map containing 1,030 positions ordered with odds of at least 1,000:1 has been defined as a tool for mapping these markers, and for future RH mapping in the rat. More than 500 genes which have been mapped in mouse and/or human were localized with respect to the rat RH framework, allowing the construction of detailed rat-mouse and rat-human comparative maps and illustrating the power of the RH approach for comparative mapping.
Whole genome sequences of several organisms have been completely determined by advances of genome sequencing projects. Consequently, many of the novel gene and protein structures have been also identified. However, many of genes or proteins have not been annotated with biological functions. Thus the next major challenge of genome sequencing project are to characterize the biological function of each protein and to elucidate its roll in various intracellular processes, such as translation, splicing, post-translational modification, post-translational transfer, and so on [1]. For identification of gene functions and elucidation of intracellular processes, many protein-protein interactions, which are not only physical interactions between two proteins, but also functional or genetic interactions, have been rapidly identifying and their experimental results have been accumulating in the public databases such as MIPS and YPD [2, 3]. We propose the computer software. Two major functions of this software are:
This paper describes on a developmental system WebPACADE which is an extended version of a deductive database system PACADE for the analysis of protein 3D structure. It enables structural similarity searches on various proteins in the level of secondary structure. Results of a similarity search can be displayed graphically in a WWW browser.