Background: With the advent of "omics" ( e. g. genomics, transcriptomics, proteomics and phenomics), studies can produce enormous amounts of data. Managing this diverse data and integrating with other biological data are major challenges for the bioinformatics community. Comprehensive new tools are needed to store, integrate and analyze the data efficiently.Description: The PhenoGen Informatics website http://phenogen.uchsc.edu is a comprehensive toolbox for storing, analyzing and integrating microarray data and related genotype and phenotype data. The site is particularly suited for combining QTL and microarray data to search for "candidate" genes contributing to complex traits. In addition, the site allows, if desired by the investigators, sharing of the data. Investigators can conduct "in-silico" microarray experiments using their own and/or "shared" data.Conclusion: The PhenoGen website provides access to tools that can be used for high-throughput data storage, analyses and interpretation of the results. Some of the advantages of the architecture of the website are that, in the future, the present set of tools can be adapted for the analyses of any type of high-throughput "omics" data, and that access to new tools, available in the public domain or developed at PhenoGen, can be easily provided.
IBM Research and the University of Colorado collaborated on their submission to the inaugural Genomics track at TREC 2003. IBM Research has extensive experience in natural language processing, text analysis, and large-scale systems [9, 13, 3, 5, 16, 10]. IBM also has numerous research and business activities in the broad areas of bioinformatics and bio-medical information processing [14, 8]. IBM Research is currently developing BioTeKS, a middleware system for text analysis, mining, and information retrieval in the bio-medical domain. The University of Colorado (CU) has been working in the area of bioinformatics and text analysis in the bio-medical domain for a number of years and has made substantial contributions to the field [7, 11, 15, 12]. CU contributed their domain expertise to enhance the BioTeKS system and jointly we designed and evaluated experiments while preparing our track submissions.
We studied contrast and variability in a corpus of gene names to identify potential heuristics for use in performing entity identification in the molecular biology domain. Based on our findings, we developed heuristics for mapping weakly matching gene names to their official gene names. We then tested these heuristics against a large body of Medline abstracts, and found that using these heuristics can increase recall, with varying levels of precision. Our findings also underscored the importance of good information retrieval and of the ability to disambiguate between genes, proteins, RNA, and a variety of other referents for performing entity identification with high precision.