Existing methods based on homology rely on current research in genome analysis using n-grams (i.e. breaking the genome up into "words or "syllables"), protein motifs, and other bio-linguistic techniques have shown promise. In particular, as new protein structures and functions are identified, these bio-linguistic approaches can reach across multiple genomes to identify similar genes, elucidating their functions. Likewise new genes or disease gene variations identified through sequencing of individuals can be compared to known genes for identification of changes to their "normal" functions. In this review, we describe algorithms for searching biological databases using the n-gram analysis. Our results demonstrate that these algorithms are more sensitive than those currently available for both genomics and proteomics analysis, allowing a more accurate portrayal of similarity of gene function. The algorithm's capabilities extend to the comparison of biological sequences using phylogenetic and bio-chemical properties that enable the results to be significant from perspective of structure and function of genomic and proteomic data analysis. Recent years have seen an explosive growth in the speed and capacity of data collection and storage devices. The biological databases are experiencing an unprecedented growth where they are doubling every fifteen months. The algorithms described are amenable to parallelization with effective domain database partitioning. This makes them an attractive alternative for searching protein databases by developing high-speed functionally partitioned searches.