The ARB (from Latin arbor, tree) project was initiated almost 10 years ago. The ARB program package comprises a variety of directly interacting software tools for sequence database maintenance and analysis which are controlled by a common graphical user interface. Although it was initially designed for ribosomal RNA data, it can be used for any nucleic and amino acid sequence data as well. A central database contains processed (aligned) primary structure data. Any additional descriptive data can be stored in database fields assigned to the individual sequences or linked via local or worldwide networks. A phylogenetic tree visualized in the main window can be used for data access and visualization. The package comprises additional tools for data import and export, sequence alignment, primary and secondary structure editing, profile and filter calculation, phylogenetic analyses, specific hybridization probe design and evaluation and other components for data analysis. Currently, the package is used by numerous working groups worldwide.
Comparative sequence analysis of small subunit rRNA is currently one of the most important methods for the elucidation of bacterial phylogeny as well as bacterial identification. Phylogenetic investigations targeting alternative phylogenetic markers such as large subunit rRNA, elongation factors, and ATPases have shown that 16S rRNA-based trees reflect the history of the corresponding organisms globally. However, in comparison with three to four billion years of evolution the phylogenetic information content of these markers is limited. Consequently, the limited resolution power of the marker molecules allows only a spot check of the evolutionary history of microorganisms. This is often indicated by locally different topologies of trees based on different markers, data sets or the application of different treeing approaches. Sequence peculiarities as well as methods and parameters for data analysis were studied with respect to their effects on the results of phylogenetic investigations. It is shown that only careful data analysis starting with a proper alignment, followed by the analysis of positional variability, rates and character of change, testing various data selections, applying alternative treeing methods and, finally, performing confidence tests, allows reasonable utilization of the limited phylogenetic information.