
This chapter presents a review of the mathematical techniques available to construct phylogenies and to represent reticulate evolution. Phylogenies can be estimated using distance-based, maximum parsimony, or maximum likelihood methods. Bayesian methods have recently become available to construct phylogenies. Reticulate evolution includes horizontal gene transfer between taxa, hybridization events, and homoplasy. Genetic recombination also creates reticulate evolution within lineages. Several methods are now available to construct reticulated networks of various kinds. Twelve such methods and the accompanying software are described in this review chapter.
The Pathway Tools software allows a group of scientists to create, update, and publish on the Web an evolving knowledge resource describing the genome and biochemical networks of the organism. Such a knowledge resource will minimize duplication of experimental effort, ensure that all relevant knowledge will be brought to bear on interpreting new experimental results, and permit system-level computational analyses. Creation of a new Pathway/Genome Database (PGDB) by Pathway Tools includes inference of fungal metabolic pathways and pathway hole fillers (genes that code for enzymes missing from predicted pathways). Pathway Tools also infers the transport reactions present in an organism. A collection of interactive editing tools allows refinement of a PGDB by adding or modifying gene functions or pathways to capture knowledge from the biomedical literature. Pathway Tools provides a variety of query and visualization capabilities including a genome browser, displays of biochemical pathways, and a visualization of the cellular biochemical network. The Omics Viewer paints multiple types of functional genomics data onto that cellular network diagram. Comparative genomics capabilities allow comparison with other fungal Pathway/Genome Databases.
In a companion chapter in this volume, Wilson et al. (this volume, chapter by Wilson et al.) provide a detailed account of the experimental design and statistical analysis of microarray data. Their chapter is of interest to researchers planning microarray experiments capable of yielding data that can be statistically analyzed to insure reliable levels of confidence. In contrast, the present chapter emphasizes what can be done with the gathered data so as to simplify the huge task of interpreting the expression levels of tens of thousands of genes. In the companion chapter the authors assume the availability of statistical programs that are often used in the design of experiments. In this chapter we explore in greater detail the algorithms that process the collected data to obtain further information about cell behavior. Many of the algorithms described here aim at grouping similar data. We also explore microarray usage that is not addressed in the companion chapter.
Secreted proteins play critical biological roles in fungal species. Here we review and assess computational protocols for the identification of secreted proteins using their amino acid sequences. Protein sequences are screened for the presence of secretory signal peptides and the lack of features that prevent the delivery of proteins to the extracellular space, such as transmembrane segments, C-terminal ER-retention signals, or glycosylphosphatidylinositol (GPI) anchors. We apply such techniques to the complete genomes of 10 fungal species, identifying their putative complete sets of secreted proteins (their secretomes). Particular attention is given to predictions for the yeast secretome, which can be validated using the curated subcellular localizations of proteins from yeast. We make distinctions between the soluble and non-soluble portions of secretomes, discussing the roles of putative GPI proteins in the fungal cell wall.
A growing pool of genomic data is being archived to online public databases. These have the potential to impact both the diagnosis of genetically linked diseases as well as to aid in defining their genesis. The current in-silico tools for cytogenetic analysis have generally been targeted to the academic and industrial communities for use in research into the origins of disease. In comparison, few tools have been developed to integrate the multiple online resources for the clinical setting. We have addressed this deficit through a web-delivered application, LARaLINK 2.0: Loci Analysis for Rearrangements Link version 2.0 http://LARaLINK.bioinformatics.wai/ne.edu:8080/unigene controlled hierarchical vocabulary for mining cDNA and microarray expression data. This tool now provides researchers and clinicians with the means to effectively use cytogenetic data to rapidly assess disease association. The investigator is delivered a defined set of candidate disease genes together with the supporting evidence for their expression and disrupted phenotypes.
Computational biology has revolutionized biological and medical research. In the last two decades, a large number of computer methods have been developed to analyze DNA, RNA and protein sequences. These computer methods are playing a vital role in extracting useful information from sequences of genomes. These computational methods have been developed by different academic groups all over the world to serve the biological community. The methods are available as stand alone programs or on-line web servers. Most of these software packages are available free for academicians (freeware). In this chapter, we have described the major computational methods available for biologists to extract information from sequences. This chapter covers computational methods f r i) genome annotation; ii) comparative genomics; iii) protein structure prediction; iv) functional classification of proteins; and v) identification of potential vaccine candidates. These software packages are not available from a single source so it is not easy for users to obtain the software of their interest. In order to overcome this problem, attempts have been made to collect and compile a list of free biological software programs that includes software at EMBL and Indiana University. A catalog of biological software (Biocatalog) is also available on the internet (Rodriguez-Tome 1998). Recently, a repository of free software in biology has been created at Institute of Microbial Technology, Chandigarh, India which contains more than 800 free software packages.
The advent of microarray technology has significantly changed the way we can quantitatively measure and observe gene expression at the mRNA level within a given biological sample of interest, allowing for the monitoring of tens to hundreds of thousands of genes within a single experiment. The two main array platforms are spotted two-colour arrays and one-colour in situ-synthesized arrays. Microarrays are used for a wide range of applications including gene annotation, investigation of gene-gene interactions, elucidation of gene regulatory networks and gene-expression profiling of Saccharomyces cerevisiae and other fungal organisms. Academic researchers and both the pharmaceutical and agricultural industries have an enormous interest in developing microarrays both as diagnostic tools and for use in basic research into how pathogens, such as fungi, interact with their host. Microarray experiments generate vast quantities of raw gene expression data, therefore good experimental design and statistical analysis is required for the extraction of accurate and useful information regarding the expression of genes. In this review we firstly provide an overview of the arrival and development of microarray technology. We then focus on the issues surrounding experimental design and the processing of microarray images, followed by a discussion on methods for cleaning and normalizing raw gene expression data and a final discussion of the importance statistical analysis plays in identifying differentially expressed genes.
We created an application called Sight, a Java (TM)-based package that provides a user-friendly interface to generate and connect agents for automatic genomic data mining without requiring programming skills from the user. Sight was originally developed to automate analysis of the human genome and attempts to generate web agents for fungus-related Internet resources revealed that some of those resources use new methods of representing the information they report, and some servers returned multiple intermediate pages leading towards their response, which created difficulties for automated recovery of results. Consequently, it was not possible to use effectively the old version of Sight so this version of the application was adapted with a little additional programming, creating a new version for which these features of the fungal genome servers do not represent a problem. The new version of Sight (v. 3.0.0) that is tailored to servers carrying fungal databases is freely available for download from the project website at these URLs: http://bioinformatics.org/jSight/ and http://jsight.sourceforge.net/index_SF.htm.
Research paradigms in modern biology are shifting from a single gene to a genome-wide scale. Two major contributions toward this new trend are large-scale genome sequencing and bioinformatics. Recently, bioinformatics has emerged as a new science field that provides computational tools for collecting and maintaining complex biological data. Along with an exponential accumulation of sequence data, many bioinformatics software and algorithms have been developed to assist in genome scale analyses. A comprehensive knowledge of these tools can help not only to understand gene functions and genome organizations, but also to provide an opportunity to develop new tools that can answer many biological questions.
Homology modelling has become a useful tool for the prediction of protein structure when only sequence data are available. Structural information is often more valuable than sequence alone for determining protein function. Homology modelling is potentially a very useful tool for the mycologist, as the number of fungal gene sequences available has exploded in recent years, whilst the number of experimentally determined fungal protein structures remains low. Programs available for homology modelling utilise different approaches and methods to produce the final model. Within each step of the homology modelling process, many factors affect the quality of the model produced, and appropriate selection of the program can significantly improve the quality of the model. This review discusses the advantages and limitations of the currently available methods and programs and provides a starting point for novices wishing to create a structural model. We have taken a practical approach as we hope to enable any scientist to utilise homology modelling as a tool for the analysis of their protein, or genome, of interest.
N-linked glycosylation is an essential modification of secretory and membrane proteins in all eukaryotic cells. Here, we review the current metabolic pathways of N-linked oligosaccharide biosynthesis in the endoplasmic reticulum and in the Golgi apparatus for yeasts: Saccharomyces cerevisiae and Schizosaccharomyces pombe, and higher eukaryotes: plants and human. The evolutionarily conserved proteins, processed in the cytosolic and the luminal side of the ER membrane, and the unique genes and their specific functions, occurring in the Golgi complex, for each selected organism, will be collated and discussed. This precise knowledge of the glycosylation pathway contributes to better understanding of the N-linked glycoprotein biosynthesis among different species, resulting in the recently successfully engineered strains for heterologous gene expression systems for industrial and therapeutic protein production.
The term "genome annotation" includes identification of protein-coding and noncoding sequences (e.g., repeats, rDNA, and ncRNA) in genome assemblies and attaching functional information (metadata) to these annotated features. Here, we describe the basic outline of fungal nuclear and mitochondrial genome annotation as performed at the US Department of Energy Joint Genome Institute (JGI).
Existing methods based on homology rely on current research in genome analysis using n-grams (i.e. breaking the genome up into "words or "syllables"), protein motifs, and other bio-linguistic techniques have shown promise. In particular, as new protein structures and functions are identified, these bio-linguistic approaches can reach across multiple genomes to identify similar genes, elucidating their functions. Likewise new genes or disease gene variations identified through sequencing of individuals can be compared to known genes for identification of changes to their "normal" functions. In this review, we describe algorithms for searching biological databases using the n-gram analysis. Our results demonstrate that these algorithms are more sensitive than those currently available for both genomics and proteomics analysis, allowing a more accurate portrayal of similarity of gene function. The algorithm's capabilities extend to the comparison of biological sequences using phylogenetic and bio-chemical properties that enable the results to be significant from perspective of structure and function of genomic and proteomic data analysis. Recent years have seen an explosive growth in the speed and capacity of data collection and storage devices. The biological databases are experiencing an unprecedented growth where they are doubling every fifteen months. The algorithms described are amenable to parallelization with effective domain database partitioning. This makes them an attractive alternative for searching protein databases by developing high-speed functionally partitioned searches.
As scavengers of recalcitrant polymers in the nature, filamentous fungi are excellent secretors of proteins outside the growing mycelium. This characteristic has been targeted and systematically improved in industrially-exploited production strains. Over the last five years there has been a significant shift from one-gene-at-a-time approaches to wider understanding of the organism, made possible, for example by gene array and proteome technologies that can now also be applied to filamentous fungi. This has presented novel opportunities for studies into gene regulation under specific conditions such as a particular carbon source or developmental stage with a view of advancing the basic knowledge and gaining information that can be applied for strain and process improvement. Filamentous fungi offer enormous potential for efficient and large scale production of heterologous gene products. Importantly, protein secretion provides a platform for the eukaryotic style post-translational modification of gene products. Fungi are cheap to cultivate and down-stream processing is made easy with no need to break cells open for product recovery. In order to capitalize on fungi as heterologous production hosts, research is now directed to revealing the cellular mechanisms for internal protein quality control, secretion stress, functional genomics of protein expression and secretion, protein modification and linking the physiology to productivity.
Since the late 1970s it had become clear that the coding sequence of many eukaryotic genes is disrupted by genetic elements, intervening sequences, which must be removed prior to host gene function. Intervening sequences can be classified into introns and inteins. Introns are excised from the primary RNA transcript by a process termed splicing, whilst inteins are transcribed and translated together with their host protein and are removed from the unprocessed protein. Based on sequence homology, secondary structure and the splicing mechanism, introns can be classified into spliceosomal mRNA introns, group I introns and group II introns. Recent annotations of several fungal genomes revealed that introns and inteins are integral elements of fungal genes. These intervening sequences perform various important functions. They are carriers of transcription regulatory elements, contain signals for mRNA stability and export from the nucleus and participate in gene evolution.
EST (expressed sequence tag) technology has long been used for gene discovery. As more and more EST data have become publicly available, the usage of ESTs has expanded to other areas, such as in silico genetic marker discovery, in silico gene discovery, construction of gene models, alternative splicing prediction, genome annotation, expression profiling, and comparative genomics. In comparison with whole genome sequencing, EST technology is simpler and less costly, especially in the case of large genomes. Moreover, since ESTs represent "the expressed parts" of genomes, they are more immediately informative about the transcriptomes. On the other hand, ESTs are not suitable for the studies related to "the control parts" of genomes, such as promoters and transcription enhancing/inhibiting elements. In addition, information for rarely expressed genes is also difficult to mine from EST data. EST data mining requires bioinformatics resources such as databases, data retrieving tools and analysis algorithms. Bioinformatics tools are also required to deal with EST errors and contaminations. Many of these tools are freely available to the academic community. In this chapter, EST data resources, tools used for processing and mining EST data, and applications of ESTs in genomics, particularly, fungal genomics, are reviewed.
Fungal carotenoids are synthesized by the isoprenoid pathway with isopentenyl pyrophosphate as the general precursor. They are found in all divisions of the fungal realm, and several are at the edge of being exploited at an industrial scale for satisfying an increasing demand for carotenoid pigments, food and feed additives, and components of cosmetics and pharmaceuticals. Fungi as carotenoid source are highly appealing. At least prospectively, they should be easier amenable to genetic manipulation than plants, and thus will allow tailoring of specially designed substances. Genes for carotenoid synthesis were cloned from many different fungi. In order to stimulate further functional studies on genetic pathways for internal and environmental regulation of carotene synthesis, modification and degradation, an overview on the situation in the most thoroughly studied model organisms is presented. The role of carotenoids as antioxidants, light protective substances and as signalling compounds is discussed.
There is now a sufficient number of filamentous fungal genomes in the public databases to warrant at least initial comparisons with animal and plant genomes. Our interest lies in the control of multicellular morphogenesis, which is a feature of filamentous ascomycetes and basidiomycetes. Search of a representative collection of filamentous fungal genomes with gene sequences generally considered to be essential and highly conserved components of normal development in animals failed to reveal any homologies. We conclude that fungal and animal lineages diverged from their common opisthokont line well before the emergence of any multicellular arrangement, and that the unique cell biology of filamentous fungi has caused control of multicellular development in fungi to evolve in a radically different fashion from that in animals and plants.