Salamanders are the only living tetrapods capable of fully regenerating limbs. The discovery of salamander lineage-specific genes (LSGs) expressed during limb regeneration suggests that this capacity is a salamander novelty. Conversely, recent paleontological evidence supports a deeper evolutionary origin, before the occurrence of salamanders in the fossil record. Here we show that lungfishes, the sister group of tetrapods, regenerate their fins through morphological steps equivalent to those seen in salamanders. Lungfish de novo transcriptome assembly and differential gene expression analysis reveal notable parallels between lungfish and salamander appendage regeneration, including strong downregulation of muscle proteins and upregulation of oncogenes, developmental genes and lungfish LSGs. MARCKS-like protein (MLP), recently discovered as a regeneration-initiating molecule in salamander, is likewise upregulated during early stages of lungfish fin regeneration. Taken together, our results lend strong support for the hypothesis that tetrapods inherited a bona fide limb regeneration programme concomitant with the fin-to-limb transition.
What are coelacanths? Coelacanths are a curious group of fish, represented by only two extant species: the African coelacanth (Latimeria chalumnae) and the Indonesian coelacanth (Latimeria menadoensis). These large, lobe-finned fish live in and around deep-water caves off the coasts of southeastern Africa and Indonesia. Around two meters in length, the coelacanth looks like no other fish alive. In addition to its fleshy limb-like fins with their skeletal supporting structures, it has a unique bicaudal tail and a hinge on the top of its skull which allows it to expand its gape. When hunting they orient themselves vertically, allowing an electrosensitive rostral organ in their snout to assist in the detection of prey. Coelacanths are ovoviviparous, with their eggs developing and hatching in the oviduct before birth. Only two species? So, why the hype? A number of factors contributed to the fame of what might seem like a rather obscure fish. For one, this fish had been playing an epic 70 million year game of hide-and-seek. Coelacanths were a more abundant and diverse group prior to the extinction event at the end of the cretaceous period (yes, the one that killed the dinosaurs); coelacanth fossils are well represented worldwide, yet conspicuously absent in any rocks after that time. When a living coelacanth specimen was caught off the coast of South Africa and identified in 1938 by Marjorie Courtenay-Latimer and J.L.B. Smith, it was as surprising and unexpected as finding a T-rex, albeit perhaps slightly less intimidating. However, it was more than just the surprise factor that made the discovery of the coelacanth perhaps the most notable zoological find of the last century. The existence of living coelacanths offered the possibility of significant insights into the early origins of the tetrapods (Figure 1), the group comprising amphibians, reptiles, birds and mammals, which is to say, ultimately, insights into our own evolutionary origins. How so? Coelacanths emerged during a critical stage of vertebrate evolution, and are phylogenetically placed in the sarcopterygian lineage. Sarcopterygians, also known as lobe-finned vertebrates, comprise coelacanths, lungfish and tetrapods. The coelacanth lineage diverged from the tetrapods roughly 400 million years ago, making them a key resource for comparative genomics. Lungfish are located at an equally propitious phylogenetic position as coelacanths but have intractably large genomes, effectively ruling them out as candidates for whole genome sequencing. This particular window of evolutionary time is especially notable because a large number of key tetrapod innovations arose during that period. I hear the coelacanth genome has now been sequenced... Yes, the coelacanth genome makes it possible to identify, with greatly increased resolution, genomic changes that may underlie the various adaptations that accompanied the transition from life in water to a life on land. For example, we can now distinguish between genomic gains or losses that were specific to the tetrapods and those that were shared with the entire sarcopterygian lineage. So, was a coelacanth relative the first fish to crawl out of the water? The idea of the first fish to crawl onto the land is one that captures people’s imaginations. Identification of the closest living relative of the tetrapod ancestor has long vexed evolutionary biologists. The lobe-finned fishes, of which the coelacanth is a member, with their distinctive fins supported by limb-like bony arrangements were proposed early on, and the discussion moved on to establishing which lineage was the sister group to the tetrapods. The discovery of extant coelacanths made it possible to use molecular phylogenetic analysis in addition to morphological data. A number of studies tended to support the lungfish as our closest living relative, but prior to the publication of the coelacanth genome, no one had been able to rule out the possibility that both lungfish and the coelacanth were equally related to the tetrapods. We now know that the lungfish is definitively the closest living relative of the tetrapods (Figure 1). How do you make a fish into a land animal? The rise of terrestrial vertebrates is a fascinating success story of evolution. The obstacles for an invasion of the land by fish, exquisitely adapted for life in the water as they are, stretch the imagination. A body previously supported by the water column must now be able to support itself in air, necessitating limbs with strong skeletal elements rather than the delicate fin rays of bony fish; gills that efficiently extracted oxygen dissolved in water must be replaced with lungs for extracting oxygen from the air; and as water becomes a precious commodity that must be conserved, more efficient ways to excrete waste products that don’t rely on an unlimited water supply need to be found. Even the senses must be significantly overhauled to match the unique demands of the terrestrial environment. Don’t the coelacanth’s fins look almost like limbs already? Yes, sort of. But as with all lobe-finned fishes, they don’t have the proper digit field (autopod) and they possess fin rays, or dermal bone, at their distal ends; an arrangement that is good for swimming but not walking on land. Rather, the coelacanth’s fins might be thought of as containing the rudiments of the autopod structure. Loss of the fin rays occurred in this evolutionary transition and an inkling of how this could have taken place can still be found by comparing the genomes of the coelacanth and other fishes with those of tetrapods. Genetic antecedents for building the autopod proper have been found in the coelacanth genome in the form of regulatory regions (enhancer elements). In particular, the HOX-D cluster of genes plays a key role in patterning digits in the tetrapod limb. It was possible to find regulatory regions upstream of the HOX-D cluster that were shared between tetrapods and coelacanths, but which could not be found in teleost fish. When such a coelacanth regulatory sequence was placed in a transgenic mouse, it was shown to drive reporter expression in an autopod-specific pattern. This strongly suggests that the developmental program driving limb patterning and formation in modern tetrapods was indeed co-opted from a more ancient sarcopterygian developmental program. Is the coelacanth a ‘living fossil’? There has been some push-back concerning the oft-used phrase, living fossil, with regard to the coelacanth. The term was coined by Charles Darwin, and is operationally used to indicate that a species is a surviving representative of an ancient lineage that still retains some key features shared with archaic fossils. Typically such a lineage will have survived one or more mass extinctions. Examples of living fossils often cited include the sharks, ginkgo trees, metasequoia, lampshell brachiopods, horseshoe crabs, and as defined here surely the coelacanth. However, a common misconception is that the phrase implies that evolution has not acted on the organism over these long timescales, something that is clearly shown not to be true for coelacanths based on gross differences in the skeletal morphology of fossilized specimens, especially of forms prior to the Mesozoic. While it is difficult to measure the rate of morphological evolution of extinct coelacanths, analyses of the coelacanth’s protein coding genes have shown, enigmatically, that its relative rate of molecular evolution is slower than that of other fishes and tetrapods. The implications of this relative rate difference remain speculative with respect to the morphological evolution of the coelacanth. What does the future hold for the coelacanth? Many companion genome papers that report surveys of various aspects of coelacanth biology have been published or are soon to be published. The coelacanth genome will continue to play a key role in evolutionary developmental biology studies with respect to the origin of the tetrapods and their unique adaptations. Furthermore, sequence data from additional coelacanth specimens are starting to provide important insights into the genetic diversity in modern populations, insights that will be needed for future conservation efforts given its endangered status. Accurate estimates of coelacanth population sizes are still lacking, but evidence suggests they have extremely restricted ranges. Accidental captures by oilfish fishermen seem to be placing these endangered fish under increasing pressure. The Coelacanth Conservation Council and the South African Coelacanth Conservation and Genome Resource Programme were specifically launched to help research and tackle these urgent issues. Many intact specimens of coelacanths have been captured and preserved within the past few years from the eastern coast of Africa, enabling more in-depth anatomical investigations. And fossil coelacanths are continually being discovered, including a recent form from the Triassic that differs greatly from the modern day Latimeria by virtue of its fork-tailed morphology. We have lots to learn about this iconic species.
The discovery of a living coelacanth specimen in 1938 was remarkable, as this lineage of lobe-finned fish was thought to have become extinct 70 million years ago. The modern coelacanth looks remarkably similar to many of its ancient relatives, and its evolutionary proximity to our own fish ancestors provides a glimpse of the fish that first walked on land. Here we report the genome sequence of the African coelacanth, Latimeria chalumnae. Through a phylogenomic analysis, we conclude that the lungfish, and not the coelacanth, is the closest living relative of tetrapods. Coelacanth protein-coding genes are significantly more slowly evolving than those of tetrapods, unlike other genomic features. Analyses of changes in genes and regulatory elements during the vertebrate adaptation to land highlight genes involved in immunity, nitrogen excretion and the development of fins, tail, ear, eye, brain and olfaction. Functional assays of enhancers involved in the fin-to-limb transition and in the emergence of extra-embryonic tissues show the importance of the coelacanth genome as a blueprint for understanding tetrapod evolution.
Chris T. Amemiya*, Jessica Alföldi*, Alison P. Lee, Shaohua Fan, Hervé Philippe, Iain MacCallum, Ingo Braasch, Tereza Manousaki, Igor Schneider, Nicolas Rohner, Chris Organ, Domitille Chalopin, Jeramiah J. Smith, Mark Robinson, Rosemary A. Dorrington, Marco Gerdol, Bronwen Aken, Maria Assunta Biscotti, Marco Barucca, Denis Baurain, Aaron M. Berlin, Gregory L. Blatch, Francesco Buonocore, Thorsten Burmester, Michael S. Campbell, Adriana Canapa, John P. Cannon, Alan Christoffels, Gianluca De Moro, Adrienne L. Edkins, Lin Fan, Anna Maria Fausto, Nathalie Feiner, Mariko Forconi, Junaid Gamieldien, Sante Gnerre, Andreas Gnirke, Jared V. Goldstone, Wilfried Haerty, Mark E. Hahn, Uljana Hesse, Steve Hoffmann, Jeremy Johnson, Sibel I. Karchner, Shigehiro Kuraku{, Marcia Lara, Joshua Z. Levin, Gary W. Litman, Evan Mauceli{, Tsutomu Miyake, M. Gail Mueller, David R. Nelson, Anne Nitsche, Ettore Olmo, Tatsuya Ota, Alberto Pallavicini, Sumir Panji{, Barbara Picone, Chris P. Ponting, Sonja J. Prohaska, Dariusz Przybylski, Nil Ratan Saha, Vydianathan Ravi, Filipe J. Ribeiro{, Tatjana Sauka-Spengler, Giuseppe Scapigliati, Stephen M. J. Searle, Ted Sharpe, Oleg Simakov, Peter F. Stadler, John J. Stegeman, Kenta Sumiyama, Diana Tabbaa, Hakim Tafer, Jason Turner-Maier, Peter van Heusden, Simon White, Louise Williams, Mark Yandell, Henner Brinkmann, Jean-Nicolas Volff, Clifford J. Tabin, Neil Shubin, Manfred Schartl, David B. Jaffe, John H. Postlethwait, Byrappa Venkatesh, Federica Di Palma, Eric S. Lander, Axel Meyer & Kerstin Lindblad-Toh
ABSTRACTCircular and apparently trans‐spliced RNAs have recently been reported as abundant types of transcripts in mammalian transcriptome data. Both types of non‐colinear RNAs are also abundant in RNA‐seq of different tissue from both the African and the Indonesian coelacanth. We observe more than 8,000 lincRNAs with normal gene structure and several thousands of circularized and trans‐spliced products, showing that such atypical RNAs form a substantial contribution to the transcriptome. Surprisingly, the majority of the circularizing and trans‐connecting splice junctions are unique to atypical forms, that is, are not used in normal isoforms. J. Exp. Zool. (Mol. Dev. Evol.) 322B: 342–351, 2014. © 2013 Wiley Periodicals, Inc.
The identification of transcription factor binding sites (TFBSs) is a non-trivial problem as the existing computational predictors produce a lot of false predictions. Though it is proven that combining these predictions with a meta-classifier, like Support Vector Machines (SVMs), can improve the overall results, this improvement is not as significant as expected. The reason for this is that the predictors are not reliable for the negative examples from non-binding sites in the promoter region. Therefore, using negative examples from different sources during training an SVM can be one of the solutions to this problem. In this study, we used different types of negative examples during training the classifier. These negative examples can be far away from the promoter regions or produced by randomisation or from the intronic region of genes. By using these negative examples during training, we observed their effect in improving predictions of TFBSs in the yeast. We also used a modified cross-validation method for this type of problem. Thus we observed substantial improvement in the classifier performance that could constitute a model for predicting TFBSs. Therefore, the major contribution of the analysis is that for the yeast genome, the position of binding sites could be predicted with high confidence using our technique and the predictions are of much higher quality than the predictions of the original prediction algorithms.
It is known that much of the genetic change underlying morphological evolution takes place in cis-regulatory regions, rather than in the coding regions of genes. Identifying these sites in a genome is a non-trivial problem. Experimental methods for finding binding sites exist with some limitations regarding their applicability, accuracy, availability or cost. On the other hand predicting algorithms perform rather poorly. The aim of this research is to develop and improve computational approaches for the prediction of transcription factor binding sites (TFBSs) by integrating the results of computational algorithms and other sources of complementary biological evidence, with particular emphasis on the use of the Support Vector Machine (SVM). Data from two organisms, yeast and mouse, were used in this study. The initial results were not particularly encouraging, as still giving predictions of low quality. However, when the vectors labelled as non-binding sites in the training set were replaced by randomised training vectors, a significant improvement in performance was observed. This gave substantial improvement over the yeast genome and even greater improvement for the mouse data. In fact the resulting classifier was finding over 80% of the binding sites in the test set and moreover 80% of the predictions were correct.
Identifying transcription factor binding sites computationally is a hard problem as it produces many false predictions. Combining the predictions from existing predictors can improve the overall predictions by using classification methods like Support Vector Machines (SVM). But conventional negative examples (that is, example of nonbinding sites) in this type of problem are highly unreliable. In this study, we have used different types of negative examples. One class of the negative examples has been taken from far away from the promoter regions, where the occurrence of binding sites is very low, and another one has been produced by randomization. Thus we observed the effect of using different negative examples in predicting transcription factor binding sites in mouse. We have also devised a novel cross-validation technique for this type of biological problem.
Finding the location of binding sites in DNA is a difficult problem. Although the location of some binding sites have been experimentally identified, other parts of the genome may or may not contain binding sites. This poses problems with negative data in a trainable classifier. Here we show that using randomized negative data gives a large boost in classifier performance when compared to the original labeled data.
Currently the best algorithms for transcription factor binding site prediction within sequences of regulatory DNA are severely limited in accuracy. In this paper, we integrate 12 original binding site prediction algorithms, and use a ‘window’ of consecutive predictions in order to contextualise the neighbouring results. We combine either random selection or Tomek links under-sampling with SMOTE over-sampling techniques. In addition, we investigate the behaviour of four feature selection filtering methods: bi-normal separation, correlation coefficients, F-Score and a cross entropy based algorithm. Finally, we remove some of the final predicted binding sites on the basis of their biological plausibility. The results show that we can generate a new prediction that significantly improves on the performance of any one of the individual algorithms.
Currently the best algorithms for predicting transcription factor binding sites in DNA sequences are severely limited in accuracy. There is good reason to believe that predictions from different classes of algorithms could be used in conjunction to improve the quality of predictions. In this paper, we apply single layer networks, rules sets, support vector machines and the Adaboost algorithm to predictions from 12 key real valued algorithms. Furthermore, we use a ‘window’ of consecutive results as the input vector in order to contextualise the neighbouring results. We improve the classification result with the aid of under- and over-sampling techniques. We find that support vector machines and the Adaboost algorithm outperform the original individual algorithms and the other classifiers employed in this work. In particular they give a better tradeoff between recall and precision.
The identification of cis-regulatory binding sites in DNA is a difficult problem in computational biology. To obtain a full understanding of the complex machinery embodied in genetic regulatory networks it is necessary to know both the identity of the regulatory transcription factors together with the location of their binding sites in the genome. We show that using an SVM together with data sampling to classify the combination of the results of individual algorithms specialised for the prediction of binding site locations, can produce significant improvements upon the original algorithms. The resulting classifier produces fewer false positive predictions and so reduces the expensive experimental procedure of verifying the predictions.
The identification of cis-regulatory binding sites in DNA is a difficult problem in computational biology. To obtain a full understanding of the complex machinery embodied in genetic regulatory networks it is necessary to know both the identity of the regulatory transcription factors together with the location of their binding sites in the genome. We show that using an SVM together with data sampling, to integrate the results of individual algorithms specialised for the prediction of binding site locations, can produce significant improvements upon the original algorithms. These results make more tractable the expensive experimental procedure of actually verifying the predictions.
Computational prediction of cis-regulatory binding sites is widely acknowledged as a difficult task. There are many different algorithms for searching for binding sites in current use. However, most of them produce a high rate of false positive predictions. Moreover, many algorithmic approaches are inherently constrained with respect to the range of binding sites that they can be expected to reliably predict. We propose to use SVMs to predict binding sites from multiple sources of evidence. We combine random selection under-sampling and the synthetic minority over-sampling technique to deal with the imbalanced nature of the data. In addition, we remove some of the final predicted binding sites on the basis of their biological plausibility. The results show that we can generate a new prediction that significantly improves on the performance of any one of the individual prediction algorithms.
The identification of cis-regulatory binding sites in DNA is a difficult problem in computational biology. To obtain a full understanding of the complex machinery embodied in genetic regulatory networks it is necessary to know both the identity of the regulatory transcription factors and the location of their binding sites in the genome. We show that using an SVM together with data sampling to classify the combination of the results of individual algorithms specialised for the prediction of binding site locations, can produce significant improvements upon the original algorithms. The resulting classifier produces fewer false positive predictions and so reduces the expensive experimental procedure of verifying the predictions.
Original paper can be found at: http://www.biomedcentral.com/bmcsystbiol/archive --DOI : 10.1186/1752-0509-1-S1-P72
Clustering categorical data received more attention since recent years, but several aspects of the existing algorithms, such as the interpretabilities of found clusters, the impact of data selection orders, are not well solved. A novel categorical data clustering algorithm called CLUBMIS is proposed in this paper which can effectively find the interesting clusters. In addition, the clusters can be easily interpreted by the maximal frequent itemsets used in the clustering process. Different from most of the hierarchical clustering algorithm, CLUBMIS clusters datasets based on the summarized information, i.e. maximal frequent itemsets, thus it eliminates the effect of different data selection order.
Currently the best algorithms for transcription factor binding site predictions are severely limited in accuracy. However, a non-linear combination of these algorithms could improve the quality of predictions. A support-vector machine was applied to combine the predictions of 12 key real valued algorithms. The data was divided into a training set and a test set, of which two were constructed: filtered and unfiltered. In addition, a different "window" of consecutive results was used in the input vector in order to contextualize the neighbouring results. Finally, classification results were improved with the aid of under and over sampling techniques. Our major finding is that we can reduce the False-Positive rate significantly. We also found that the bigger the window, the higher the F-score, but the more likely it is to make a false positive prediction, with the best trade-off being a window size of about 7.
The identification of cis-regulatory binding sites in DNA is a difficult problem in computational biology. To obtain a full understanding of the complex machinery embodied in genetic regulatory networks it is necessary to know both the identity of the regulatory transcription factors together with the location of their binding sites in the genome. We show that using an SVM together with data sampling, to integrate the results of individual algorithms specialised for the prediction of binding site locations, can produce significant improvements upon the original algorithms. These results make more tractable the expensive experimental procedure of actually verifying the predictions.
Currently the best algorithms for transcription factor binding site prediction are severely limited in accuracy. There is good reason to believe that predictions from these different classes of algorithms could be used in conjunction to improve the quality of predictions. In previous work, we have applied single layer networks, rules sets and support vector machines on predictions from key real valued algorithms. Furthermore, we used a ‘window’ of consecutive results in the input vector in order to contextualise the neighbouring results. In this paper, we improve the classification result with the aid of a hybrid Adaboost algorithm working on the dataset with windowed inputs. In the proposed algorithm, we first apply weighted majority voting. Those data points which cannot be classified ‘easily’ using weighted majority voting are then classified using the Adaboost algorithm. We find that our method outperforms each of the original individual algorithms and the other classifiers used previously in this work.