Abstract Asian rice is one of the best documented crops in terms of genetic diversity. The domestication process, that probably started 9000 years ago in China, remains difficult to infer since the main vertical signal is blurred by horizontal signals related to gene flow among cultivars and wild relatives. Consequently, a large number of hypotheses on the domestication process of rice have been published. Besides, most of the methods used to infer these scenarios do not model all the known biological phenomena at stake. Here, we present a methodological study based on a rich stochastic model, that incorporates introgression events, incomplete lineage sorting, and mutations that happen over time. The global evolutionary scenario is represented by a phylogenetic network. Furthermore, each locus scenario is modeled according to a locus tree through the Multispecies Network Coalescent . More importantly, for inferring the phylogenetic network, we propose a new hybrid approach combining a phylogenetic network method and a machine learning technique. In particular, our hybrid approach, named S narf , benefits from advantages of a mathematical phylogenetic method, S napp N et , and from the potential of a powerful machine learning classifier, i.e. Approximate Bayesian Computation Random Forest (ABC-RF). These two methods are complementary since S napp N et reconstructs network accurately, whereas ABC-RF is able to handle a large amount of data. The originality is twofold. First, prior distributions required for ABC-RF are calibrated thanks to S napp N et ’s estimates. Secondly, ABC-RF relies on summary statistics inspired by phylogenetic network literature. We show, on simulated data, that the S narf hybrid approach enjoys very good performances. On rice real data, it infers a scenario with a unique domestication (that of Japonica), followed by three reticulation events involving early Japonica. It highlights two introgression events at the origin of Indica and cAus, and one admixture event responsible for the emergence of cBas. Author summary Today, in genomics, there is a real need for methods able to infer phylogenetic networks. A phylogenetic network is a directed graph representing events like hybridization, introgression, and horizontal gene transfer. Understanding these complex biological phenomena, essential for crop adaptation, can help breeders when facing challenges like climate change and population growth. Genome-wide diversity analysis thus requires network methods scaling for large data volumes and incorporating fundamental biological phenomena. In this context, we present a new hybrid approach, S narf , that benefits from the potential of a powerful machine learning classifier, Approximate Bayesian Computation Random Forest, and from advantages of a mathematical phylogenetic method, S napp N et . Consequently, S narf is able to handle large data-sets thanks to machine learning and is also based on a deep mathematical theory. On simulated data, our hybrid approach performs very well. When applied to real rice genomic data, it supports a scenario with a single domestication event, that of Japonica. The analysis further highlights the role of early Japonica in the origin of both Indica and circumAus. Finally, it identifies an ancient admixture event, involving circumAus in the emergence of circumBasmati. Together, these findings confirm the importance of early rice history along the Himalayan region.
AbstractBabesiosis, caused by protozoan parasites of the genusBabesia, is an emerging tick-borne disease of significance for both human and animal health.Babesiaparasites infect erythrocytes of vertebrate hosts where they develop and multiply rapidly to cause the pathological symptoms associated with the disease. The identification of variousBabesiaspecies underscores the ongoing risk of new zoonotic pathogens capable of infecting humans, a concern amplified by anthropogenic activities and environmental shifts impacting the distribution and transmission dynamics of parasites, their vectors, and reservoir hosts. One such species,BabesiaMO1, previously implicated in severe cases of human babesiosis in the midwestern United States, was initially considered closely related toB. divergens, the predominant agent of human babesiosis in Europe. Yet, uncertainties persist regarding whether these pathogens represent distinct variants of the same species or are entirely separate species. We show that although bothB.MO1 andB. divergensshare similar genome sizes, comprising three nuclear chromosomes, one linear mitochondrial chromosome, and one circular apicoplast chromosome, major differences exist in terms of genomic sequence divergence, gene functions, transcription profiles, replication rates and susceptibility to antiparasitic drugs. Furthermore, both pathogens have evolved distinct classes of multigene families, crucial for their pathogenicity and adaptation to specific mammalian hosts. Leveraging genomic information forB.MO1,B. divergens, and other members of the Babesiidae family within Apicomplexa provides valuable insights into the evolution, diversity, and virulence of these parasites. This knowledge serves as a critical tool in preemptively addressing the emergence and rapid transmission of more virulent strains.
Babesiosis, caused by protozoan parasites of the genus Babesia, is an emerging tick-borne disease of significance for both human and animal health. Babesia parasites infect erythrocytes of vertebrate hosts where they develop and multiply rapidly to cause the pathological symptoms associated with the disease. The identification of new Babesia species underscores the ongoing risk of zoonotic pathogens capable of infecting humans, a concern amplified by anthropogenic activities and environmental changes. One such pathogen, Babesia MO1, previously implicated in severe cases of human babesiosis in the United States, was initially considered a subspecies of B. divergens, the predominant agent of human babesiosis in Europe. Here we report comparative multiomics analyses of B. divergens and B. MO1 that offer insight into their biology and evolution. Our analysis shows that despite their highly similar genomic sequences, substantial genetic and genomic divergence occurred throughout their evolution resulting in major differences in gene functions, expression and regulation, replication rates and susceptibility to antiparasitic drugs. Furthermore, both pathogens have evolved distinct classes of multigene families, crucial for their pathogenicity and adaptation to specific mammalian hosts. Leveraging genomic information for B. MO1, B. divergens, and other members of the Babesiidae family within Apicomplexa provides valuable insights into the evolution, diversity, and virulence of these parasites. This knowledge serves as a critical tool in preemptively addressing the emergence and rapid transmission of more virulent strains.
The microservice architecture (MSA) emerged as an evolution of existing architectural styles with the promise of improving software quality by decomposing an app into modules that can be maintained, deployed, and scaled independently. However, the transition from a monolithic to a microservice architecture is fraught with difficulties, especially when it comes to assessing qualitative aspects, as controversial results can arise. In this paper, we present an experience report on the migration of a monolithic web application and use performance, availability and energy efficiency as quality attributes to shed light on such an architectural transition. Horizontal scaling, i.e., distributing the workload across several service instances, is applied and we study its impact.Our main findings are: i) when no app component is replicated, MSA outperforms the monolithic architecture; ii) the monolithic architecture shows performance and availability improvement when replicating the entire app; iii) the replicated MSA version reaches a ceiling when not replicating its routing part (i.e., the API gateway), showing worse response times compared to the replicated monolith; iv) when replicating the API gateway, the MSA version reaches optimal performance with fewer replicates than the monolith; v) when not replicating services, MSA consumes more CPU resources than the monolithic architecture; vi) when scaling up, the MSA version is more efficient than the replicated monolith in terms of memory usage, and it can better exploit CPU resources; vii) when not replicating services, MSA consumes more energy than the monolithic architecture, whereas when scaling up, the MSA version is more efficient than the replicated monolith; MSA version reaches a good balance between CPU and memory usage.
Babesiosis, caused by protozoan parasites of the genus Babesia , is an emerging tick-borne disease of significance for both human and animal health. Babesia parasites infect erythrocytes of vertebrate hosts where they develop and multiply rapidly to cause the pathological symptoms associated with the disease. The identification of various Babesia species underscores the ongoing risk of new zoonotic pathogens capable of infecting humans, a concern amplified by anthropogenic activities and environmental shifts impacting the distribution and transmission dynamics of parasites, their vectors, and reservoir hosts. One such species, Babesia MO1, previously implicated in severe cases of human babesiosis in the midwestern United States, was initially considered closely related to B. divergens , the predominant agent of human babesiosis in Europe. Yet, uncertainties persist regarding whether these pathogens represent distinct variants of the same species or are entirely separate species. We show that although both B. MO1 and B. divergens share similar genome sizes, comprising three nuclear chromosomes, one linear mitochondrial chromosome, and one circular apicoplast chromosome, major differences exist in terms of genomic sequence divergence, gene functions, transcription profiles, replication rates and susceptibility to antiparasitic drugs. Furthermore, both pathogens have evolved distinct classes of multigene families, crucial for their pathogenicity and adaptation to specific mammalian hosts. Leveraging genomic information for B. MO1, B. divergens , and other members of the Babesiidae family within Apicomplexa provides valuable insights into the evolution, diversity, and virulence of these parasites. This knowledge serves as a critical tool in preemptively addressing the emergence and rapid transmission of more virulent strains.
Historically, computing instructors and researchers have developed a wide variety of tools to support teaching and educational research, including exam and code testing suites and data collection solutions. Many are then community or individually maintained. However, these tools often find limited adoption beyond their creators. As a result, it is common for many of the same functionalities to be re-implemented by different instructional groups within the CS Education community. We hypothesize that this is due in part to accessibility, discoverability, and adaptability challenges, among others. Further, instructors often face institutional barriers to deployment, which can include hesitance of institutions to utilize community developed solutions that often lack a centralized authority. This working group will explore what solutions are currently available, what instructors need, and reasons behind the above-mentioned phenomenon. This will be accomplished via a literature review and survey to identify the tools that have been developed by the community; the solutions that are currently available and in use by instructors; what features are needed moving forward for classroom and research use; what support for extensions is needed to support further CS Education research; and what institutional challenges instructors and researchers are currently facing or have faced in the past in developing, deploying or otherwise using community software solutions. Finally, the working group will identify factors that limit adoption of solutions and ways to integrate and improve the accessibility, discoverability, and dissemination of existing community projects, as well as manage and overcome institutional challenges.
We present ShellOnYou, a new Computer Science education tool, and analyze its use with four successive student cohorts. Developed to help instructors manage numerous students, this web application offers auto-graded exercises to acquire practical knowledge of Unix-like operating systems from the command line. For each answer, and almost instantly, students receive a score and detailed feedback. This reactive and iterative process encourages students to resubmit answers and progressively expand their procedural knowledge. The tool can also deliver individualized statements, thereby allowing students to improve their skills by combining personal research and peer learning. As an online tool, ShellOnYou affords students access flexibility, and also easily fits in distance learning programs. We found it particularly useful when teaching students with heterogeneous Unix backgrounds. The tool is available on request. We placed four successive student cohorts in a learning situation involving this tool, and asked them to fill a survey at the end of the learning period. We combine qualitative and quantitative methods to analyze their answers to the survey. We attempt to characterize their acquisition of procedural knowledge and the building of group dynamics. Several dimensions emerge: the benefits of using the tool in a learning situation, the learning process iteration as a catalyst for renewed commitment, the tool's entertaining format and its scoring system as a motivation for regular studying, and the inherent customization of the learning pace.
We present ShellOnYou, a new Computer Science education tool, and analyze its use with four successive student cohorts. Developed to help instructors manage numerous students, this web application offers auto-graded exercises to acquire practical knowledge of Unix-like operating systems from the command line. For each answer, and almost instantly, students receive a score and detailed feedback. This reactive and iterative process encourages students to resubmit answers and progressively expand their procedural knowledge. The tool can also deliver individualized statements, thereby allowing students to improve their skills by combining personal research and peer learning. As an online tool, ShellOnYou affords students access flexibility, and also easily fits in distance learning programs. We found it particularly useful when teaching students with heterogeneous Unix backgrounds. The tool is available on request. We placed four successive student cohorts in a learning situation involving this tool, and asked them to fill a survey at the end of the learning period. We combine qualitative and quantitative methods to analyze their answers to the survey. We attempt to characterize their acquisition of procedural knowledge and the building of group dynamics. Several dimensions emerge: the benefits of using the tool in a learning situation, the learning process iteration as a catalyst for renewed commitment, the tool's entertaining format and its scoring system as a motivation for regular studying, and the inherent customization of the learning pace.
Historically, computing instructors and researchers have developed a wide variety of tools to support teaching and educational research, including exam and code testing suites and data collection solutions. However, these tools often find limited adoption beyond their creators. As a result, it is common for many of the same functionalities to be re-implemented by different instructional groups within the Computing Education community. We hypothesise that this is due in part to discoverability, availability, and adaptability challenges. Further, instructors often face institutional barriers to deployment, which can include hesitance of institutions to rely on community developed solutions that often lack a centralised authority and may be community or individually maintained. To this end, our working group explored what solutions are currently available, what instructors needed, and the reasons behind the above-mentioned phenomenon. To do so, we reviewed existing literature and surveyed the community to identify the tools that have been developed by the community; the solutions that are currently available and in use by instructors; what features are needed moving forward for classroom and research use; what support for extensions is needed to support further Computing Education research; and what institutional challenges instructors and researchers are currently facing or have faced in using community software solutions. Finally, the working group identified factors that limited adoption of solutions. This work proposes ways to integrate and improve the availability, discoverability, and dissemination of existing community projects, as well as ways to manage and overcome institutional challenges.
For various species, high quality sequences and complete genomes are nowadays available for many individuals. This makes data analysis challenging, as methods need not only to be accurate, but also time efficient given the tremendous amount of data to process. In this article, we introduce an efficient method to infer the evolutionary history of individuals under the multispecies coalescent model in networks (MSNC). Phylogenetic networks are an extension of phylogenetic trees that can contain reticulate nodes, which allow to model complex biological events such as horizontal gene transfer, hybridization and introgression. We present a novel way to compute the likelihood of biallelic markers sampled along genomes whose evolution involved such events. This likelihood computation is at the heart of a Bayesian network inference method called SnappNet, as it extends the Snapp method inferring evolutionary trees under the multispecies coalescent model, to networks. SnappNet is available as a package of the well-known beast 2 software. Recently, the MCMC_BiMarkers method, implemented in PhyloNet, also extended Snapp to networks. Both methods take biallelic markers as input, rely on the same model of evolution and sample networks in a Bayesian framework, though using different methods for computing priors. However, SnappNet relies on algorithms that are exponentially more time-efficient on non-trivial networks. Using simulations, we compare performances of SnappNet and MCMC_BiMarkers. We show that both methods enjoy similar abilities to recover simple networks, but SnappNet is more accurate than MCMC_BiMarkers on more complex network scenarios. Also, on complex networks, SnappNet is found to be extremely faster than MCMC_BiMarkers in terms of time required for the likelihood computation. We finally illustrate SnappNet performances on a rice data set. SnappNet infers a scenario that is consistent with previous results and provides additional understanding of rice evolution.
Phylogenetic networks are rooted directed acyclic graphs used to depict the evolution of a set of species in the presence of reticulate events. Reconstructing these networks from molecular data is challenging and current algorithms fail to scale up to genome-wide data. In this paper, we introduce a new width measure intended to help design faster parameterized algorithms for this task. We study its relation with other width measures and problems in graph theory and finally prove that deciding it is NP-complete, even for very restricted classes of networks.
This paper focuses on veto supertree methods, ie methods that aim at producing a conservative synthesis of the relationships agreed upon by all source trees. We propose desirable properties that a supertree should satisfy in this framework, namely the non-contradiction property (PC) and the induction property (PI). The former requires that the supertree does not contain relationships that contradict one or a combination of the source topologies, while the latter requires that all topological information contained in the supertree is present in a source tree or collectively induced by several source trees. We provide simple examples to illustrate their relevance and that allow a comparison with previously advocated properties. We show that these properties can be checked in polynomial time for any given rooted supertree. Moreover, we introduce the PhySIC method (phylogenetic signal with induction and non-contradiction). For k input trees spanning a set of n taxa, this method produces a supertree that satisfies the above-mentioned properties in O(kn3 + n4) polynomial computing time. The polytomies of the produced supertree are also tagged by labels indicating areas of conflict as well as areas with insufficient overlap. As a whole, PhySIC enables the user to quickly summarize consensual information of a set of trees and localize groups of taxa for which the data requires consolidation. Lastly, we illustrate the behaviour of PhySIC on primate datasets of various sizes, and propose a supertree covering 95% of all primate extant genera. The PhySIC algorithm is available at http://atgc.lirmm.fr/cgi-bin/PhySIC/physic.cgi.
Comment un objet de divertissement, le jeu, et tout particulierement le jeu video, peut-il devenir un support d’apprentissage ? Le jeu qui abandonne la frivolite en devenant un jeu serieux, contradiction pour certains, est-il encore un jeu ? Gamifier, est-ce vraiment offrir une experience ludique ? La rencontre permettra d’explorer les enjeux de ces mouvements, et ce que la gamification peut apporter au-dela du marketing. Cette rencontre est disponible dans notre WebTV / WebRadio.
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Maximum Compatible Tree Vincent Berry
Motivation: Developing methods for computing phylogenetic networks from biological data is an important problem posed by molecular evolution and much work is currently being undertaken in this area. Although promising approaches exist, there are no tools available that biologists could easily and routinely use to compute rooted phylogenetic networks on real datasets containing tens or hundreds of taxa. Biologists are interested in clades, i.e. groups of monophyletic taxa, and these are usually represented by clusters in a rooted phylogenetic tree. The problem of computing an optimal rooted phylogenetic network from a set of clusters, is hard, in general. Indeed, even the problem of just determining whether a given network contains a given cluster is hard. Hence, some researchers have focused on topologically restricted classes of networks, such as galled trees and level-k networks, that are more tractable, but have the practical draw-back that a given set of clusters will usually not possess such a representation. Results: In this article, we argue that galled networks (a generalization of galled trees) provide a good trade-off between level of generality and tractability. Any set of clusters can be represented by some galled network and the question whether a cluster is contained in such a network is easy to solve. Although the computation of an optimal galled network involves successively solving instances of two different NP-complete problems, in practice our algorithm solves this problem exactly on large datasets containing hundreds of taxa and many reticulations in seconds, as illustrated by a dataset containing 279 prokaryotes. Availability: We provide a fast, robust and easy-to-use implementation of this work in version 2.0 of our tree-handling software Dendroscope, freely available from http://www.dendroscope.org. Contact: huson@informatik.uni-tuebingen.de
Motivation:A reconciliation is an annotation of the nodes of a gene tree with evolutionary events-for example, speciation, gene duplication, transfer, loss, etc.-along with a mapping onto a species tree. Many algorithms and software produce or use reconciliations but often using different reconciliation formats, regarding the type of events considered or whether the species tree is dated or not. This complicates the comparison and communication between different programs.Results:Here, we gather a consortium of software developers in gene tree species tree reconciliation to propose and endorse a format that aims to promote an integrative-albeit flexible-specification of phylogenetic reconciliations. This format, named recPhyloXML, is accompanied by several tools such as a reconciled tree visualizer and conversion utilities.Availability and implementation:http://phylariane.univ-lyon1.fr/recphyloxml/.
CompPhy2 is a real time collaborative web platform for visualization, comparison and annotation of phylogenetic trees with a responsive user friendly interface. It is the only phylogenetic web platform allowing several users to work synchronously or asynchronously on a same dataset. CompPhy2 brings new features such as side by side comparison, tree annotation, custom display, supertree inference,
Edible bananas result from interspecific hybridization between Musa acuminata and Musa balbisiana, as well as among subspecies in M. acuminata. Four particular M. acuminata subspecies have been proposed as the main contributors of edible bananas, all of which radiated in a short period of time in southeastern Asia. Clarifying the evolution of these lineages at a whole-genome scale is therefore an important step toward understanding the domestication and diversification of this crop. This study reports the de novo genome assembly and gene annotation of a representative genotype from three different subspecies of M. acuminata. These data are combined with the previously published genome of the fourth subspecies to investigate phylogenetic relationships. Analyses of shared and unique gene families reveal that the four subspecies are quite homogenous, with a core genome representing at least 50% of all genes and very few M. acuminata species-specific gene families. Multiple alignments indicate high sequence identity between homologous single copy-genes, supporting the close relationships of these lineages. Interestingly, phylogenomic analyses demonstrate high levels of gene tree discordance, due to both incomplete lineage sorting and introgression. This pattern suggests rapidradiation within Musa acuminata subspecies that occurred after the divergence with M. balbisiana. Introgression between M. a. ssp. malaccensis and M. a. ssp. burmannica was detected across the genome, though multiple approaches to resolve the subspecies tree converged on the same topology. To support evolutionary and functional analyses, we introduce the PanMusa database, which enables researchers to exploration of individual gene families and trees.
Cospeciation studies aim at investigating whether hosts and symbionts speciate simultaneously or whether the associations diversify through host shifts. This problem is often tackled through reconciliation analyses that map the symbiont phylogeny onto the host phylogeny by mixing different types of diversification events. These reconciliations can be difficult to interpret and are not always biologically realistic. Researchers have underlined that the biogeographic histories of both hosts and symbionts influence the probability of cospeciation and host switches, but up to now no reconciliation software integrates geographic data. We present a new functionality in the Mowgli software that bridges this gap. The user can provide geographic information on both the host and symbiont extant and ancestral taxa. Constraints in the reconciliation algorithm have been implemented to generate biologically realistic codiversification scenarios. We apply our method to the fig/fig wasp association and infer diversification scenarios that differ from reconciliations ignoring geographic information. In addition, we updated the reconciliation viewer SylvX to visualize ancestral character states on the phylogenetic trees and highlight parts of reconciliations that are geographically inconsistent when not accounting for geographic constraints. We suggest that the comparison of reconciliations obtained with and without such constraints can help solving ambiguities in the biogeographic histories of the partners. With the development of robust methods in historical biogeography, and the advent of next-generation sequencing that leads to better-resolved trees, a geography-aware reconciliation method represents a substantial advance that is likely to be useful to researchers studying the evolution of biotic interactions and biogeography.
CompPhy2 is a real time collaborative web platform for visualization, comparison and annotation of phylogenetic trees with a responsive user friendly interface. It is the only phylogenetic web platform allowing several users to work synchronously or asynchronously on a same dataset. CompPhy2 brings new features such as side by side comparison, tree annotation, custom display, supertree inference,
Olivier Gascuel合作论文数Methodes et Algorithmes pour la Bioinformatique
LIRMM4