enhance accessibility and early exposure of its High-Performance Computing (HPC) resources to students, faculty, and staff, Florida State University's Research Computing Center (RCC) started several educational and outreach initiatives from 2022 to 2024 that we will present in this paper. The key initiatives are the Interdisciplinary Data Humanities Initiative (IDHI) and the HPC Driver's Ed course. IDHI focuses on supporting users in humanities, social sciences, and the arts (HSSA) with the following services: project consultation resources, software installation, and the development of discipline-specific educational materials. IDHI has supported a range of HSSA departments within the last two years utilizing all service components that previously facilitated HPC research in STEM and furthered HPC education. The HPC Driver's Ed course aims to expand asynchronous instructional materials through the development of video-based lectures and examinations to facilitate beginner users' onboarding process. As the HPC Driver's Ed program was initially released in April 2024, we lack statistical data of uptake by our user community.
Recently Graphics Processing Units (GPUs) have been used to speed up very CPU-intensive gravitational microlensing simulations. In this work, we use the Xeon Phi coprocessor to accelerate such simulations and compare its performance on a microlensing code with that of NVIDIA’s GPUs. For the selected set of parameters evaluated in our experiment, we find that the speedup by Intel’s Knights Corner coprocessor is comparable to that by NVIDIA’s Fermi family of GPUs with compute capability 2.0, but less significant than GPUs with higher compute capabilities such as the Kepler. However, the very recently released second generation Xeon Phi, Knights Landing, is about 5.8 times faster than the Knights Corner, and about 2.9 times faster than the Kepler GPU used in our simulations. We conclude that the Xeon Phi is a very promising alternative to GPUs for modern high performance microlensing simulations.
Since its introduction in 2001, MrBayes has grown in popularity as a software package for Bayesian phylogenetic inference using Markov chain Monte Carlo (MCMC) methods. With this note, we announce the release of version 3.2, a major upgrade to the latest official release presented in 2003. The new version provides convergence diagnostics and allows multiple analyses to be run in parallel with convergence progress monitored on the fly. The introduction of new proposals and automatic optimization of tuning parameters has improved convergence for many problems. The new version also sports significantly faster likelihood calculations through streaming single-instruction-multiple-data extensions (SSE) and support of the BEAGLE library, allowing likelihood calculations to be delegated to graphics processing units (GPUs) on compatible hardware. Speedup factors range from around 2 with SSE code to more than 50 with BEAGLE for codon problems. Checkpointing across all models allows long runs to be completed even when an analysis is prematurely terminated. New models include relaxed clocks, dating, model averaging across time-reversible substitution models, and support for hard, negative, and partial (backbone) tree constraints. Inference of species trees from gene trees is supported by full incorporation of the Bayesian estimation of species trees (BEST) algorithms. Marginal model likelihoods for Bayes factor tests can be estimated accurately across the entire model space using the stepping stone method. The new version provides more output options than previously, including samples of ancestral states, site rates, site dN/dS rations, branch rates, and node dates. A wide range of statistics on tree parameters can also be output for visualization in FigTree and compatible software.
What is the probability that Sweden will win next year's world championships in ice hockey? If you're a hockey fan, you probably already have a good idea, but even if you couldn't care less about the game, a quick perusal of the world championship medalists for the last 15 years (Table 7.1) would allow you to make an educated guess. Clearly, Sweden is one of only a small number of teams that compete successfully for the medals. Let's assume that all seven medalists the last 15 years have the same chance of winning, and that the probability of an outsider winning is negligible. Then the odds of Sweden winning would be 1:7 or 0.14. We can also calculate the frequency of Swedish victories in the past. Two gold medals in 15 years would give us the number 2:15 or 0.13, very close to the previous estimate. The exact probability is difficult to determine but most people would probably agree that it is likely to be in the vicinity of these estimates.
The main limiting factor in Bayesian MCMC analysis of phylogeny is typically the efficiency with which topology proposals sample tree space. Here we evaluate the performance of seven different proposal mechanisms, including most of those used in current Bayesian phylogenetics software. We sampled 12 empirical nucleotide data sets-ranging in size from 27 to 71 taxa and from 378 to 2,520 sites-under difficult conditions: short runs, no Metropolis-coupling, and an oversimplified substitution model producing difficult tree spaces (Jukes Cantor with equal site rates). Convergence was assessed by comparison to reference samples obtained from multiple Metropolis-coupled runs. We find that proposals producing topology changes as a side effect of branch length changes (LOCAL and Continuous Change) consistently perform worse than those involving stochastic branch rearrangements (nearest neighbor interchange, subtree pruning and regrafting, tree bisection and reconnection, or subtree swapping). Among the latter, moves that use an extension mechanism to mix local with more distant rearrangements show better overall performance than those involving only local or only random rearrangements. Moves with only local rearrangements tend to mix well but have long burn-in periods, whereas moves with random rearrangements often show the reverse pattern. Combinations of moves tend to perform better than single moves. The time to convergence can be shortened considerably by starting with a good tree, but this comes at the cost of compromising convergence diagnostics based on overdispersed starting points. Our results have important implications for developers of Bayesian MCMC implementations and for the large group of users of Bayesian phylogenetics software.
Aim Oceanic islands represent a special challenge to historical biogeographers because dispersal is typically the dominant process while most existing methods are based on vicariance. Here, we describe a new Bayesian approach to island biogeography that estimates island carrying capacities and dispersal rates based on simple Markov models of biogeographical processes. This is done in the context of simultaneous analysis of phylogenetic and distributional data across groups, accommodating phylogenetic uncertainty and making parameter estimates more robust. We test our models on an empirical data set of published phylogenies of Canary Island organisms to examine overall dispersal rates and correlation of rates with explanatory factors such as geographic proximity and area size.Location Oceanic archipelagos with special reference to the Atlantic Canary Islands.Methods The Canary Islands were divided into three island-groups, corresponding to the main magmatism periods in the formation of the archipelago, while non-Canarian distributions were grouped into a fourth 'mainland-island'. Dispersal between island groups, which were assumed constant through time, was modelled as a homogeneous, time-reversible Markov process, analogous to the standard models of DNA evolution. The stationary state frequencies in these models reflect the relative carrying capacity of the islands, while the exchangeability (rate) parameters reflect the relative dispersal rates between islands. We examined models of increasing complexity: Jukes-Cantor (JC), Equal-in, and General Time Reversible (GTR), with or without the assumption of stepping-stone dispersal. The data consisted of 13 Canarian phylogenies: 954 individuals representing 393 taxonomic (morphological) entities. Each group was allowed to evolve under its own DNA model, with the island-model shared across groups. Posterior distributions on island model parameters were estimated using Markov Chain Monte Carlo (MCMC) sampling, as implemented in MrBayes 4.0, and Bayes Factors were used to compare models.Results The Equal-in step, the GTR, and the GTR step dispersal models showed the best fit to the data. In the Equal-in and GTR models, the largest carrying capacity was estimated for the mainland, followed by the central islands and the western islands, with the eastern islands having the smallest carrying capacity. The relative dispersal rate was highest between the central and eastern islands, and between the central and western islands. The exchange with the mainland was rare in comparison.Main conclusions Our results confirm those of earlier studies suggesting that inter-island dispersal within the Canary Island archipelago has been more important in explaining diversification within lineages than dispersal between the continent and the islands, despite the close proximity to North Africa. The low carrying capacity of the eastern islands, uncorrelated with their size or age, fits well with the idea of a historically depauperate biota in these islands but more sophisticated models are needed to address the possible influence of major recent extinction events. The island models explored here can easily be extended to address other problems in historical biogeography, such as dispersal among areas in continental settings or reticulate area relationships.
An import issue for numerical weather prediction modes (NWP) is the time it takes to produce a valid forecast. One factor, which greatly influences this simulation time is the size of the time step. However, time step size is often limited by the numerical stability of the used advection schemes. Available schemes include semiimplicit Eulerian and semi-Lagrangian schemes. In principal, semi-Lagrangian formulations result in irregular communications on parallel architectures. In this paper we describe automatic code generation for a semi-implicit scheme with a semi-Lagrangian formulation. We describe how code can be generated from a mathematical specification of the advection model, the embedding of the formulations in the CTADEL code generation tool and we show the parallelization of the code. Finally, we show results from preliminary experiments we have conducted with the generated code and the reference code from a production NWP on a number of different architectures.
The use of semi-Lagrangian formulations in numerical weather predication models (NWP) allows for an increase in time step size. Use of this method can increase performance of these models. However, on parallel architectures, communication between processors can become a huge bottleneck, limiting speedup. Furthermore, the communication pattern is dependent on the application's execution. We discus a novel strategy, called Halo on Demand, which dynamically drives the communication between the processors by examining the content of the data at runtime in order to reduce communication costs. With an extensive performance analysis of the execution of the model we show that our strategy can decrease communication time and thus decrease total execution time.
The size of a time step is important for numerical weather prediction models (NWP) since forecasts need to be available within the fraction of time that may considered to be valid. However, time step size is often limited by the numerical stability of the used advection schemes. Available schemes include semi-implicit Eulerian and semi-Lagrangian schemes. In principal, semi-Lagrangian formulations result in irregular communications on parallel architectures. In this paper we describe automatic code generation for a semi-implicit scheme with a semi-Lagrangian formulation. We describe how code can be generated from a mathematical specification of the advection model and we show results from preliminary experiments we have conducted with the generated code and the reference code from a production NWP on a number of different architectures.
In this paper we present an overview of current and on-going research on theCTADEL problem-specific code generator. TheCTADEL system provides an automated means of generating specific high performance scientific codes, optimized for a number of different architectures. We address problems like implicit equations and a SemiLagrangian method for semi-implicit schemes and show some experiments with the generated codes and the handwritten references codes.
Traditional design and implementation of large atmospheric models is a difficult, tedious and error prone task. With the CTADEL project we investigate a new method of code generation, where the designer describes the model in an abstract high-level specification language which is translated into highly optimized Fortran code. This is applied to a convection scheme as used in a numerical weather prediction model (NWP). We address problems like how to generate efficient code for conditional expressions by using polymorphic templates. Finally, we compare the generated code for the convection scheme with the hand-written reference code.
Traditional design and implementation of large atmospheric models is a difficult, tedious and erroneous task. With the Ctadel project we propose a new method of code generation, where the designer describes the model in an abstract high-level specification language which is translated into highly optimized Fortran code. In this paper we show the abilities of this method on a coupled ocean-atmosphere model, in which we have to deal with multi-resolution domains and different timesteps. We, briefly, describe a new concept in compiler design, the use of templates for code generation, to elevate the burden of choosing architecture optimized numerical routines.
In this paper we describe how to extend CTADEL, a Problem Solving Environment, in order to generate code for a turbulence scheme, in our case, within a numerical weather prediction model (NWP). Common for these schemes is the presence of implicit equations. We describe how to generate efficient codes for a particular class of implicit differential equations, that is encountered in the turbulence scheme of the HIRLAM NWP model. This extension to CTADEL enables the code generation for a particular solution method for these implicit equations. This solution method is also used in the original hand-written (reference) code. We address problems like how to recognize the type of equations and how to generate efficient code for a prescribed solution method. Finally, we compare the generated code for the turbulence scheme with the hand-written reference code.
Erven Rohou合作论文数1