The effective population size (Ne) is a major factor determining allele frequency changes in natural and experimental populations. Temporal methods provide a powerful and simple approach to estimate short-term Ne. They use allele frequency shifts between temporal samples to calculate the standardized variance, which is directly related to Ne. Here we focus on experimental evolution studies that often rely on repeated sequencing of samples in pools (Pool-seq). Pool-seq is cost-effective and often outperforms individual-based sequencing in estimating allele frequencies, but it is associated with atypical sampling properties: Additional to sampling individuals, sequencing DNA in pools leads to a second round of sampling, which increases the variance of allele frequency estimates. We propose a new estimator of Ne, which relies on allele frequency changes in temporal data and corrects for the variance in both sampling steps. In simulations, we obtain accurate Ne estimates, as long as the drift variance is not too small compared to the sampling and sequencing variance. In addition to genome-wide Ne estimates, we extend our method using a recursive partitioning approach to estimate Ne locally along the chromosome. Since the type I error is controlled, our method permits the identification of genomic regions that differ significantly in their Ne estimates. We present an application to Pool-seq data from experimental evolution with Drosophila and provide recommendations for whole-genome data. The estimator is computationally efficient and available as an R package at https://github.com/ThomasTaus/Nest.
Motivation: Recent advances in high-throughput sequencing (HTS) have made it possible to monitor genomes in great detail. New experiments not only use HTS to measure genomic features at one time point but also monitor them changing over time with the aim of identifying significant changes in their abundance. In population genetics, for example, allele frequencies are monitored over time to detect significant frequency changes that indicate selection pressures. Previous attempts at analyzing data from HTS experiments have been limited as they could not simultaneously include data at intermediate time points, replicate experiments and sources of uncertainty specific to HTS such as sequencing depth. Results: We present the beta-binomial Gaussian process model for ranking features with significant non-random variation in abundance over time. The features are assumed to represent proportions, such as proportion of an alternative allele in a population. We use the beta-binomial model to capture the uncertainty arising from finite sequencing depth and combine it with a Gaussian process model over the time series. In simulations that mimic the features of experimental evolution data, the proposed method clearly outperforms classical testing in average precision of finding selected alleles. We also present simulations exploring different experimental design choices and results on real data from Drosophila experimental evolution experiment in temperature adaptation. Availability and implementation: R software implementing the test is available at https://github.com/handetopa/BBGP . Contact: hande.topa@aalto.fi , agnes.jonas@vetmeduni.ac.at , carolin.kosiol@vetmeduni.ac.at , antti.honkela@hiit.fi Supplementary information: Supplementary data are available at Bioinformatics online.
Generating ensembles from multiple individual classifiers is a popular approach to raise the accuracy of the decision. As a rule for decision making, majority voting is a usually applied model. In this paper, we generalize classical majority voting by incorporating probability terms pn,k to constrain the basic framework. These terms control whether a correct or false decision is made if k correct votes are present among the total number of n. This generalization is motivated by object detection problems, where the members of the ensemble are image processing algorithms giving their votes as pixels in the image domain. In this scenario, the terms pn,k can be specialized by a geometric constraint. Namely, the votes should fall inside a region matching the size and shape of the object to vote together. We give several theoretical results in this new model for both dependent and independent classifiers, whose individual accuracies may also differ. As a real world example, we present our ensemble-based system developed for the detection of the optic disc in retinal images. For this problem, experimental results are shown to demonstrate the characterization capability of this system. We also investigate how the generalized model can help us to improve an ensemble with extending it by adding a new algorithm.
In this paper we propose a method using a generalization of the weighted majority voting scheme to locate the optic disc (OD) in retinal images automatically. The location with the maximal sum of the weights of OD center candidates falling into a disc of radius predefined in the clinical protocol is chosen for optic disc. We have worked out a weighted voting scheme, where besides the weights, an additional (e.g. geometrical) condition has to be taken into account in making the final decision. We can achieve better overall performance with this generalized weighted voting system than with the weighted majority voting and each individual algorithm.
In this paper we propose a new voting scheme for generalizing the classical majority voting system. In contrast to the classical voting method we can make good decision if the number of classifiers assigning the correct class label is less than the half of the overall number of classifiers. This new method can be applied to such problems when the decision is not only logical but is also required to satisfy a pre-defined (e. g. geometrical) condition. In addition, we will show the results corresponding to the concept of pattern of success and failure using the existence theorem for the combinatorial (0,1)-matrices.
In this paper we propose a method for locating the optic disc (OD) in retinal images automatically using a generalization of majority voting scheme. Applying more different optic disc detectors for voting we can achieve better performance for the automatic detection system than for each individual algorithm. The location with maximum number of OD center candidates falling within a radius predefined clinically can be used to localize the OD center. In contrast to the classical voting system we can make good decision if the number of algorithms detecting the optic disc correctly is less than the half of the overall number of algorithms.