The aim of this study was to investigate the maternal genealogical pattern of chicken breeds sampled in Europe. Sequence polymorphisms of 1256 chickens of the hypervariable region (D-loop) of mitochondrial DNA (mtDNA) were used. Median-joining networks were constructed to establish evolutionary relationships among mtDNA haplotypes of chickens, which included a wide range of breeds with different origin and history. Chicken breeds which have had their roots in Europe for more than 3000 years were categorized by their founding regions, encompassing Mediterranean type, East European type and Northwest European type. Breeds which were introduced to Europe from Asia since the mid-19th century were classified as Asian type, and breeds based on crossbreeding between Asian breeds and European breeds were classified as Intermediate type. The last group, Game birds, included fighting birds from Asia. The classification of mtDNA haplotypes was based on Liu et al.'s (2006) nomenclature. Haplogroup E was the predominant clade among the European chicken breeds. The results showed, on average, the highest number of haplotypes, highest haplotype diversity, and highest nucleotide diversity for Asian type breeds, followed by Intermediate type chickens. East European and Northwest European breeds had lower haplotype and nucleotide diversity compared to Mediterranean, Intermediate, Game and Asian type breeds. Results of our study support earlier findings that chicken breeds sampled in Europe have their roots in the Indian subcontinent and East Asia. This is consistent with historical and archaeological evidence of chicken migration routes to Europe.
The study aimed to evaluate the genetic diversity of Tanzanian chicken populations through phylogenetic relationship, and to trace the history of Tanzanian indigenous chickens. Five ecotypes of Tanzanian local chickens (Ching'wekwe, Kuchi, Morogoro-medium, Pemba and Unguja) from eight regions were studied. Diversity was assessed based on morphological measurements and 29 microsatellite markers recommended by ISAG/FAO advisory group on animal genetic diversity. A principal component analysis (PCA) of morphological measures distinguished individuals most by body sizes and body weight. Morogoro Medium, Pemba and Unguja were grouped together, while Ching'wekwe stood out because of their disproportionate short shanks and ulna bones. Kuchi formed an independent group owing to their comparably long body sizes. Microsatellite analysis revealed three clusters of Tanzanian chicken populations. These clusters encompassed i) Morogoro-medium and Ching'wekwe from Eastern and Central Zones ii) Unguja and Pemba from Zanzibar Islands and iii) Kuchi from Lake Zone regions, which formed an independent cluster. Sequence polymorphism of D-loop region was analysed to disclose the likely maternal origin of Tanzanian chickens. According to reference mtDNA haplotypes, the Tanzanian chickens that were sampled encompass two haplogroups of different genealogical origin. From haplotype network analysis, Tanzanian chickens probably originated on the Indian subcontinent and in Southeast Asia. The majority of Kuchi chickens clustered in a single haplogroup, which was previously found in Shamo game birds sampled from Shikoku Island of Japan in the Kochi Prefecture. Analysis of phenotypic and molecular data, as well as the linguistic similarity of the breed names, suggests a recent introduction of the Kuchi breed to Tanzania.
Often, it is not difficult to train a network to generate a number- but what confidence can we have in that number? We would like to compute a probability density whose mean is the target value, and the variance is the confidence.
Copyright holders have been investigating technological solutions to prevent distribution of copyrighted materials in peer-to-peer file sharing networks. A particularly popular technique consists in "poisoning" a specific item (movie, song, or software title) by injecting a massive number of decoys into the peer-to-peer network, to reduce the availability of the targeted item. In addition to poisoning, pollution, that is, the accidental injection of unusable copies of files in the network, also decreases content availability. In this paper, we attempt to provide a first step toward understanding the differences between pollution and poisoning, and their respective impact on content availability in peer-to-peer file sharing networks. To that effect, we conduct a measurement study of content availability in the four most popular peer-to-peer file sharing networks, in the absence of poisoning, and then simulate different poisoning strategies on the measured data to evaluate their potential impact. We exhibit a strong correlation between content availability and topological properties of the underlying peer-to-peer network, and show that the injection of a small number of decoys can seriously impact the users' perception of content availability.
Anomaly detection approaches have become critically important to enhance decision-making systems, especially regarding the process of risk reduction in the economic performance of an organisation and the consumer costs. Previous studies on anomaly detection have examined mainly abnormalities that translate into fraud, such as fraudulent credit card transactions or fraud in insurance systems. However, anomalies represent irregularities in system patterns data, which may arise from deviations, adulterations or inconsistencies. Further, its study encompasses not only fraud, but also any behavioural abnormalities that signal risks. This paper proposes a literature review of methods and techniques to detect anomalies on diverse financial systems using a five-step technique. In our proposed method, we created a classification framework using codes to systematize the main techniques and knowledge on the subject, in addition to identifying research opportunities. Furthermore, the statistical results show several research gaps, among which three main ones should be explored for developing this area: a common database, tests with different dimensional sizes of data and indicators of the detection models' effectiveness. Therefore, the proposed framework is pertinent to comprehending an existing scientific knowledge base and signals important gaps for a research agenda considering the topic of anomalies in financial systems.
Most approaches in forecasting merely try to predict the next value of the time series. In contrast, this paper presents a framework to predict the full probability distribution. It is expressed as a mixture model: the dynamics of the individual states is modeled with so-called (potentially non-linear neural networks), and the dynamics between the states is modeled using a Markov approach. The full density predictions are obtained by a weighted superposition of the individual densities of each expert. This model class is called hidden Markov experts.Results are presented for daily S&P500 data. While the predictive accuracy of the mean does not improve over simpler models, evaluating the prediction of the full density shows a clear out-of-sample improvement both over a simple GARCH(1,1) model (which assumes Gaussian distributed returns) and over a gated experts model (which expresses the weighting for each state non-recursively as a function of external inputs). Several interpretations are given: the blending of supervised and unsupervised learning, the discovery of states, the combination of forecasts, the specialization of experts, the removal of outliers, and the persistence of volatility.
This paper presents 'hidden Markov experts', a framework for predicting conditional probability distributions of future values of a time series. On daily S&P500 data, the out-of- sample performance is compared to several baselines including GARCH and 'gated experts'. The evaluation of the full density shows improvement over all competitors. Since the performance for point-predictions is comparable to the other methods, the main advantage of hidden Markov experts is their use for conditional density forecasting. Copyright © 2000 John Wiley & Sons, Ltd.
With the recent dramatic increase in electronic access to documents, text categorization—the task of assigning topics to a given document—has moved to the center of the information sciences and knowledge management. This article uses the structure that is present in the semantic space of topics in order to improve performance in text categorization: according to their meaning, topics can be grouped together into “meta-topics”, e.g., gold, silver, and copper are all metals. The proposed architecture matches the hierarchical structure of the topic space, as opposed to a flat model that ignores the structure. It accommodates both single and multiple topic assignments for each document. Its probabilistic interpretation allows its predictions to be combined in a principled way with information from other sources. The first level of the architecture predicts the probabilities of the meta-topic groups. This allows the individual models for each topic on the second level to focus on finer discriminations within the group. Evaluating the performance of a two-level implementation on the Reuters-22173 testbed of newswire articles shows the most significant improvement for rare classes.
Addressing the problem of non-normal portfolio returns, we introduce a novel approach for estimating the distribution of portfolio returns considering higher order mutual information. It allows us to extend the standard variance-covariance framework and efficiently re-compute measures of market risk such as the standard Value-at-Risk or any other probability density based measure. The approach combines two clean and transparent methodologies-independent component analysis and finite Gaussian mixture distributions-and is formulated algorithmically in three steps.
This article introduces a new tool for exploratory data analysis and data mining called Scale-Sensitive Gated Experts (SSGE) which can partition a complex nonlinear regression surface into a set of simpler surfaces (which we call features). The set of simpler surfaces has the property that each element of the set can be efficiently modeled by a single feedforward neural network. The degree to which the regression surface is partitioned is controlled by an external scale parameter. The SSGE consists of a nonlinear gating network and several competing nonlinear experts. Although SSGE is similar to the mixture of experts model of Jacobs et al. [10] the mixture of experts model gives only one partitioning of the input-output space, and thus a single set of features, whereas the SSGE gives the user the capability to discover families of features. One obtains a new member of the family of features for each setting of the scale parameter. In this paper, we derive the Scale-Sensitive Gated Experts and demonstrate its performance on a time series segmentation problem. The main results are: 1) the scale parameter controls the granularity of the features of the regression surface, 2) similar features are modeled by the same expert and different kinds of features are modeled by different experts, and 3) for the time series problem, the SSGE finds different regimes of behavior, each with a specific and interesting interpretation.
This volume selects the best contributions from the Fourth International Conference on Neural Networks in the Capital Markets (NNCM). The conference brought together academics from several disciplines with strategists and decision makers from the financial industries.The various chapters present and compare new techniques from many areas including data mining, information systems, machine learning, and statistical artificial intelligence. The volume focuses on evaluating their usefulness for problems in computational finance and financial engineering.Applications — risk management; asset allocation; dynamic trading and hedging; forecasting; trading cost control. Markets — equity; foreign exchange; bond; commodity; derivatives; Approaches — data mining; statistical AI; machine learning; Monte Carlo simulation; bootstrapping; genetic algorithms; nonparametric methods; fuzzy logic.The chapters emphasizes in-depth and comparative evaluation with established approaches.
This paper assumes identity covariance matrices. Furthermore, eachtrade is fully assigned to a single cluster. We compare this approach to diagonal and to full covariancestructure with probabilistic assignments.Trade profit was held back in the clustering process. It turns out that the clusters differ significantly intheir profit and risk characteristics. Using conditional distributions, we summarize features of profitabletrading styles and contrast them with losing strategies. We find...
This study uncovers trading styles in the transaction records of US Treasury bond futures. We use statistical clustering techniques to group together trades that are similar. Trade profit was held back in the clustering process. Results show that clusters differ significantly in their profit and risk characteristics. Some clusters uncover "technical" trading rules. Using the information about the individual accounts, we describe the assignments of accounts to clusters by entropy, and model the transitions of a given account through clusters by a first order Markov model.
The paper discusses the application of a signal processing technique known as independent component analysis (ICA), also called blind source separation, to multivariate financial time series. The key idea of ICA is to linearly map observed multivariate time series (such as a portfolio of stocks) into a new space of components that are statistically independent. The authors apply ICA to daily returns of the 28 largest Japanese stocks and compare the ICA results to principal component analysis. Their results indicate that the estimated ICs fall into two categories, (i) infrequent but large shocks (responsible for the major changes in the stock prices), and (ii) frequent but rather small fluctuations (contributing little to the overall level of the stocks). They show that the overall stock price can be reconstructed surprisingly well by thresholding the weighted ICs and using, on average, only one such shock per quarter. In contrast, when using shocks derived from principal components instead of independent components, the reconstructed price does not resemble the original one. The technique of ICA is shown to be a potentially powerful method to analyze and understand driving mechanisms in financial time series.
In time series problems, noise can be divided into two categories: dynamic noise which drives the process, and observational noise which is added in the measurement process, but does not influence future values of the system.In this framework, empirical volatilities (the squared relative returns of prices) exhibit a significant amount of observational noise. To model and predict their time evolution adequately, we estimate state space models that explicitly include observational noise. We obtain relaxation times for shocks in the logarithm of volatility. We compare these results with ordinary autoregres-sive models and find that autoregressive models underestimate the relaxation times by about two orders of magnitude due to their ignoring the distinction between observational and dynamic noise. This new interpretation of the dynamics of volatility in terms of relaxators in a state space model carries over to stochastic volatility models and to GARCH models, and is useful for several problems in finance, including risk management and the pricing of derivative securities.