We define a general notion of entropy in elementary, algebraic terms. Based on that, weak forms of a scalar product and a distance measure are derived. We give basic properties of these quantities, generalize the Cauchy-Schwarz inequality, and relate our approach to the theory of scoring rules. Many supporting examples illustrate our approach and give new perspectives on established notions, such as the likelihood, the Kullback-Leibler divergence, the uncorrelatedness of random variables, the scalar product itself, the Tichonov regularization and the mutual information.
All characterizations of the Shannon entropy include the so-called chain rule, a formula on a hierarchically structured probability distribution, which is based on at least two elementary distributions. We show that the chain rule can be split into two natural components, the well-known additivity of the entropy in case of cross-products and a variant of the chain rule that involves only a single elementary distribution. The latter is given as a proportionality relation and, hence, allows a vague interpretation as self-similarity, hence intrinsic property of the Shannon entropy. Analogous characterizations are given for the Rényi entropy and its limits, the min-entropy and the Hartley entropy.
We provide sufficient conditions of Pólya type which guarantee the positive definiteness of an isotropic 2 × 2-matrix-valued function in R and R 3. Several isotropic bivariate covariance models have been proposed in literature, where all components of the covariance matrix are of the same parametric family, such as the bivariate Matérn model. Based on the Pólya type conditions, we introduce two novel bivariate parametric covariance models of this class, the powered exponential (or stable) covariance model and the generalized Cauchy covariance model. Both models allow for flexible smoothness, variance, scale, and cross-correlation parameters. The smoothness parameters are in ( 0 , 1 ]. Additionally, the bivariate generalized Cauchy model allows for distinct long range parameters. We also show that the univariate spherical model can be generalized to the bivariate case within the above class only in a trivial way.
Stochastic models of point patterns in space and time are widely used to issue forecasts or assess risk, and often they affect societally relevant decisions. We adapt the concept of consistent scoring functions and proper scoring rules, which are statistically principled tools for the comparative evaluation of predictive performance, to the point process setting, and place both new and existing methodology in this framework. With reference to earthquake likelihood model testing, we demonstrate that extant techniques apply in much broader contexts than previously thought. In particular, the Poisson log-likelihood can be used for theoretically principled comparative forecast evaluation in terms of cell expectations. We illustrate the approach in a simulation study and in a comparative evaluation of operational earthquake forecasts for Italy.
In the last decade, a number of methods have been suggested to deal with large amounts of genetic data in genomic predictions. Yet, steadily growing population sizes and the suboptimal use of computational resources are pushing the practical application of these approaches to their limits. As an extension to the C/CUDA library miraculix , we have developed tailored solutions for the computation of genotype matrix multiplications which is a critical bottleneck in the empirical evaluation of many statistical models. We demonstrate the benefits of our solutions at the example of single-step models which make repeated use of this kind of multiplication. Targeting modern Nvidia ® GPUs as well as a broad range of CPU architectures, our implementation significantly reduces the time required for the estimation of breeding values in large population sizes. miraculix is released under the Apache 2.0 license and is freely available at https://github.com/alexfreudenberg/miraculix .
Body core temperature (BCT) is an important characteristic for the vitality of pigs. Suboptimal BCT might indicate or lead to increased stress or diseases. Thermal imaging technologies offer the opportunity to determine BCT in a non-invasive, stress-free way, potentially reducing the manual effort. The current approaches often use multiple close-up images of different parts of the body to estimate the rectal temperature, which is laborious under practical farming conditions. Additionally, images need to be manually annotated for the regions of interest inside the manufacturer's software. Our approach only needs a single (top view) thermal image of a piglet to automatically estimate the BCT. We first trained a convolutional neural network for the detection of the relevant areas, followed by a background segmentation using the Otsu algorithm to generate precise mean, median, and max temperatures of each detected area. The best fit of our method had an R-2 = 0.774. The standardized setup consists of a "FLIROnePro" attached to an Android tablet. To sum up, this approach could be an appropriate tool for animal monitoring under commercial and research farming conditions.
Pseudo-variograms appear naturally in the context of multivariate Brown-Resnick processes, and are a useful tool for analysis and prediction of multivariate random fields. We give a necessary and sufficient criterion for a matrix-valued function to be a pseudo-variogram, and further provide a Schoenberg-type result connecting pseudo-variograms and multivariate correlation functions. By means of these characterizations, we provide extensions of the popular univariate space-time covariance model of Gneiting to the multivariate case.
So far, the pseudo cross-variogram is primarily used as a tool for the structural analysis of multivariate random fields. Mainly applying recent theoretical results on the pseudo cross-variogram, we use it as a cornerstone in the construction of valid covariance models for multivariate random fields. In particular, we extend known univariate constructions to the multivariate case, and generalize existing multivariate models. Furthermore, we provide a general construction principle for conditionally negative definite matrix-valued kernels, which we use to reinterpret previous modeling proposals.
Understanding how attentional resources are deployed in visual processing is a fundamental and highly debated topic. As an alternative to theoretical models of visual search that propose sequences of separate serial or parallel stages of processing, we suggest a queueing processing structure that entails a serial transition between parallel processing stages. We develop a continuous-time queueing model for standard visual search tasks to formalize and implement this notion. Specified as a finite-time, single-line, multiserver queueing system, the model accounts for both accuracy and response time (RT) data in visual search on a distributional level. It assumes two stages of processing. Visual stimuli first go through a massively parallel preattentive stage of feature encoding. They wait if necessary and then enter a limited-capacity attentive stage serially where multiple processing channels ("servers") integrate features of several stimuli in parallel. A core feature of our model is the serial transition from the unlimited-capacity preattentive processing stage to the limited-capacity attentive processing stage. It enables asynchronous attentive processing of multiple stimuli in parallel and is more efficient than a simple chain of two successive, strictly parallel processing stages. The model accounts for response errors by means of two underlying mechanisms, namely, imperfect processing of the servers and, in addition, incomplete search adopted by the observer to maximize search efficiency under an accuracy constraint. For statistical inference, we develop a Monte-Carlo-based parameter estimation procedure, using maximum likelihood (ML) estimation for accuracy-related parameters and minimum distance (MD) estimation for RT-related parameters. We fit the model to two large empirical data sets from two types of visual search tasks. The model captures the accuracy rates almost perfectly and the observed RT distributions quite well, indicating a high explanatory power. The number of independent parallel processing channels that explain both data sets best was five. We also perform a Monte-Carlo model uncertainty analysis and show that the model with the correct number of parallel channels is selected for more than 90% of the simulated samples.
Article STADS – Wie eine Studierendeninitiative der Universität Mannheim eine Brücke zwischen Theorie und Praxis schlägt was published on September 30, 2022 in the journal Mitteilungen der Deutschen Mathematiker-Vereinigung (volume 30, issue 3).
Fractal behavior and long-range dependence have been observed in an astonishing number of physical systems. Either phenomenon has been modeled by self-similar random functions, thereby implying a linear relationship between fractal dimension, a measure of roughness, and Hurst coefficient, a measure of long-memory dependence. This letter introduces simple stochastic models which allow for any combination of fractal dimension and Hurst exponent. We synthesize images from these models, with arbitrary fractal properties and power-law correlations, and propose a test for self-similarity. PACS numbers: 02.50.Ey, 02.70-c, 05.40-a, 05.45.Df
Point process models are widely used tools to issue forecasts or assess risks. In order to check which models are useful in practice, they are examined by a variety of statistical methods. We transfer the concept of consistent scoring functions, which are principled statistical tools to compare forecasts, to the point process setting. The results provide a novel approach for the comparative assessment of forecasts and models and encompass some existing testing procedures.
Context Breeding programs aim at improving the genetic characteristics of livestock populations with respect to productivity, fitness and adaptation, while controlling negative effects such as inbreeding or health and welfare issues. As breeding is affected by a variety of interdependent factors, the analysis of the effect of certain breeding actions and the optimisation of a breeding program are highly complex tasks. Aims This study was conducted to display the potential of using stochastic simulation to analyse, evaluate and compare breeding programs and to show how the Modular Breeding Program Simulator (MoBPS) simulation framework can further enhance this. Methods In this study, a simplified version of the breeding program of Göttingen Minipigs was simulated to analyse the impact of genotyping and optimum contribution selection in regard to both genetic gain and diversity. The software MoBPS was used as the backend simulation software and was extended to allow for a more realistic modelling of pig breeding programs. Among others, extensions include the simulation of phenotypes with discrete observations (e.g. teat count), variable litter sizes, and a breeding value estimation in the associated R-package miraculix that utilises a graphics processing unit. Key results Genotyping with the subsequent use of genomic best linear unbiased prediction (GBLUP) led to substantial increases in genetic gain (15.3%) compared with a pedigree-based BLUP, while reducing the increase of inbreeding by 24.8%. The additional use of optimum genetic selection was shown to be favourable compared with the plain selection of top boars. The use of graphics processing unit-based breeding value estimation with known heritability was ~100 times faster than the state-of-the-art R-package rrBLUP. Conclusions The results regarding the effect of both genotyping and optimal contribution selection are in line with well established results. Paired with additional new features such as the modelling of discrete phenotypes and adaptable litter sizes, this confirms MoBPS to be a unique tool for the realistic modelling of modern breeding programs. Implications The MoBPS framework provides a powerful tool for scientists and breeders to perform stochastic simulations to optimise the practical design of modern breeding programs to secure standardised breeding of high-quality animals and answer associated research questions.
Principal Component Analysis (PCA) is a well known procedure to reduce intrinsic complexity of a dataset, essentially through simplifying the covariance structure or the correlation structure. We introduce a novel algebraic, model-based point of view and provide in particular an extension of the PCA to distributions without second moments by formulating the PCA as a best low rank approximation problem. In contrast to hitherto existing approaches, the approximation is based on a kind of spectral representation, and not on the real space. Nonetheless, the prominent role of the eigenvectors is here reduced to define the approximating surface and its maximal dimension. In this perspective, our approach is close to the original idea of Pearson (1901) and hence to autoencoders. Since variable selection in linear regression can be seen as a special case of our extension, our approach gives some insight, why the various variable selection methods, such as forward selection and best subset selection, cannot be expected to coincide. The linear regression model itself and the PCA regression appear as limit cases.
Additional file 4 Supplementary Table 4. Functional annotation of SNPs dependent on the FST value between two colonies.
Background Göttingen Minipigs (GMP) is the smallest commercially available minipig breed under a controlled breeding scheme and is globally bred in five isolated colonies. The genetic isolation harbors the risk of stratification which might compromise the identity of the breed and its usability as an animal model for biomedical and human disease. We conducted whole genome re-sequencing of two DNA-pools per colony to assess genomic differentiation within and between colonies. We added publicly available samples from 13 various pig breeds and discovered overall about 32 M loci, ~ 16 M. thereof variable in GMPs. Individual samples were virtually pooled breed-wise. F ST between virtual and DNA pools, a phylogenetic tree, principal component analysis (PCA) and evaluation of functional SNP classes were conducted. An F-test was performed to reveal significantly differentiated allele frequencies between colonies. Variation within a colony was quantified as expected heterozygosity. Results Phylogeny and PCA showed that the GMP is easily discriminable from all other breads, but that there is also differentiation between the GMP colonies. Dependent on the contrast between GMP colonies, 4 to 8% of all loci had significantly different allele frequencies. Functional annotation revealed that functionally non-neutral loci are less prone to differentiation. Annotation of highly differentiated loci revealed a couple of deleterious mutations in genes with putative effects in the GMPs . Conclusion Differentiation and annotation results suggest that the underlying mechanisms are rather drift events than directed selection and limited to neutral genome regions. Animal exchange seems not yet necessary. The Relliehausen colony appears to be the genetically most unique GMP sub-population and could be a valuable resource if animal exchange is required to maintain uniformity of the GMP.
Since the calculation of a genomic relationship matrix needs a large number of arithmetic operations, fast implementations are of interest. Our fastest algorithm is more accurate and 25× faster than a AVX double precision floating-point implementation.
The R-package MoBPS provides a computationally efficient and flexible framework to simulate complex breeding programs and compare their economic and genetic impact. Simulations are performed on the base of individuals. MoBPS utilizes a highly efficient implementation with bit-wise data storage and matrix multiplications from the associated R-package miraculix allowing to handle large scale populations. Individual haplotypes are not stored but instead automatically derived based on points of recombination and mutations. The modular structure of MoBPS allows to combine rather coarse simulations, as needed to generate founder populations, with a very detailed modeling of todays' complex breeding programs, making use of all available biotechnologies. MoBPS provides pre-implemented functions for common breeding practices such as optimum genetic contributions and single-step GBLUP but also allows the user to replace certain steps with personalized and/or self-written solutions.
Matern hard-core processes are classical examples for point processes obtained by dependent thinning of (marked) Poisson point processes. We present a generalization of the Matern models which encompasses recent extensions of the original Matern hard-core processes. It generalizes the underlying point process, the thinning rule, and the marks attached to the original process. Based on our model, we introduce processes with a clear interpretation in the context of max-stable processes. In particular, we prove that one of these processes lies in the max-domain of attraction of a mixed moving maxima process.