Having attended the First International Conference on Teaching Statistics (ICOTS) in Sheffield in 1982 (and a few since then), I found it informative to look back and see what has happened in statistical education during the past three decades. In this chapter, I comment on some consistent themes and my perception of how the focus has changed over time, particularly from the viewpoint of training beyond the introductory course for students in other disciplines. Some comment is also made on particularly influential publications and other relevant meetings overseen by the International Association of Statistical Education (IASE). My vantage point is from involvement in, and then responsibility for, teaching biometry in an agriculture department for two decades, followed by another decade in university and statistical society management where I was not directly involved in the delivery of such coursework training.
Estimates of line performance for many traits derived using data from field trials vary across trials due to variation caused by the genotype by environment interaction (GEI) effects. The relative magnitude of the genetic (G) to the total or phenotype (G+GE) variation, commonly called line-mean heritability (2), is a statistic used to estimate the reliability of the line performance estimators. There are three ways of interpreting (22). Firstly, and most commonly, it is interpreted as the proportion of the total or phenotypic variance among line means due to the genetic variation among the lines. Secondly, it is the expected value of the correlation of line mean genotype values estimated from two identical trials. In addition, the square root of the line-mean heritability is the correlation between line means estimated from the field trials and their “true” genetic value. This is known as accuracy in animal breeding. Finally, it is the regression coefficient of the line mean values estimated from two identical trials. How well do the estimates of line means from the field trials predict future estimates of the line means? In general, all three interpretations of line-mean heritability measure the repeatability of genotype performance from trial to trial. In the context of field trials, the line-mean heritability is probably best referred to as genetic repeatability.
Association mapping currently relies on the identification of genetic markers. Several technologies have been adopted for genetic marker analysis, with single nucleotide polymorphisms (SNPs) being the most popular where a reasonable quantity of genome sequence data are available. We describe several tools we have developed for the discovery, annotation, and visualization of molecular markers for association mapping. These include autoSNPdb for SNP discovery from assembled sequence data; TAGdb for the identification of gene specific paired read Illumina GAII data; CMap3D for the comparison of mapped genetic and physical markers; and BAC and Gene Annotator for the online annotation of genes and genomic sequences.
We propose a hybrid classification system for predicting peptide binding to major histocompatibility complex (MHC) molecules. This system combines Support Vector Machine (SVM) and Stabilized Matrix Method (SMM). Its performance was assessed using ROC analysis, and compared with the individual component methods using statistical tests. The preliminary test on four HLA alleles provided encouraging evidence for the hybrid model. The datasets used for the experiments are publicly accessible and have been benchmarked by other researchers.
The main goal of pharmacogenomics is to study the effects of genetic variation on patient responses to therapies. Its applications range from the evaluation of safety and efficacy of treatment to the optimization of therapies and therapeutic regimens. Pharmacogenomics is becoming increasingly important in immunology, for the development of new generation vaccines, immunotherapies and transplantation. The human immune system is a complex and adaptive learning system which operates at multiple levels: molecules, cells, organs, organisms, and groups of organisms. Immunologic research, both basic and applied, needs to deal with this complexity. We increasingly use mathematical modeling and computational simulation in the study of the immune system and immune responses. Thus, quantitative models that appropriately capture the complexity in architecture and function of the immune system are an integral component of the personalized medicine efforts. In silico models of the immune system can provide answers to a variety of questions, including understanding the general behavior of the immune system, the course of disease, effects of treatment, analysis of cellular and molecular interactions, and simulation of laboratory experiments. We herein present the ImmunoGrid project that integrates molecular and system level models of the immune system and processes for in silico studies of the immune function. The ImmunoGrid simulator uses Grid technologies, enabling computational simulation of the immune system at the natural scale, perform a large number of simulated experiments, capture the diversity of the immune system between individuals, and provide a basis for therapeutic approaches tailored to the individual genetic make-up. Keywords: Pharmacogenomics applications, pharmacogenomics in personalized medicine, computational modeling, immune system simulation, ImmunoGrid
SELECTION REPRESENTS a costly and important part of sugarcane breeding programs. Previous research showed that cane yields in small single-row plots are affected strongly by competition effects, and that a high weighting in selection indices should be placed on CCS in small single-row plots to maximise gains for economic value. This led to a new selection system being suggested, involving initial screening of large numbers of clones in 5 m plots with heavy selection pressure for CCS followed by two stages of selection in multi-row plots. A stochastic simulation model using assumptions on relevant parameters (genetic, error, competition, and GE variances, genetic correlations) was developed to predict gains from alternative selection systems. Field trials were conducted in the Burdekin region to assess realised gains from alternative selection trial designs to validate and refine assumptions important in the model. The model was then used to predict genetic gains in selection systems with a wide range of configurations (e.g. plot size, replicate number, number of sites, selection criteria, selection intensity in each stage, and number of stages). Based on the results, it was recommended that three stages of clonal selection (following current family selection in stage 1 be performed in core breeding programs. This should involve firstly screening clones in small (1 row × 5 m) plots, with a selection index biased strongly toward CCS, but also including cane yield estimated via visual grade. Selected clones should then be evaluated in two further stages – the first one consisting of 4-row plots at four sites with a single replicate per clone per site. Clones selected from this stage should then be evaluated in 4-row plots at four sites but with two replicates per site. The recommended system has been introduced in the Burdekin selection system for further practical evaluation and it is recommended that it be assessed in other regions. The research conducted here also emphasised the importance of using optimal selection indices in single-row plots, in order to maximise gains from selection.
A Phenome Map is a representation of all the regions of a genome that influence heritable phenotypic variation for a trait, and a Phenome Atlas consists of the integration of all available phenome maps with a description of the methodologies that were used to produce the maps. A Phenome Atlas Toolbox is a set of tools and methodologies for producing the Phenome Atlas. The Wheat Phenome Atlas (WPA) will be an integration of phenotypic data (17 million data points for 80 traits from 10,000 international field trials collected during more than 40 years) generated by CIMMYT and partners on approximately 13,000 wheat lines (for which pedigrees are known) with greater than 26 million DArT marker data points obtained by genotyping these lines. To generate this amount of phenotypic data would cost over $500 million today.
Sugarcane crop residues (‘trash’) have the potential to supply nitrogen (N) to crops when they are retained on the soil surface after harvest. Farmers should account for the contribution of this N to crop requirements in order to avoid over-fertilisation. In very wet tropical locations, the climate may increase the rate of trash decomposition as well as the amount of N lost from the soil–plant system due to leaching or denitrification. A field experiment was conducted on Hydrosol and Ferrosol soils in the wet tropics of northern Australia using 15N-labelled trash either applied to the soil surface or incorporated. Labelled urea fertiliser was also applied with unlabelled surface trash. The objective of the experiment was to investigate the contribution of trash to crop N nutrition in wet tropical climates, the timing of N mineralisation from trash, and the retention of trash N in contrasting soils. Less than 6% of the N in trash was recovered in the first crop and the recovery was not affected by trash incorporation. Around 6% of the N in fertiliser was also recovered in the first crop, which was less than previously measured in temperate areas (20–40%). Leaf samples taken at the end of the second crop contined 2–3% of N from trash and fertilizer applied at the beginning of the experiment. Although most N was recovered in the 0–1.5 m soil layer there was some evidence of movement of N below this depth. The results showed that trash supplies N slowly and in small amounts to the succeeding crop in wet tropics sugarcane growing areas regardless of trash placement (on the soil surface or incorporated) or soil type, and so N mineralisation from a single trash blanket is not important for sugarcane production in the wet tropics.
Finite mixture models are being increasingly used to model the distributions of a wide variety of random phenomena. While normal mixture models are often used to cluster data sets of continuous multivariate data, a more robust clustering can be obtained by considering the t mixture model-based approach. Mixtures of factor analyzers enable model-based density estimation to be undertaken for high-dimensional data where the number of observations n is very large relative to their dimension p. As the approach using the multivariate normal family of distributions is sensitive to outliers, it is more robust to adopt the multivariate t family for the component error and factor distributions. The computational aspects associated with robustness and high dimensionality in these approaches to cluster analysis are discussed and illustrated.
The main purpose of this article is to gain an insight into the relationships between variables describing the environmental conditions of the Far Northern section of the Great Barrier Reef, Australia, Several of the variables describing these conditions had different measurement levels and often they had non-linear relationships. Using non-linear principal component analysis, it was possible to acquire an insight into these relationships. Furthermore. three geographical areas with unique environmental characteristics could be identified. Copyright (c) 2005 John Wiley & Sons, Ltd.
The primary aim of this paper is to provide an introduction to a three-mode method of clustering and the useful role it can fulfil in clustering chemical three-mode data. The analysis of trace elements present in body tissues of diseased blue crabs caught along the coast of North Carolina serves as an example. The clustering method succeeded in separating the diseased crabs from healthy controls, lending support to the hypothesis that the trace elements were the origin of the blue crabs' disease. Copyright (C) 2005 John Wiley & Sons, Ltd.
When studying genotype X environment interaction in multi-environment trials, plant breeders and geneticists often consider one of the effects, environments or genotypes, to be fixed and the other to be random. However, there are two main formulations for variance component estimation for the mixed model situation, referred to as the unconstrained-parameters (UP) and constrained-parameters (CP) formulations. These formulations give different estimates of genetic correlation and heritability as well as different tests of significance for the random effects factor. The definition of main effects and interactions and the consequences of such definitions should be clearly understood, and the selected formulation should be consistent for both fixed and random effects. A discussion of the practical outcomes of using the two formulations in the analysis of balanced data from multi-environment trials is presented. It is recommended that the CP formulation be used because of the meaning of its parameters and the corresponding variance components. When managed (fixed) environments are considered, users will have more confidence in prediction for them but will not be overconfident in prediction in the target (random) environments. Genetic gain (predicted response to selection in the target environments from the managed environments) is independent of formulation.
In broader catchment scale investigations, there is a need to understand and ultimately exploit the spatial variation of agricultural crops for an improved economic return. In many instances, this spatial variation is temporally unstable and may be different for various crop attributes and crop species. In the Australian sugar industry, the opportunity arose to evaluate the performance of 231 farms in the Tully Mill area in far north Queensland using production information on cane yield (t/ha) and CCS (a fresh weight measure of sucrose content in the cane) accumulated over a 12-year period. Such an arrangement of data can be expressed as a 3-way array where a farm × attribute × year matrix can be evaluated and interactions considered. Two multivariate techniques, the 3-way mixture method of clustering and the 3-mode principal component analysis, were employed to identify meaningful relationships between farms that performed similarly for both cane yield and CCS. In this context, farm has a spatial component and the aim of this analysis was to determine if systematic patterns in farm performance expressed by cane yield and CCS persisted over time. There was no spatial relationship between cane yield and CCS. However, the analysis revealed that the relationship between farms was remarkably stable from one year to the next for both attributes and there was some spatial aggregation of farm performance in parts of the mill area. This finding is important, since temporally consistent spatial variation may be exploited to improve regional production. Alternatively, the putative causes of the spatial variation may be explored to enhance the understanding of sugarcane production in the wet tropics of Australia.
Previous research has reported both agreements and serious anomalies in relationships between production attributes of sugarcane varieties in variety trials (VTs) and commercial production (CP). This paper examines VT and CP data for tonnes of cane per hectare (TCH) and sugar content (CCS). Data, analysed by REML, included 107 VTs and 54 CP mill years for 9 varieties from the mill districts of Mulgrave, Babinda, and Tully for harvest years 1982–99. Important consistencies included high TCH of Q152, high CCS of Q117 and Q120, and low CCS of H56-752. Significant anomalies existed with respect to TCH for Q113, Q117, Q120, Q122, Q138, and H56-752 and to CCS for Q113 and Q124. Investigation of these anomalies was assisted by access to independent REML analyses of CP data for 65 692 individual Tully cane blocks from 1988 to 1999 and by the knowledge of persons familiar with the preferential uses of varieties by farmers. Minor anomalies were due to limited year or mill area data. Q124 TCH was deemed to be decreased and its CCS increased by severe disease in Babinda CP in the extremely wet 1998 and 1999 seasons. Other serious anomalies have credible but unsubstantiated explanations. The most convincing, for Q113, Q117, Q138, and H56-752, are that these varieties were deployed unevenly with regard to late season harvesting, predominant use or avoidance on high fertility soils, or use confined to low fertility sandy soils, respectively. Uneven deployment results in confounding of these effects in the varietal CP statistics at mill area level. It is concluded that VTs cannot be enhanced to anticipate or evaluate most effects of uneven deployment. They give adequate predictions of relative CP performance for varieties deployed evenly across confounding influences. Routine analyses of individual block CP data would be useful and enhanced by addition of relevant information to the block records.
AbstractBackground: Previous studies of stability and relapse after orthodontic treatment report short-term stability is generally followed by slow relapse to the original condition. What these studies do not report is whether this relapse is continuous or interspersed with periods of improvement or stability.Methods: A subjective 0–10 index of malocclusion was used to record post-treatment stability and relapse over 10 to 12 years following fixed appliance orthodontic treatment of 24 patients. The severity scores were plotted on timelines.Results: Episodes of change, both favourable and unfavourable, were interspersed with episodes of stability.Conclusions: Changes in the first 3 and 12 months post-treatment are indicative of the 10 to 12 years post-treatment outcomes. This index may provide a useful instrument to analyze patients and/or their study models longitudinally.