Yield forecasting depends on accurate tree fruit counts and mean size estimation. This information is generally obtained manually, requiring many hours of work. Artificial vision emerges as an interesting alternative to obtaining more information in less time. This study aimed to test and train YOLO pre-trained models based on neural networks for the detection and count of pears and apples on trees after image analysis; while also estimating fruit size. Images of trees were taken during the day and at night in apple and pear trees while fruits were manually counted. Trained models were evaluated according to recall, precision and F1score. The correlation between detected and counted fruits was calculated while fruit size estimation was made after drawing straight lines on each fruit and using reference elements. The precision, recall and F1score achieved by the models were up to 0.86, 0.83 and 0.84, respectively. Correlation coefficients between fruit sizes measured manually and by images were 0.73 for apples and 0.80 for pears. The proposed methodologies showed promising results, allowing forecasters to make less time consuming and accurate estimates compared to manual measurements. Highlights: The number of fruits in apple and pear trees, could be estimated from images with promising results. The possibility of estimating the fruit numbers from images could reduce the time spent on this task, and above all, the costs. This allow growers to increase the number of trees sampled to make yield forecasts.
Characterization of plant material conserved in germplasm banks allows the study and analysis of the genetic variability within a collection. When germplasm banks have a large number of accessions, field evaluation should be performed using assays with manageable accession subsets. Common checks connecting the different assays are required to compare these accession subsets. In this study, the Generalized Procrustes Analysis was proposed as a basis for obtaining a factorial plane where all individuals are projected. This technique is applied to genotypes common to all assays, iteratively generating scale factors and rotation matrices. Accessions only belonging to a given assay are considered supplementary elements. This proposal was illustrated using datasets of 54 maize accessions from the Pergamino Active Germplasm Bank of the Experimental Station at the Instituto Nacional de Tecnología Agropecuaria (INTA) in Argentina. The proposal achieved highly satisfactory results. Highlights: In field evaluation of large germplasm collections, the material must be divided into manageable experimental trials, in which different accession subsets are evaluated in different environments. A new algorithm based on Generalized Procrustes Analysis (GPA) allowed to find the consensus of several configurations of individuals connected by common checks. The characterization data analysis strategy was illustrated using a set of accessions from the Argentine Maize Germplasm Bank. The new proposal stands as a useful tool for evaluate germplasm collections, providing good results with easy implementation and considering the multivariate structure of the data set.
In field evaluation of large germplasm collections, the experimental material must be divided into manageable experimental trials. In each trial, a set of entries is evaluated by several descriptors. The sets of individuals are different between trials, but an important characteristic of this design is the presence of a certain number of entries involved in all trials, these ones are used as connections between them. The information obtained from these trials is presented by incomplete matrices, because not all individuals are evaluated in all the conditions (trials). Considering this problem, the present algorithm is based on performing a Generalized Procrustes Analysis on the individuals that are common to all conditions, considering the rest of the individuals in each condition as supplementary individuals. This algorithm allows to obtain a consensus configuration when the initial matrices are not complete but are connected by a group of individuals in common. The algorithm was illustrated by a simulation of the experimental situation with three conditions and 16 individuals.
Sample- and gene- based hierarchical cluster analyses have been widely adopted as tools for exploring gene expression data in high-throughput experiments. Gene expression values (read counts) generated by RNA sequencing technology (RNA-seq) are discrete variables with special statistical properties, such as over-dispersion and right-skewness. Additionally, read counts are subject to technology artifacts as differences in sequencing depth. This possesses a challenge to finding distance measures suitable for hierarchical clustering. Normalization and transformation procedures have been proposed to favor the use of Euclidean and correlation based distances. Additionally, novel model-based dissimilarities that account for RNA-seq data characteristics have also been proposed. Adequacy of dissimilarity measures has been assessed using parametric simulations or exemplar datasets that may limit the scope of the conclusions. Here, we propose the simulation of realistic conditions through creation of plasmode datasets, to assess the adequacy of dissimilarity measures for sample-based hierarchical clustering of RNA-seq data. Consistent results were obtained using plasmode datasets based on RNA-seq experiments conducted under widely different conditions. Dissimilarity measures based on Euclidean distance that only considered data normalization or data standardization were not reliable to represent the expected hierarchical structure. Conversely, using either a Poisson-based dissimilarity or a rank correlation based dissimilarity or an appropriate data transformation, resulted in dendrograms that resemble the expected hierarchical structure. Plasmode datasets can be generated for a wide range of scenarios upon which dissimilarity measures can be evaluated for sample-based hierarchical clustering analysis. We showed different ways of generating such plasmodes and applied them to the problem of selecting a suitable dissimilarity measure. We report several measures that are satisfactory and the choice of a particular measure may rely on the availability on the software pipeline of preference.
The establishment of relationships among taxa is an essential step in the process of cataloging and evaluation of material conserved in a germplasm bank. The...
The establishment of relationships among taxa is an essential step in the process of cataloging and evaluation of material conserved in a germplasm bank. There are several evaluation methods according to the types of the characters in the study. When the registration of the characters should be repeated in diverse environments and times, it is necessary to separate the genetic variability of the taxa from the variability due to the environment, and from the possible genotype*environment interaction variability. Consequently, pure phylogenetic relationships may be established. In this work, the feasibility of application of two strategies of statistical analysis to give solution to this problem is studied comparatively. The first one is a traditional Principal Component Analysis applied on the average characters. The second one is a set of more complex methods where each datum is originated by three ways: individuals, variables and environmental conditions, as the Multiple Factorial Analysis and the Generalized Procrustes Analysis. While the resulting configurations were all equivalent, three-way methods allow the interpretation of genotype*environment.
A set of 34 quinoa populations from the Northwest Argentina region was characterised using quantitative and qualitative phenotypic traits in an experiment conducted in the province of Jujuy, Argentina. A selection of quinoa descriptors from the Bioversity International (former IBPGR) list was applied, and data were analyzed using descriptive and multivariate techniques. Morphological and phenological traits variation was observed among accessions collected in contrasted ecogeographic zones of this Andean region. On the basis of quantitative traits, both the principal component analysis and the Cluster Analysis differentiated between accessions from the highlands, transition zone, central dry valleys and eastern valleys. On the other hand, the principal coordinates analysis based on qualitative traits only discriminated accessions from transition zone and eastern valleys. The correlation between both characterisations was fairly low suggesting that individual characterisations offer information that can be complementary. The accessions from the highlands and dry valleys presented the more advanced domesticated traits, while accessions from transition zone and eastern valleys showed traits more similar to wild-type related Chenopods from the Andean region. These differences are discussed on the basis of previous hypotheses about the domestication and crop diffusion processes from the southern Andes suggested for this species.
El establecimiento de relaciones entre taxones es un paso esencial en el proceso de catalogacion y evaluacion del material conservado en un Banco de Germoplasma. Existen distintos metodos de evaluacion en funcion del tipo de caracteres estudiados. Cuando el registro de caracteres se repite en el tiempo y en distintos ambientes, se debe separar la variabilidad intrinsecamente genetica entre los taxones de aquella que se debe al ambiente, y mas aun, de la posible variabilidad debida a la interaccion genotipo*ambiente para el posterior establecimiento de relaciones puramente filogeneticas. En el presente trabajo se estudia comparativamente la factibilidad de aplicacion de dos estrategias de analisis estadistico para dar solucion a este problema. La primera corresponde al analisis tradicional donde se realiza un Analisis de Componentes Principales sobre los caracteres promedios a lo largo de los diferentes ambientes; y la segunda son metodos mas complejos en los cuales cada dato es originado por tres modos: individuos, variables y condiciones ambientales, tales como el Analisis Factorial Multiple y el Analisis de Procrustes Generalizado. Si bien las configuraciones resultantes fueron todas equivalentes, los metodos de tres vias permiten la interpretacion de la interaccion genotipo*ambiente. Rev. FCA UNCUYO. 2012. 44(1): 49-64. ISSN impreso 0370-4661. ISSN (en linea) 1853-8665.
Los modelos de crecimiento de frutos describen la evolución de su tamaño a lo largo del período de desarrollo. Con fines de pronóstico, estos modelos permiten estimar en forma anticipada el tamaño que alcanzarán los frutos al momento de la cosecha. Para lograr estimaciones insesgadas del tamaño de frutos a cosecha es necesario un diseño adecuado de muestreo en la etapa de recolección de datos. El objetivo del presente trabajo fue determinar el tamaño óptimo de muestra, compuesta por árboles (n) y frutos (m), para establecer modelos de crecimiento de frutos de naranjo 'Valencia late', que permitan estimar la distribución de tamaño a la cosecha. Se trabajó con el diámetro ecuatorial de frutos previo a la cosecha, proveniente de dos huertos comerciales ubicados en la provincia de Corrientes, Argentina, durante tres temporadas. Mediante modelos mixtos se estimaron las componentes de varianzas entre árboles y frutos, y posteriormente a partir de dos tipos de metodologías se determinó el tamaño de muestra óptimo. La variabilidad entre frutos fue superior a la variabilidad entre árboles. Para la determinación del patrón de crecimiento de frutos de naranjo 'Valencia late' mediante un muestreo bietápico, se sugiere seleccionar 7 árboles y 30 frutos de cada árbol, para lograr estimaciones del diámetro ecuatorial de frutos con una precisión entre el 2 y 3%.
M. Marticorena, S. Bramardi, and R. Defacio. 2010. Characterization of maize populations in different environmental conditions by means of Three-Mode Principal Components Analysis. Cien. Inv. Agr. 37(3): 91-103. Characterization of 31 native populations of maize conserved at the germplasm bank of the INTA Pergamino Experimental Station, Argentina, was achieved by evaluating 10 quantitative attributes in two different environmental situations. The experimental design generated three-way or three-mode data, repeated observations of a set of attributes for a set of individuals in different conditions. The information was displayed in a three-dimensional array, and the structure of the data was explored using Three-Mode Principal Component Analysis, the Tucker-2 Model. A group of populations was identified that displayed homogeneous behavior in the two environments with respect to the following traits: ear length, prolificacy, grains per meter, yield and 1000 kernel weight. However, another group of populations displayed opposite behavior for the traits of plant height and ear insertion height in the different environment conditions and is indicative of the existence of genotype-environment interaction. In conclusion, Three-Mode Principal Component Analysis is an important tool for characterizing plant genetic resources when their phenotypic values are likely to be affected by environmental conditions.
Los modelos de crecimiento de frutos describen la evolución de su tamaño a lo largo del período de desarrollo. Con fines de pronóstico, estos modelos permiten estimar en forma anticipada el tamaño que alcanzarán los frutos al momento de la cosecha. Para lograr estimaciones insesgadas del tamaño de frutos a cosecha es necesario un diseño adecuado de muestreo en la etapa de recolección de datos. El objetivo del presente trabajo fue determinar el tamaño óptimo de muestra, compuesta por árboles (n) y frutos (m), para establecer modelos de crecimiento de frutos de naranjo 'Valencia late', que permitan estimar la distribución de tamaño a la cosecha. Se trabajó con el diámetro ecuatorial de frutos previo a la cosecha, proveniente de dos huertos comerciales ubicados en la provincia de Corrientes, Argentina, durante tres temporadas. Mediante modelos mixtos se estimaron las componentes de varianzas entre árboles y frutos, y posteriormente a partir de dos tipos de metodologías se determinó el tamaño de muestra óptimo. La variabilidad entre frutos fue superior a la variabilidad entre árboles. Para la determinación del patrón de crecimiento de frutos de naranjo 'Valencia late' mediante un muestreo bietápico, se sugiere seleccionar 7 árboles y 30 frutos de cada árbol, para lograr estimaciones del diámetro ecuatorial de frutos con una precisión entre el 2 y 3%.Fruit growth models are used, among other applications, to describe the evolution of fruit size throughout its development period. For forecasting, these models allow estimate in advance the size that will achieve the fruits at harvest. To have growth curves that provide unbiased estimate of the fruit size at harvest is necessary to design an appropriate sampling at the data collection stage for its construction. The objective of the present work was to determine the optimal sample size, consisting of trees (n) and fruits (m), to establish the growth patterns of orange fruits 'Valencia late' to estimate the size distribution at harvest. It was used the equatorial diameter of fruits before harvest, from two orchards located in the province of Corrientes, Argentina, during three seasons. Through mixed models were estimated components of variances between trees and fruits, and then from two types of methodologies was determined optimal sample size. The variability among fruits was higher than the variability between trees. To determine the growth pattern of orange fruit 'Valencia late' by sampling in two-stages, it is suggested to select 7 trees and 30 fruit per tree, for the estimative to reach equatorial diameter of fruit with an accuracy between 2 and 3%.
In this paper, a study of the relationship between genetic patterns, obtained by the combination of mtDNA-RFLP and PCR-amplified inter-δ sequence DNA polymorphism analysis, and relevant enological phenotypic data (fermentative power, specific productivity, volatile and total acidity) was carried out on Argentinean Saccharomyces cerevisiae isolates from north Patagonia. The use of a powerful statistical tool, Generalized Procrustes analysis, allowed us to weigh the relationship for each isolate in particular, denoting a good enough degree of agreement between molecular and physiological data for most of the population analysed. The inclusion of a physiological feature, as the killer sensitivity biotype, within identification methods resulted in a higher degree of discrimination among isolates and in better correlation between both characterizations. The combined use of methods based on molecular polymorphisms and killer biotype could be applied so as not to miss any isolate with differential enological properties in selection protocols.