Eggplant (Solanum melongena L.) shows remarkable diversity in fruit shape, making it an excellent model for studying shape variation. Eggplant fruit shape influences consumer preference and plays an important role in the classification of commercial varieties and germplasm. Despite its importance, existing classification systems are limited to description without quantitative criteria, differ by country or region and fail to fully capture the diversity of eggplant fruit shapes. In the present study, thirteen shape categories were identified using a decision tree model with Gini index-based variable selection. Ten key attributes that largely determine fruit shape were identified and high accuracy (92.59%) classification rules were generated. Five other methods, including random forest, XGBoost, SVM, K-means and GMM, were also applied to fruit shape classification, but they proved less robust for classification compared to the decision tree. The shape modeling informed the key attribute selection for the QTL-seq and GWAS analyses. Four QTLs controlling Fruit Shape Index (FSI) and Proximal Angle Micro (PAMi) were detected using GWAS and QTL-seq. The candidate gene SmFSI3.1/SmFL, a member of the SUN/IQD family, was over-expressed in tomato and resulted in elongated fruits, indicating the positive roles of this gene in regulating fruit elongation in eggplant. In summary, we developed an accurate and reproducible model for classifying eggplant fruit shapes, which is of significance for eggplant breeding and variety classification. Moreover, we verified the function of the causal gene responsible for fsi3.1/fl3.1 locus, providing a foundation for understanding the genetic regulation of fruit shape in eggplant.
Tomato fruit weight is primarily controlled by stable genetic effects and domestication-/improvement-associated loci, including prominent chromosome 5 signals, and integrating SNP, INDEL, and SV diversity improves candidate-locus discovery, biological interpretation, and genomic prediction across environments. Tomato fruit weight (fw) is a major breeding target shaped by domestication, crop improvement, and environmental variation. We investigated the genetic architecture and genomic predictability of fw in the Varitome population representing the tomato domestication continuum, including Solanum pimpinellifolium (SP), S. lycopersicum var. cerasiforme (SLC), and cultivated S. lycopersicum (SLL). Genome-wide SNPs, insertions and deletions (INDELs), and structural variants (SVs) were analyzed individually and in combination to assess their contributions to association mapping, candidate locus discovery, and genomic prediction. Population structure based on all three variant classes clearly separated the domestication groups. Genome-wide association analyses identified 15 SNP, 10 INDEL, and 10 SV loci significantly associated with fw across environments. Several associations co-localized with known fw genes, including fw2.2/CNR, fw3.2/KLUH, fw11.3/CSR, lc/WUSCHEL, and fas/CLAVA3 supporting the biological relevance of the detected signals. Chromosome (Chr) 5 was particularly notable, with a 390-bp deletion at chr5_45,551,023 detected across all four environments and an INDEL at chr5_55,361,233 detected across three environments, suggesting stable improvement-associated candidate loci for fw variation validation. Multiple loci exhibited clear allele-frequency shifts from SP through SLC to SLL, consistent with selection during domestication and improvement, whereas others were detected almost exclusively in cultivated germplasm, suggesting more recent breeding-associated origins. Fw displayed strong genetic control, with genotype explaining more than 92
Understanding the impact of domestication on deleterious mutations has fascinated evolutionary biologists and breeders alike. A 'cost of domestication' has been reported for some organisms through accumulation of gene disruptions or radical amino acid changes. However, recent evidence paints a more complex picture of this phenomenon in different domesticated species. In this study, we used genomic sequences of 253 tomato accessions to investigate the evolution of deleterious mutations and genomic structural variants (SVs) through tomato domestication history. We apply phylogeny-based methods to identify deleterious mutations in the domesticated tomato as well as its semi-wild and wild relatives. Our results implicate a downward trend throughout domestication in the number of genetic variants, regardless of their functional impact. This suggests that demographic factors have reduced overall genetic diversity, leading to lower deleterious load and SVs as well as loss of some beneficial alleles during tomato domestication. However, we detected an increase in proportions of nonsynonymous and deleterious alleles (relative to synonymous and neutral nonsynonymous alleles, respectively) during the initial stage of tomato domestication in Ecuador. Additionally, deleterious alleles in the commonly cultivated tomato seem to be more frequent than expected under a neutral hypothesis of molecular evolution. Our analyses also revealed frequent deleterious alleles in several well-studied tomato genes, probably involved in response to biotic and abiotic stress as well as fruit development and flavour regulation. To provide a practical guide for breeding experiments, we created TomDel, a public searchable database of 21,162 potentially deleterious alleles identified in this study (hosted on the Solanaceae Genomic Network; https://solgenomics.net/).
Abstract The process of plant domestication is often protracted, involving underexplored intermediate stages with important implications for the evolutionary trajectories of domestication traits. Previously, tomato domestication history has been thought to involve two major transitions: one from wild Solanum pimpinellifolium L. to a semidomesticated intermediate, S. lycopersicum L. var. cerasiforme (SLC) in South America, and a second transition from SLC to fully domesticated S. lycopersicum L. var. lycopersicum in Mesoamerica. In this study, we employ population genomic methods to reconstruct tomato domestication history, focusing on the evolutionary changes occurring in the intermediate stages. Our results suggest that the origin of SLC may predate domestication, and that many traits considered typical of cultivated tomatoes arose in South American SLC, but were lost or diminished once these partially domesticated forms spread northward. These traits were then likely reselected in a convergent fashion in the common cultivated tomato, prior to its expansion around the world. Based on these findings, we reveal complexities in the intermediate stage of tomato domestication and provide insight on trajectories of genes and phenotypes involved in tomato domestication syndrome. Our results also allow us to identify underexplored germplasm that harbors useful alleles for crop improvement.
Running head: 1 Network analyses of tomato fruit shape regulation 2 3 Author for correspondence: 4 Esther van der Knaap 5 Ohio State University, Ohio Agricultural Research and Development Center 6 Department of Horticulture and Crop Science 7 1680 Madison Ave 8 Wooster OH 44691 9 330-263-3822 10 Vanderknaap.1@osu.edu 11 Plant Physiology Preview. Published on May 4, 2015, as DOI:10.1104/pp.15.00379
Within the cultivated tomato germplasm, sun, ovate and fs8.1 are the three predominant QTLs controlling fruit elongation. Although SUN and OVATE have been cloned, their role in plant growth and development are not well understood. To compare and contrast the effects of the three QTLs in a homogeneous background, we developed near isogenic lines (NILs) in the wild species Solanum pimpinellifolium LA1589 background. We carried out detailed morphological characterization of reproductive and vegetative organs in the single, double and triple NILs and determined the epistatic interactions of the three loci affecting fruit shape. The phenotypic evaluations demonstrated that the three loci regulate unique aspects of ovary and fruit elongation and in different temporal manners. The strongest effect on organ shape was caused by sun. In addition to fruit shape, sun also affected leaf and sepal elongation and stem thickness. The synergistic interaction between sun and ovate or fs8.1 suggested that the pathways involving SUN, OVATE and the gene(s) underlying fs8.1 may converge at a common node. The results of an extensive profiling analysis suggested that the degree of fruit elongation was not related to the accumulation of any of the classical hormones.
SUN controls elongated tomato (Solanum lycopersicum) shape early in fruit development through changes in cell number along the different axes of growth. The gene encodes a member of the IQ domain family characterized by a calmodulin binding motif. To gain insights into the role of SUN in regulating organ shape, we characterized genome-wide transcriptional changes and metabolite and hormone accumulation after pollination and fertilization in wild-type and SUN fruit tissues. Pericarp, seed/placenta, and columella tissues were collected at 4, 7, and 10 d post anthesis. Pairwise comparisons between SUN and the wild type identified 3,154 significant differentially expressed genes that cluster in distinct gene regulatory networks. Gene regulatory networks that were enriched for cell division, calcium/transport, lipid/hormone, cell wall, secondary metabolism, and patterning processes contributed to profound shifts in gene expression in the different fruit tissues as a consequence of high expression of SUN. Promoter motif searches identified putative cis-elements recognized by known transcription factors and motifs related to mitotic-specific activator sequences. Hormone levels did not change dramatically, but some metabolite levels were significantly altered, namely participants in glycolysis and the tricarboxylic acid cycle. Also, hormone and primary metabolite networks shifted in SUN compared with wild-type fruit. Our findings imply that SUN indirectly leads to changes in gene expression, most strongly those involved in cell division, cell wall, and patterning-related processes. When evaluating global coregulation in SUN fruit, the main node represented genes involved in calcium-regulated processes, suggesting that SUN and its calmodulin binding domain impact fruit shape through calcium signaling.
Classification and characterization of the shape of plant organs are important tools for plant biologists, breeders and growers. Here we use boundary measurements, i.e. contour morphometric data, of scanned tomato fruits in conjunction with elliptic Fourier shape modeling and Bayesian classification techniques to find the optimum number of shape categories. Our findings show that there are nine computationally and visually distinct tomato shape categories: ellipsoid, flat, heart, long, long rectangular, rectangular, round, obovoid, and oxheart. Analyses of fruits from a diverse set of tomato accessions demonstrate that some varieties carry fruits that conform to predominantly one shape category while others carry fruits that conform to multiple shape categories. In particular the categories oxheart and long rectangular feature fruit that tend to equivalently fit several categories of shape, while the flat and obovoid categories contain fruit that consistently conform exclusively to a single category. The findings show that elliptic Fourier shape modeling and Bayesian classification provide an excellent tool for further in depth analyses of fruit shape variation that may occur across varieties and/or result from growth under different environmental conditions.
Region segmentation and edge detection are standard image processing operations.Clustering can be used for region segmentation.However, often clustering results depend on the selection of various parameters, such as the number of clusters, or the clustering algorithm used.The framework presented here employs the result of edge detection on the original image, as well as on the clustering results of the same image, to automatically select (according to some agreement measure) the optimal number of clusters, and the corresponding (best) segmentation.The framework supports an extended pixel representation in which other information, such as texture, can be incorporated in addition to edge and region information.To illustrate this framework, the edge guided clustering algorithm presented here, uses the Canny edge detection approach to guide region identification through fuzzy k-means clustering.Experimental results on benchmark images for which manual segmentation is available as reference illustrate the effectiveness of this approach.
Survival analysis is a procedure of data analysis focusing on time until an event occurs. In the medical field, predicted life expectancy is a highly significant factor in the decision making process for both the patient and the medical practitioner i.e. when making decision on palliative care and hospice referral, initiation of medications, and avoidance of aggressive therapies. The conventional statistical approach faces many challenges in handling the nature of the survival analysis datasets which often are censored data, and the difficulties in managing the complex, non-linear relationships between the prognostic factors and the patient's tumor progression. Also the statistical approach omits the need in prediction of the patient's prognosis since it does not take into account that all patients are individual and unique cases. The aim of this study is to develop a survival prediction model for breast cancer patients using Fuzzy Classifier (FC). The FC method applied is a new approach to classifying datasets with imbalanced and overlapping problems which is particularly effective in managing survival data since the data is widely known as imbalanced in nature and very rarely normally distributed. The results from a comparative study on FC, PNN and CART using Wisconsin breast cancer datasets are presented, where FC classification yields better results than the other two methods.
This paper introduces a new technique for feature selection and illustrates it on a real data set. Namely, the proposed approach creates subsets of attributes based on two criteria: (1) individual attributes have high discrimination (classification) power; and (2) the attributes in the subset are complementary that is, they misclassify different classes. The method uses information from a confusion matrix and evaluates one attribute at a time.
Three methods for attribute reduction in conjunction with Neural Networks, Naive Bayes, and k-Nearest Neighbor classifiers are investigated here when classifying a particularly challenging data set. The difficulty encountered with this data set is mainly due to the high dimensionality and to some inbalance between classes. As a result of this research, a subset of only 8 attributes (out of 34) is identified leading to a 92.7% classification accuracy. The confusion matrix analysis identifies class 7 as the one poorly learned across all combinations of attributes and classifiers. This information can be further used to upsample this underrepresented class or to investigate a classifier less sensitive to imbalance.
Classes of real world datasets have various properties (such as imbalance, size, complexity, and class distribution) that make the classification task more difficult. We investigate the robustness of six classification techniques over data having various combinations of the above mentioned properties. One artificial domain and six real world datasets are used in these experiments. Results of our analysis point to the frequency-based classifiers (such as the fuzzy and the Bayes classifiers) as being more robust over various imbalance, size, complexity, and training distribution.
Several issues arise when we consider building classifiers in general, and fuzzy classifiers in particular. These issues include but are not limited to attribute/feature selection, adoption of a specific approach/algorithm, evaluate the classifier performance, etc. We consider the opportunities that such classifiers have to offer and contrast them with the challenges they pose.
This research shows work in progress in comparing the genomes of two related organisms, namely the two North American strains of Histoplasma Capsulatum, in an effort to understand why one of them (G217B) is more infectious then the other (Wu24). For this we employ bioinformatics analysis tools along with Matlab and Per1 programs. The results shown here indicate that the two strains have many genes in common (70%). Also othologs, paralogs, and gene families are reported. Further, corresponding contigs having common genes are found.
This research proposes a segmentation algorithm for detecting nuclei and other regions of interests in biological tissues such as kidney and liver cells. The algorithm maximizes the edge coincidence between the edge obtained by Canny method and the edge resulted from the k-means clustering. The algorithm's effectiveness is illustrated on two images for which gegmentation of four components is required: nuclei, red blood cell, tubules/sinusoids, and cytoplasm.
Anca Ralescu合作论文数School of Computing Sciences and Informatics
College of Engineering9