Over the past two decades, omics and big data have shifted plant molecular biology from single-gene, hypothesis-driven studies to systems-level, data-driven discovery. As datasets expand in scale and diversity, bioinformatics software has become essential for routine analysis and interpretation. However, the efficiency of data exploration and evidence integration has not kept pace with data growth, leaving many datasets underutilized and only slowly translated into biological insight. A central bottleneck is the widening gap between the limited data analysis skills of many experimental biologists and the increasing complexity of biological data. TBtools was developed to narrow this gap by providing low-barrier, interactive functions for common plant omics tasks, and it has been broadly adopted. Here, we use TBtools as a decade-long case study to discuss why certain local tools achieve broad adoption in plant omics research, distill eight actionable design recommendations, and propose four capacity pillars for next-generation local workbenches: project-level data management, reproducible workflow construction, elastic remote computing, and AI-assisted navigation and automation. Together, these lessons provide a practical roadmap for accelerating the translation of omics data into biological insights.
The lychee industry is vital to agricultural economies, boosting the livelihood of farmers and regional growth. However, instability of flowering causes yield fluctuations, severely limiting industry sustainability. Stable pistil development in female flowers is essential for yield improvement, yet its molecular regulation remains poorly understood. Although APETALA2 (AP2) transcription factors regulate floral organ differentiation and pistil development, their functional role in woody perennials such as lychee is uncharacterized. In this study, two AP2 genes (LITCHI007109 and LITCHI010784) were found to exhibit high and specific expression in carpels. LITCHI007109, designated as LcANT1, is an ortholog of Arabidopsis AINTEGUMENTA (ANT). We next systematically identified the direct downstream target genes of LcANT1, the set of which were significantly enriched in biological processes related to floral organ development and carpel morphology. Notably, the carpel development-related gene LITCHI024703 (LcREV) exhibited a high level of co-expression with LcANT1. We found that the LcANT1 protein can directly bind to the promoter region of LcREV. Further evolutionary analysis indicates that the ANT-REV regulatory module is highly conserved in angiosperms, especially in Sapindaceae. Our findings establish a novel theoretical framework for understanding female flower development in lychee and offer critical gene resources and regulatory networks for molecular breeding strategies aimed at developing high-yield, stable cultivars.
BACKGROUND:Centromeres are chromosomal loci epigenetically specified by the histone variant CENH3, where kinetochores assemble to ensure accurate chromosome segregation during cell division. Their repetitive and rapidly evolving DNA has long impeded large-scale characterization. Advances in long-read sequencing now enable complete genome assemblies across species and within populations, providing opportunities to investigate how centromeres evolve and diversify over timescales from thousands to millions of years. RESULTS:Here, we generate near-telomere-to-telomere genome assemblies for eggplant, African eggplant, and wild pepper. Using CENH3 ChIP-seq, we delineate functional centromeric chromatin in these assemblies and in the cultivated pepper 'CA59', tomato 'Heinz 1706', and a wild tomato accession. These genomes harbor satellite-free centromeres across all chromosomes except chromosome 3 in tomato and its wild progenitor. Instead, centromeres are primarily composed of Ty3/Gypsy LTR retrotransposons, whose clade composition, abundance, recent activity, and spatial distribution differ among species. Centromere size scales with genome size in Solanaceae crops. Comparisons of closely related genomes reveal frequent centromere positional shifts driven by pericentromeric inversions and centromere repositioning. Synteny decays more rapidly around centromeres, consistent with elevated breakage within CENH3-binding regions. Finally, centromere haplotypes vary within species, exemplified by multiple haplotypes on four African eggplant chromosomes. CONCLUSIONS:These findings highlight the remarkable evolutionary dynamics and within-species variation of centromeres in Solanaceae crops, revealing distinct species-specific organizational patterns. This study positions Solanaceae as a promising model for comparative analyses of plant centromere evolution and provides a foundation for future research exploring how centromere variation contributes to phenotypic diversity.
Bitter gourd (Momordica charantia) is a widely cultivated vegetable and medicinal crop for its nutritional and pharmaceutical properties. Yet, the genomic basis of its domestication and key agronomic traits remains largely unexplored. Here, we present a high-quality, chromosome-level genome assembly specifically representing a smooth bitter gourd cultivar (Y1745), with a total assembly size of 327.14 Mb, a contig N50 of 11.33 Mb, and a scaffold N50 of 23.47 Mb. Comparative genomics reveals lineage-specific expansions of gene families related to stress responses and secondary metabolism. Population genomic analysis of 192 globally representative accessions identifies three distinct genetic groups (China, Southeast Asia, and South Asia Subcontinent), reflecting complex demographic histories, and uncovers strong signatures of selection associated with domestication and regional adaptation. In the genome-wide association study (GWAS) of 35 traits, we identify 893 significant trait-associated loci. Notably, two candidate genes (McWRKY and McFPF1-like 1) are associated with flowering time regulation, and one candidate gene (McEXLB1) with fruit size determination. Our work provides valuable genomic resources for understanding bitter gourd domestication and offers potential genetic targets for breeding.
Eggplant (Solanum melongena L.) is a globally important Solanaceae crop, yet trait-relevant genomic variants remain poorly characterized. Here, we perform population genomic analyses of 226 eggplant accessions sampled mainly from a major domestication center spanning Southeast Asia and South China, and find that genetic relationships closely track geographic origin. We generate chromosome-scale assemblies for 11 representative accessions using long-read sequencing and integrate six published genomes to build a pangenome resource. Using this resource, association scans identify a 12.4 Mb inversion on chromosome 10 segregating at 50.44% frequency that is strongly associated with fruit color, likely through hitchhiking with SmMYB1. We also detect variants associated with bacterial wilt resistance, including a premature stop codon in SmCYP82D47 and copy number variations in SmEPS1 and SmRoq1 homologs. Together, our results illuminate the evolution and phenotypic impact of large structural variants and provide genomic resources for eggplant genetics and breeding.
A high-quality reference genome requires not only accurate DNA sequences but also well-defined gene-structure annotations. However, many existing tools depend predominantly on automated pipelines that perform poorly when confronted with complex gene architectures, such as overlapping loci, alternative splicing patterns, and lowly expressed isoforms, resulting in incomplete or inaccurate annotations. To overcome these limitations, we developed GSAman, a standalone, ready-to-use tool that enables intuitive, what-you-see-is-what-you-get (WYSIWYG) editing of gene-structure annotations. In contrast to web-based platforms such as Apollo2, which generally require server deployment and do not provide full offline functionality, GSAman delivers a fully local, responsive interface for real-time annotation refinement, thereby improving accessibility across research settings. GSAman supports both fine-scale curations of individual genes and large-scale annotation of entire genomes. By enabling precise curation of gene models across varied genomic contexts, it directly facilitates downstream applications, including pan-genome construction, gene family evolutionary analyses, and precision crop enhancement. Using a telomere-to-telomere rice genome (MH63) annotation project as a case study, we demonstrate the practical utility of GSAman in producing a complete and accurate reference annotation, improving Benchmarking Universal Single-Copy Orthologs (BUSCO) completeness to 99.63% following manual curation. We believe GSAman will serve as a critical resource for advancing functional genomics across diverse species.
Identifying plant disease-resistance genes is essential for understanding the plant immune system and accelerating the breeding of disease-resistant crops. There is a pressing need for a method capable of accurately identifying plant disease-resistance genes on a genome-wide scale. In this study, we propose evolutionary scale modeling for LRR (ESM-LRR), a deep protein language model designed to accurately predict LRR domains, which are substantially variable structures in disease-resistance proteins. ESM-LRR achieved its highest F1 score of 0.80 on a test set using 90% identity as the matching threshold. Building on ESM-LRR, we developed R-Predictor, a plant disease-resistance gene predictor to simultaneously annotate 15 diverse domain topologies, covering characterized resistance genes across the whole genome. R-Predictor integrates 4 modules, each employing superior methods that outperform existing methods (achieving F1 scores of 0.89 for RLKs and 0.88 for NLRs), demonstrating its high accuracy and practicality in annotating plant disease-resistance genes. R-Predictor integrated with gene expression profiles to identify candidate R genes associated with grape gray mold and downy mildew, outperforming existing methods and detecting dozens of candidate R genes. Overall, this study presents a novel approach to advancing our understanding of plant immunity and facilitating crop breeding for disease resistance.
Centromeric and pericentromeric regions of most eukaryotic genomes are highly repetitive and strongly recombination-suppressed, confounding efforts to resolve genetic variation, population structure and phenotypic associations. Pepper (Capsicum annuum) centromeres are nearly devoid of satellite repeats, facilitating assembly and population-level comparison of centromeric regions. Here we integrate 9 near-complete genome assemblies, CENH3 ChIP-seq profiles from 26 diverse accessions, and resequencing and phenotypic data from ~400 cultivated and wild accessions to investigate population-level diversity and phenotypic relevance of pepper peri/centromeric regions. Functional centromere positions are largely fixed on 8 of 12 chromosomes, whereas the remaining 4 carry distinct centromeric epialleles shaped mainly by centromere repositioning and pericentromeric inversions. Pepper centromeres are embedded within ultra-long centromere-spanning haplotype (cenhap) blocks, ranging from 29.8 to 112.9 Mb and collectively covering 23.96% of the genome; each block contains only 1-4 major haplotypes. Some cenhaps may act as supergene-like units and are strongly associated with fruit traits, probably because recombination-suppressed intervals harbour multiple fruit-related genes, including OFP and F-box genes. F2 segregation assays further reveal transmission distortion of chromosomes carrying alternative cenhaps. Together, these findings highlight peri/centromeric regions as underrecognized reservoirs of agronomically important variation.
Fusarium wilt, caused by soil-borne Fusarium oxysporum f. sp. cubense (Foc), can damage banana crops globally. Herein, we first examined the different responses between a susceptible ('Baxijiao', BX) and resistant ('Yueyoukang1', YK) cultivar to Foc 4 via RNA-Seq and immunofluorescence labeling technique. The KEGG analysis highlighted the abundance of plant-pathogen interaction, plant hormone signal transduction, and MAPK signaling pathways-plant. We found that, following Foc 4 infection, GO items relevant to the cell wall were enriched in abundance. Most differentially expressed AGPs were downregulated in YK than in BX before and/or after pathogen infection. Interestingly, MaAGP15 was downregulated in the YK before and after the pathogen infection. The BX plants exhibited higher epitope levels of AGPs recognized by JIM16, LM2, and JIM8 antibodies than YK before or after pathogen infection. Functional characterization of the MaAGP15 protein revealed that it was localized to the cell membrane. The MaAGP15 promoter contains hormone and transcription factor-responsive cis-elements (E-boxes, W-boxes), and its GO terms are associated with cell wall integrity and secondary cell wall. Overexpression of MaAGP15 increased Arabidopsis susceptibility to Fo5176 and led to elevated levels of AGPs recognized by the JIM8 and LM2 antibodies. Our study identified MaAGP15 as a negative regulator of the banana immune system against Foc 4 stress.
AP2/ERF (APETALA2/Ethylene Responsive Factor) transcription factors, a class of plant-specific transcription factors, play a pivotal role in plant growth, development, metabolism, and stress response. The pineapple (Ananas comosus (L.) Merr.), a perennial fruit, belongs to the Bromeliaceae family. It is an economically important crop worldwide, which is consumed as fresh fruit, canned fruit, a fiber source, and even pharmaceutical raw material. We identified 75 AcoAP2/ERF genes in the pineapple genome, with four manually curated. They were distributed evenly on 23 chromosomes, except on LG20 and LG23. Sequence lengths, molecular weights, and intron numbers were diverse. The majority of pineapple AcoAP2/ERF genes were localized in nuclear while seven AcoERFs were located in mitochondrial or chloroplast. All pineapple AcoAP2/ERF genes possess an AP2 domain and are divided into 10 clades. Most originate from whole-genome or segmental duplication instead of transposon events. Utilizing pineapple calluses as experimental material, qRT–PCR analysis revealed that the expression of the majority of AcoAP2/ERF genes was induced in response to abscisic acid (ABA), gibberellic acid (GA), ethylene (ET), and naphthalene acetic acid (NAA). In this study, we cloned the promoter sequence of the AcoERF24 gene and divided it into three fragments to construct individual vectors. These vectors were subsequently introduced into Arabidopsis thaliana for β-glucuronidase (GUS) activity analysis, revealing variations in activity levels among the different fragments. This study not only deepens our understanding of the AcoAP2/ERF genes family in olives but also provides an important basis for subsequent studies on the regulation of AcoERF24 gene expression and biological functions.
Sexual systems in animals and plants are remarkably diverse, with dioecy having evolved independently in numerous lineages. In plants, dioecy often evolved more recently than in the best-studied animal systems, making plants especially important for understanding how separate sexes evolved independently from functionally hermaphrodite ancestors. Despite long-standing theories of developmental trade-offs in sex allocation, the underlying genetic mechanisms remain elusive. Here, we show that the XY sex determination system in the dioecious plant species Eurycorymbus cavaleriei in Sapindaceae involves two Y-linked mutations that act jointly within the developmental male-female trade-off: YUNΔ , a truncated allele that lowers the dosage of the D-class MADS-box gene YUN , and SUNMAO , a novel sRNA locus that silences the X-linked SUN allele. In females, SUN stabilizes the HD-ZIP transcription factor KUN, which is a known sex determinant in another dioecious plant, thereby promoting femaleness by increasing YUN expression; loss of SUN expression, together with the effect of YUNΔ , shifts development toward males. Two interlocking regulatory loops in this “SKY” module (SUN-KUN-YUN) fine-tunes YUN dosage. This dioecious system in E. cavaleriei likely evolved by sequential mutations in genes acting in the predicted male-female trade-off system, with their close linkage reflecting a translocation, and later recombination-suppressing inversions. ### Competing Interest Statement The authors have declared no competing interest.
Sepals and petals form the peels of pineapple fruits, which influence the size of cavities below the surface of the fruits (so-called “fruit eye”) and subsequently the fruit quality and edible rate. In this study, to investigate the underlying mechanisms controlling septal-petal formation in pineapple, we utilized a mutant of variety ‘Yulinglong’ with petaloid sepals for comparative analyses with the wild type. Phenotypic and microscopic observations confirmed the either partially or completely petalized structure of the sepals of the mutant. Comparative gene expression analysis identified two MADS-box family members AcPI and AcAP3 that are potentially associated with the petaloid sepals. Heterologous overexpression experiments in Arabidopsis and tobacco validated their functions in controlling the identity and organogenesis of sepals/petals, as well as confirmed their role in transforming sepals to petals. Protein-protein interaction experiments and gene expression profiling suggested that AcPI and AcAP3 may coordinately determine floral organogenesis in pineapple flower bud primordia differentiation. The results provide important insights into the molecular regulation of floral organ identity and peel structure formation in pineapple, which may be harnessed to improve fruit quality and edible rate for pineapple.
Mango is the second most important tropical fruit crop. Due to ever-changing environmental conditions, world mango production is facing challenges such as diseases (anthracnose and mango malformation), physiological disorders (alternate bearing), low fruit setting, poor fruit quality, short shelf life, and climate change adaptation. Breeding efforts are hindered by the long juvenile period, outdated breeding system, and high heterozygosity, resulting in a slow pace of mango improvement programs. However, over the last decade, significant advances in high-quality genome assemblies, pangenomics, genetic mapping, multiomics data, and phenomics of large populations have accelerated crop genetics and breeding. Here, we summarize recent progress on the origin and domestication of mango, advancements in genome assemblies, development of genetic maps, functional and comparative genomics, evolutionary insights, and assessments of global phenotypic and genotypic diversity, including species at risk. We also discuss the integration of multiomics approaches with quantitative genetics for crop improvement. Furthermore, we highlight the key research gaps that limit breeding efficiency and propose integrative strategies combining pangenomics, multiomics, and machine learning with improved transformation protocols and multienvironment testing to accelerate the development of climate-resilient, high-quality mango cultivars.
Most genomic studies start by mapping sequencing data to a reference genome. The quality of reference genome assembly, genetic relatedness to the studied population, and the mapping method employed directly impact variant calling accuracy and subsequent genomic analyses, introducing reference bias and resulting in erroneous conclusions. However, the impacts of reference bias have gained limited attention. This study compared population genomic analyses using four different reference genomes of mango (Mangifera indica), including the two haploid assemblies of haplotype-resolved telomere-to-telomere (T2T) genome assembly, a pangenome, and an older version of the reference genome available on NCBI. The choice of reference genome dramatically impacted the mapping efficiency and resulted in notable differences in calling the genetic variants, particularly structural variations (SVs). Phylogenetic analysis was more sensitive to the reference genome compared to genetic differentiation. Population genomic analyses of artificial selection in domestication and SV hotspot regions varied across reference genomes. Notably, the gene enrichment analyses showed significant differences in the top enriched biological processes depending on the reference genome used. Overall, the mango pangenome outperformed the other reference genomes across various metrics, followed by T2T reference genomes, as they captured greater diversity and effectively reduced reference bias. Our findings highlight the role of the mango pangenome in reducing reference bias and underscore the critical role of reference genome selection, suggesting that it is one of the most important factors in population genomic studies.
Fasciclin-like arabinogalactan proteins (FLAs) have been shown to improve plant tolerance to salt stress. However, their role in cold tolerance (CT) remains unclear. Here, we report that banana MaFLA27 positively regulates CT in Arabidopsis. MaFLA27-overexpression (OE) caused the upregulation of differentially expressed arabinogalactan proteins (AGPs) and genes involved in the biosynthesis of cellulose, lignin, and xylan, as well as the degradation of pectin and xyloglucan. Correspondingly, MaFLA27-OE plants exhibited increased cell wall thickness, enhanced cellulose lignin and starch granule content, elevated levels of partially homogalacturonans recognized by JIM5 and JIM7 antibodies, xyloglucan components recognized by CCRC-M39/104 and LM15 antibodies, LM14 antibody binding AGPs. In contrast, transgenic plants showed a decreased degree of pectin methyl-esterification and accumulated less reactive oxygen species after cold acclimation when compared to wild-type plants. A higher number of pectin methylesterases and cellulose and xylan biosynthesis genes were elevated after cold acclimation. Additionally, both Arabidopsis mutant cesa8 and cellulose inhibitor-treated plants displayed decreased freezing tolerance. Our data suggested that MaFLA27-OE in Arabidopsis may perceive and transmit low-temperature stress signals to the cellulose synthase complexes, activating cellulose synthesis and enhancing cold tolerance. These findings reveal a previously unreported cold-tolerance function of FLAs and highlight associated cell wall-mediated tolerance mechanisms.
Comparisons of complete genome assemblies offer a direct procedure for characterizing all genetic differences among them. However, existing tools are often limited to specific aligners or optimized for specific organisms, narrowing their applicability, particularly for large and repetitive plant genomes. Here, we introduce SVGAP, a pipeline for structural variant (SV) discovery, genotyping, and annotation from high-quality genome assemblies at the population level. Through extensive benchmarks using simulated SV datasets at individual, population, and phylogenetic contexts, we demonstrate that SVGAP performs favorably relative to existing tools in SV discovery. Additionally, SVGAP is one of the few tools to address the challenge of genotyping SVs within large assembled genome samples, and it generates fully genotyped VCF files. Applying SVGAP to 26 maize genomes revealed hidden genomic diversity in centromeres, driven by abundant insertions of centromere-specific LTR-retrotransposons. The output of SVGAP is well-suited for pan-genome construction and facilitates the interpretation of previously unexplored genomic regions.
Background/Objectives: CRISPR-Cas9 (Clustered Regularly Interspaced Short Palindromic Repeats)-associated protein 9 is now widely used in agriculture and medicine. Off-target effects can lead to unexpected results that may be harmful, and these effects are a common concern in both research and therapeutic applications. Methods: In this study, using pineapple as the gene-editing material, eighteen target sequences with varying numbers of PAM (Protospacer-Adjacent Motif) sites were used to construct gRNA vectors. Fifty mutant lines were generated for each target sequence, and the off-target rates were counted. Results: Selecting sequences with multiple flanking PAM sites as editing targets resulted in a lower off-target rate compared to those with a single PAM site. Target sequences with two 5 '-NGG ("N" represents any nucleobase, followed by two guanine "G") PAM sites at the 3 ' end exhibited greater specificity and a higher probability of binding with the Cas9 protein than those only with one 5 '-NGG PAM site at the 3 ' end. Conversely, although the target sequence with a 5 '-NAG PAM site (where "N" is any nucleobase, followed by adenine "A" and guanine "G") adjacent and upstream of an NGG PAM site had a lower off-target rate compared to sequences with only an NGG PAM site, their off-target rates were still higher than those of sequences with two adjacent 5 '-NAG PAM sites. Among the target sequences of pineapple mutant lines (AcACS1, AcOT5, AcCSPE6, AcPKG11A), more deletions than insertions were found. Conclusions: We found that target sequences with multiple flanking PAM sites are more likely to bind with the Cas9 protein and induce mutations. Selecting sequences with multiple flanking PAM sites as editing targets can reduce the off-target effects of the Cas9 enzyme in pineapple. These findings provide a foundation for improving off-target prediction and engineering CRISPR-Cas9 complexes for gene editing.
DREB/ERF transcription factors play pivotal roles in plant development; however, their structural characteristics, DNA-binding preferences, and functional roles in highly heterozygous woody plants remain insufficiently understood. Using lychee (Litchi chinensis) as a model, we identified 95 DREB/ERF genes subdivided into ten phylogenetic groups. DNA affinity purification sequencing (DAP-seq) of 45 representative members uncovered 65 194 binding sites with subfamily-specific motifs: C(G/A)CCG(A/C)C for DREB and CGCCG(C/T)C for ERF subfamilies. Each group exhibited unique binding motif preferences, aligning with their protein structures and essential peptide positions. Notably, LITCHI017494 directly regulated terpenoid biosynthesis and aroma formation by activating tandemly repeated LcTPS genes. Furthermore, single nucleotide polymorphisms (SNPs) in LITCHI017494's binding sites altered the binding efficiency of two flowering-related genes (LcSVP and LcVOZ) in early- and late-maturing haplotypes, revealing a mechanism underlying flowering and fruit maturation period. Overall, with experimental evidence, this study provides a comprehensive binding profile of the DREB/ERF family in lychee, revealing intricate transcriptional regulatory networks and serving as a crucial resource for transcription factor research within complex genomic contexts, especially in the DREB/ERF gene family.