Land plants underpin civilization and planetary health, yet their genomic diversity remains largely uncharted. Current resources are unstandardized and scarce, lacking reference genomes for 95% of genera, 70% of families, and 51% of orders, impeding evolutionary and functional insight. We thus propose the PLANeT initiative, an international effort to generate high-quality, standardized genomes across the plant tree of life. Integrating artificial intelligence (AI) with genomics, we will decode conserved principles to advance fundamental plant biology, biodiversity conservation, crop improvement, and natural product discovery. Engaging around 100 labs to train 1,000 scientists, we will tackle pivotal questions for a sustainable future.
Abstract Summary Pangenome graphs capture extensive genetic diversity but introduce analytical challenges due to the redundant representation of structural variations (SVs). While existing tools effectively address cross-sample redundancy or cross-locus redundancy, none specifically target the intra-locus allelic redundancy inherent to pangenome graphs. Here, we present PanSVmerger, an open-source tool designed to consolidate redundant multiallelic SVs within individual loci using three complementary clustering strategies: adaptive k-mer-based Jaccard distance, global alignment distance via VSEARCH, and length distribution. Validation on HPRC pangenome data demonstrates that PanSVmerger effectively reduces multiallelic complexity (e.g., AC ≥ 3 loci from 62.4% to 4.7% using Strategy A) with a modest trade-off: Recall decreased from 97.13% to 93.58%, while precision improved from 94.95% to 96.56%, yielding an overall F1-score of 95.05%. These results demonstrate that PanSVmerger effectively consolidates redundant allele representations with only a minimal loss of sensitivity, making it well-suited for downstream applications that require clean, non-redundant variants. Availability and implementation PanSVmerger is implemented in Python 3.8+ and freely available under the MIT license at GitHub: https://github.com/tingting100/PanSVmerger . The software requires vcflib, bcftools, and optionally VSEARCH. Comprehensive documentation and tutorials are provided.
In gymnosperms such as Ginkgo biloba, the regulatory role of large introns remains unclear. To address this, we conducted integrative multi-omics analyses of chromatin accessibility, three-dimensional chromosomal architecture, histone modifications, DNA methylation, RNA polymerase II occupancy, and gene expression in Ginkgo, complemented by laboratory experiments. We first identified a plant-specific factor, plant-Loop Factor (p-LoopF), which shares ~40% sequence similarity with human CCCTC-binding factor (CTCF). p-LoopF binds specifically to the consensus motif recognized by human CTCF and is significantly associated with chromatin looping, but it lacks canonical CTCF functional features, including motif orientation dependence, chromatin insulation activity, and statistically significant correlations with cohesin subunits. Through integrative multi-omics analyses, we propose a regulatory hypothesis in which p-LoopF-associated chromatin loops are correlated with the recruitment of distal enhancer-like regulatory regions, with large introns serving as a key regulatory context for these interactions. p-LoopF also localizes to promoters and distal intergenic regions, correlating with transcriptional regulation and local chromatin organization. We characterized large introns as regions enriched for chromatin loops, p-LoopF binding sites, and enhancer-like elements, which are strongly associated with the transcriptional regulation of their host genes. Additionally, active histone marks and DNA demethylation were enriched near the boundaries of large introns, particularly around splice sites, suggesting that splicing regulation differs between large and small introns.
Epigenetic inheritance is fundamental to human development and disease, yet the mechanisms governing the transmission of DNA methylation across generations remain incompletely understood. In this study, we performed haplotype-resolved, whole-genome DNA methylation profiling in a healthy three-generation Chinese family, leveraging high-depth Oxford Nanopore Technologies (ONT) and PacBio HiFi long-read sequencing, anchored to a proband-specific telomere-to-telomere (T2T) genome assembly. We observed globally conserved bimodal methylation landscapes across all individuals and generations. Stratified analyses revealed clear functional compartmentalization of methylation marks, characterized by distinct hypomethylation in centromeres and hypermethylation in retrotransposons and repetitive elements. Chromosome-resolved analysis of ribosomal DNA (rDNA) arrays demonstrated a domain-specific methylation pattern with hypomethylation in the transcriptional core and hypermethylation in the intergenic spacer, with evidence for age-associated epigenetic drift in the transcriptional core domain. Through de novo identification and validation, we mapped 23 high-confidence imprinting control regions (ICRs) showing robust parent-of-origin-specific methylation, all overlapping known imprinted genes and enriched for regulatory element signatures. Haplotype-resolved X chromosome analysis further uncovered sex- and allele-specific methylation patterns linked to X inactivation dynamics. Together, this pedigree-scale, high-resolution study delineates the landscape and principles of intergenerational DNA methylation inheritance, revealing both conserved and dynamic features shaping the human epigenome.
Abstract Engineering autoluminescent plants, especially horticultural crops, has recently emerged as a promising research area, with one current approach involving the transgenic introduction of fungal bioluminescence pathway genes. Although autoluminescent plants such as tobacco and petunia have been created, it remains unclear whether all horticultural plants have the potential to be autoluminescent after genetic modification, especially those having flowers with dark colors and thick cuticles. To understand whether, and what if any, floral traits affect autoluminescence potential, we assess the autoluminescence characteristics of 25 representative Phalaenopsis cultivars after transient transformation of fungal bioluminescence pathway genes, alongside several key morphological and biochemical traits. Our results demonstrate that autoluminescence characteristics are correlated with floral color lightness, organ textures and epidermal cell types. In contrast, the content of the substrates of luciferin—caffeic acid and tyrosine, and the infiltration ease of inoculation solution into floral organs after injection, have limited effects on autoluminescence characteristics. Autoluminescence intensity can be reasonably predicted using five floral traits investigated, as 80.4% of variation can be explained by these traits. Our study not only identifies specific Phalaenopsis cultivars with high potential for developing autoluminescent lines but also provides a selection framework applicable to other horticultural crops.
Three-dimensional point cloud (3DPC) data capture detailed geometric and structural plant traits beyond the capability of 2D imaging. When combined with artificial intelligence (AI), it offers a powerful, non-invasive tool for plant phenotyping, which is crucial for driving advancements in plant breeding and agriculture. However, challenges related to data complexity, limited datasets, and model generalization hinder 3DPC’s widespread adoption. To provide a comprehensive overview and guide future research in this area, we conducted a systematic literature review (SLR) following Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) guidelines by analysing 381 papers published between January 2017 and October 2025 from major databases. Our review examines the advantages, current status, limitations, and future directions of AI applications in 3DPC-based plant phenotyping. Our findings indicate a rapid increase in publications since 2022, with deep learning (DL) methods, especially pointwise MLP-based networks, driving much of this growth, with a notable recent surge in Transformer-based, Graph-based, and particularly Hybrid models that combine their strengths. Furthermore, novel methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) are emerging as powerful tools for 3D reconstruction and scene synthesis. Time-of-Flight (ToF) and Structure from Motion and Multi-View Stereo (SfM-MVS) technologies remain the predominant 3DPC data acquisition techniques. Research in this area focuses on trees/shrubs and cereals, typically involving single-species studies. Although the overall use of public datasets remains low (18.9%), their adoption has significantly increased since 2020. Key limitations identified include: (1) a lack of standardized data collection and formats, (2) insufficient model robustness and generalization, especially from lab to field, (3) high computational demands, and (4) a reliance on species-specific models. The future of AI-driven 3DPC phenotyping hinges on overcoming these bottlenecks. Priority should be given to: developing field-deployable, computationally efficient models; exploring the potential of the foundation model; establishing diverse and standardized public datasets; and strengthening the integration of 3D phenomics with genomics to bridge the genotype-to-phenotype gap. This review provides a foundational roadmap to guide research in plant phenomics, crop breeding, and plant science.
The gliding patagium represents a key adaptation for mammalian flight, but its cellular development remains unexplored. Using single-nucleus RNA sequencing of embryonic flying squirrel patagium and dorsal skin, we construct a single-cell atlas of patagium development and identify two distinct fibroblast subpopulations (Fp2 and Fr) highly enriched in the patagium. These fibroblasts are characterized by the patagium upregulation of Wnt5a, Fgf7, and Fgf10, and are associated with patagium morphogenesis through dermal-epidermal putative communication interactions between dermal fibroblasts (Fp2 and Fr) and epithelial basal keratinocytes. Specifically, Fp2 fibroblasts are potentially involved in distal dermal condensation and epithelial thickening together with elevated Wnt5a expression, while both Fp2 and Fr fibroblasts could play a role in epithelial polarization and thickening through Fgf7 and Fgf10, as suggested by ex vivo assays. Our data suggest that gliding patagium development results from the co-option of conserved WNT and FGF signaling pathways within a specialized fibroblast-epithelial context, illustrating how modifications of conserved developmental programs give rise to derived morphological traits.
Nanopore direct RNA sequencing (DRS) offers distinct advantages for transcriptome analysis over the traditional high-throughput RNA sequencing methods by preserving native RNA modifications, eliminating polymerase chain reaction bias, and simplifying the workflow. However, its high basecalling error rate remains a significant hurdle. Here we introduce Coral, a dual context-aware nanopore DRS basecaller that uses a Transformer-based encoder-decoder architecture to capture contextual dependencies at both the signal and sequence levels, substantially improving accuracy. Coral achieves up to a 6.17% improvement in accuracy on human RNA samples compared to Oxford Nanopore Technologies' Dorado basecaller. This improved accuracy enables the detection of 26% more annotated transcript isoforms. Coral also enhances the downstream haplotype phasing, reducing switch errors by up to 78.8% and Hamming errors by 76%, while phasing 36% more single nucleotide polymorphisms.
Posture is a critical phenotypic trait that reflects crop growth and serves as an essential indicator for both agricultural production and scientific research. Accurate pose estimation enables real-time tracking of crop growth processes, but in field environments, challenges such as variable backgrounds, dense planting, occlusions, and morphological changes hinder precise posture analysis. To address these challenges, we propose PFLO (Pose Estimation Model of Field Maize Based on YOLO Architecture), an end-to-end model for maize pose estimation, coupled with a novel data processing method to generate bounding boxes and pose skeleton data from a"keypoint-line"annotated phenotypic database which could mitigate the effects of uneven manual annotations and biases. PFLO also incorporates advanced architectural enhancements to optimize feature extraction and selection, enabling robust performance in complex conditions such as dense arrangements and severe occlusions. On a fivefold validation set of 1,862 images, PFLO achieved 72.2% pose estimation mean average precision (mAP50) and 91.6% object detection mean average precision (mAP50), outperforming current state-of-the-art models. The model demonstrates improved detection of occluded, edge, and small targets, accurately reconstructing skeletal poses of maize crops. PFLO provides a powerful tool for real-time phenotypic analysis, advancing automated crop monitoring in precision agriculture.
After 500 million years of evolution, extant land plants compose the following two sister groups: the bryophytes and the vascular plants. Despite their small size and simple structure, bryophytes thrive in a wide variety of habitats, including extreme conditions. However, the genetic basis for their ecological adaptability and long-term survival is not well understood. A comprehensive super-pangenome analysis, incorporating 123 newly sequenced bryophyte genomes, reveals that bryophytes possess a substantially greater diversity of gene families than vascular plants. This includes a higher number of unique and lineage-specific gene families, originating from extensive new gene formation and continuous horizontal transfer of microbial genes over their long evolutionary history. The evolution of bryophytes' rich and diverse genetic toolkit, which includes new physiological innovations like unique immune receptors, likely facilitated their spread across different biomes. These newly sequenced bryophyte genomes offer a valuable resource for exploring alternative evolutionary strategies for terrestrial success.
The COP9 signalosome (CSN) is a highly conserved protein complex in eukaryotes, with CSN5 serving as its critical catalytic subunit. However, the role of CSN5 in plant immunity is largely unexplored. Here, we found that suppression of OsCSN5 in rice enhances resistance against the fungal pathogen Magnaporthe oryzae and the bacterial pathogen Xanthomonas oryzae pv. oryzae ( Xoo ) without affecting growth. OsCSN5 is ubiquitinated and degraded by the E3 ligase OsPUB45. Overexpression of OsPUB45 increased resistance against M. oryzae and Xoo , while dysfunction of OsPUB45 decreased resistance. In addition, OsCSN5 stabilized OsCUL3a to promote the degradation of a positive regulator OsNPR1. Overexpression of OsPUB45 compromised accumulation of OsCUL3a, leading to stabilization of OsNPR1, whereas mutations in OsPUB45 destabilized OsNPR1. These findings suggest that OsCSN5 stabilizes OsCUL3a to facilitate the degradation of OsNPR1, preventing its constitutive activation without infection. Conversely, OsPUB45 promotes the degradation of OsCSN5, contributing to immunity activation upon pathogen infection.
Bioinformatics analysis often requires the filtering of multi-datasets, based on frequency or frequency of occurrence, for decisions on retention or deletion. Existing tools for this purpose often present a challenge with complex installation, which necessitate custom coding, thereby impeding efficient data processing activities. To address this issue, Filterx, a user-friendly command line tool that written in C language, was developed that supports multi-condition filtering, based on frequency or occurrence. This tool enables users to complete the data processing tasks through a simple command line, greatly reducing both workload and data processing time. In addition, future development of this tool could facilitate its integration into various bioinformatics data analysis pipelines.
Plant phenomics has become one of the most significant scientific fields in recent years. However, typical phenotyping procedures have low accuracy, low throughput, and are labor-intensive and time-consuming. Large-scale phenotypic collection equipment, on the other hand, is pricy, rigid, and inconvenient. The advancement of phenomics has been hampered by these restrictions. Lightweight picture collection equipment can now be used to capture plant phenotypic data thanks to the development of deep learning-based image identification. For the purpose of training the model, this approach needs high-quality annotated datasets. In this study, we used a handheld camera to gather multi-angle, multi-time series images and an unmanned aerial vehicle (UAV) to create a maize image phenotyping database (MIPDB). Over 30,000 high-resolution photos are available in the MIPDB, with 17,631 of those images having been carefully tagged with point-line method. The MIPDB can be accessed by the general public at <http://phenomics.agis.org.cn>. We anticipate that the availability of this superior dataset will stimulate a new revolution in crop breeding and advance deep learning-based phenomics research. ### Competing Interest Statement The authors have declared no competing interest.
Recent advancements in genome assembly have greatly improved the prospects for comprehensive annotation of Transposable Elements (TEs). However, existing methods for TE annotation using genome assemblies suffer from limited accuracy and robustness, requiring extensive manual editing. In addition, the currently available gold-standard TE databases are not comprehensive, even for extensively studied species, highlighting the critical need for an automated TE detection method to supplement existing repositories. In this study, we introduce HiTE, a fast and accurate dynamic boundary adjustment approach designed to detect full-length TEs. The experimental results demonstrate that HiTE outperforms RepeatModeler2, the state-of-the-art tool, across various species. Furthermore, HiTE has identified numerous novel transposons with well-defined structures containing protein-coding domains, some of which are directly inserted within crucial genes, leading to direct alterations in gene expression. A Nextflow version of HiTE is also available, with enhanced parallelism, reproducibility, and portability.
Long reads that cover more variants per read raise opportunities for accurate haplotype construction, whereas the genotype errors of single nucleotide polymorphisms pose great computational challenges for haplotyping tools. Here we introduce KSNP, an efficient haplotype construction tool based on the de Bruijn graph (DBG). KSNP leverages the ability of DBG in handling high-throughput erroneous reads to tackle the challenges. Compared to other notable tools in this field, KSNP achieves at least 5-fold speedup while producing comparable haplotype results. The time required for assembling human haplotypes is reduced to nearly the data-in time.
MOTIVATION:Seeding is a rate-limiting stage in sequence alignment for next-generation sequencing reads. The existing optimization algorithms typically utilize hardware and machine-learning techniques to accelerate seeding. However, an efficient solution provided by professional next-generation sequencing compressors has been largely overlooked by far. In addition to achieving remarkable compression ratios by reordering reads, these compressors provide valuable insights for downstream alignment that reveal the repetitive computations accounting for more than 50% of seeding procedure in commonly used short read aligner BWA-MEM at typical sequencing coverage. Nevertheless, the exploited redundancy information is not fully realized or utilized. RESULTS:In this study, we present a compressive seeding algorithm, named CompSeed, to fill the gap. CompSeed, in collaboration with the existing reordering-based compression tools, finishes the BWA-MEM seeding process in about half the time by caching all intermediate seeding results in compact trie structures to directly answer repetitive inquiries that frequently cause random memory accesses. Furthermore, CompSeed demonstrates better performance as sequencing coverage increases, as it focuses solely on the small informative portion of sequencing reads after compression. The innovative strategy highlights the promising potential of integrating sequence compression and alignment to tackle the ever-growing volume of sequencing data. AVAILABILITY AND IMPLEMENTATION:CompSeed is available at https://github.com/i-xiaohu/CompSeed.
MicroRNAs (miRNAs) and short RNA fragments (18-25 nt) are crucial biomarkers in biological research and disease diagnostics. However, their accurate and rapid detection remains a challenge, largely due to their low abundance, short length, and sequence similarities. In this study, we report on a highly sensitive, one-step RNA O-circle amplification (ROA) assay for rapid and accurate miRNA detection. The ROA assay commences with the hybridization of a circular probe with the test RNA, followed by a linear rolling circle amplification (RCA) using dUTP. This amplification process is facilitated by U-nick reactions, which lead to an exponential amplification for readout. Under optimized conditions, assays can be completed within an hour, producing an amplification yield up to the microgram level, with a detection limit as low as 0.15 fmol (6 pM). Notably, the ROA assay requires only one step, and the results can be easily read visually, making it user-friendly. This ROA assay has proven effective in detecting various miRNAs and phage ssRNA. Overall, the ROA assay offers a user-friendly, rapid, and accurate solution for miRNA detection.
Increasing the accuracy of the nucleotide sequence alignment is an essential issue in genomics research. Although classic dynamic programming (DP) algorithms (e.g., Smith-Waterman and Needleman-Wunsch) guarantee to produce the optimal result, their time complexity hinders the application of large-scale sequence alignment. Many optimization efforts that aim to accelerate the alignment process generally come from three perspectives: redesigning data structures [e.g., diagonal or striped Single Instruction Multiple Data (SIMD) implementations], increasing the number of parallelisms in SIMD operations (e.g., difference recurrence relation), or reducing search space (e.g., banded DP). However, no methods combine all these three aspects to build an ultra-fast algorithm. In this study, we developed a Banded Striped Aligner (BSAlign) library that delivers accurate alignment results at an ultra-fast speed by knitting a series of novel methods together to take advantage of all of the aforementioned three perspectives with highlights such as active F-loop in striped vectorization and striped move in banded DP. We applied our new acceleration design on both regular and edit distance pairwise alignment. BSAlign achieved 2-fold speed-up than other SIMD-based implementations for regular pairwise alignment, and 1.5-fold to 4-fold speed-up in edit distance-based implementations for long reads. BSAlign is implemented in C programing language and is available at https://github.com/ruanjue/bsalign.
ABSTRACT Error-correcting codes (ECCs) employed in the state-of-the-art DNA digital storage (DDS) systems suffer from a trade-off between error-correcting capability and the proportion of redundancy. To address this issue, in this study, we introduce soft-decision decoding approach into DDS by proposing a DNA-specific error prediction model and a series of novel strategies. We demonstrate the effectiveness of our approach through a proof-of-concept DDS system based on Reed-Solomon (RS) code, named as Derrick. Derrick shows significant improvement in error-correcting capability without involving additional redundancy in both in vitro and in silico experiments, using various sequencing technologies such as Illumina, PacBio and Oxford Nanopore Technology (ONT). Notably, in vitro experiments using ONT sequencing at a depth of 7× reveal that Derrick, compared with the traditional hard-decision decoding strategy, doubles the error-correcting capability of RS code, decreases the proportion of matrices with decoding-failure by 229-fold, and amplifies the potential maximum storage volume by impressive 32 388-fold. Also, Derrick surpasses ‘state-of-the-art’ DDS systems by comprehensively considering the information density and the minimum sequencing depth required for complete information recovery. Crucially, the soft-decision decoding strategy and key steps of Derrick are generalizable to other ECCs’ decoding algorithms.