
Premise:Accessioning herbarium specimens is labor intensive, yet remains vital for research in ecology, evolution, and conservation. As institutional support for herbaria declines, efficient tools are needed to streamline this process. The R package BarnebyLives was developed to assist collectors by supplementing collection notes, verifying taxonomic data, conducting quality checks, generating labels, and submitting digital records. Methods and Results:BarnebyLives integrates geospatial data from U.S. government sources to provide jurisdictional and site information and checks taxonomic names using in-house spell checkers, International Plant Names Index (IPNI) author standards, and Kew's Plants of the World Online. Optional features include generating Google Maps driving directions. The tool outputs data in tabular and spatial formats for review before producing LaTeX-based labels and shipping manifests. Conclusions:BarnebyLives improves data accuracy, ensures up-to-date taxonomy, and significantly reduces the time and effort required to accession herbarium specimens in the United States.
Premise:In North America, Phragmites australis (common reed) has drawn a great deal of research attention. Non-native P. australis subsp. australis is a noxious weed that has locally displaced native P. australis subsp. americanus in some areas. Although morphological features can distinguish the two subspecies, molecular tools often are required to confirm identifications. Additionally, the existence of natural intrasubspecific hybrids presents novel management challenges. Hybrid Phragmites is difficult to detect, and it has become standard practice to apply molecular tools to survey for hybrids. Methods:We applied several molecular techniques-microsatellite, DArTseq (a type of genotyping-by-sequencing), restriction fragment length polymorphism (PCR-RFLP), and next-generation sequencing-to characterize P. australis at the landscape scale in Minnesota and Wisconsin and to search for hybrids. Results:We obtained molecular data for Phragmites plants sampled from 341 stands, ultimately characterizing 98 stands as native and 236 as non-native. Plants from two adjacent stands in Washington County, Minnesota, were confirmed to be hybrids. Discussion:These are the first confirmed hybrids from the Upper Midwest/western Great Lakes region. We also discuss the relative cost and effectiveness of the various molecular methods and offer recommendations for future studies.
Premise Modern plant breeding requires robust analysis of complex, multi-environment datasets to identify superior genotypes. Although R offers powerful statistical packages, their use often requires programming expertise. Existing graphical user interface (GUI)-based tools are either limited in scope or not openly accessible, leaving breeders without a comprehensive solution.Methods and Results We introduce PbAT (Plant breeding Analytical Tools), an open-access Shiny application that unifies multiple R packages into a seamless, code-free workflow. PbAT enables trial design, data curation, experimental design analysis, stability testing, multivariate methods, and mating design analysis. Outputs include variance components, best linear unbiased estimators (BLUEs), best linear unbiased predictors (BLUPs), ANOVA tables, and publication-ready plots, accompanied by automated plain-language interpretations and model equations. Its modular architecture supports future integration of additional tools. PbAT is available online as a web application (https://pbat.online) and as an R package (https://github.com/abhijithkpgen/PbAT).Conclusions By integrating powerful statistical methods with intuitive interfaces, PbAT streamlines plant breeding analytics, enhances reproducibility, and accelerates data-driven decision-making in crop improvement.
Abstract Premise The growing demand for wildflower seeds in ecological restoration requires reliable species identification, yet current market products often contain heterogeneous species. As seed identification is labor‐intensive and requires advanced botanical knowledge, we evaluated multiple segmentation and classification approaches to determine which combination performed better for identifying species and quantifying their frequencies under real‐life constraints. Methods We generated seed images from common wildflower species using a flatbed scanner. Artificial intelligence (AI) and non‐AI segmentation pipelines were compared. We evaluated over 20 classifiers on tabulated features and convolutional neural networks (CNNs) on single seed image inputs. Open‐set conditions were simulated using “mock seed mixes” with previously unseen taxa using decision thresholds. Results With tabulated data, the XGBoost and Multi‐layer Perceptron (MLP) models exceeded F1 > 0.96, while AutoGluon's CNN and ensemble models reached F1 > 0.97. ResNet‐50 achieved an F1 score over 0.99 on single‐seed images obtained with Cellpose segmentation. Random forest achieved the highest accuracy (0.92) with unseen species in open‐set classification. Discussion Deep learning segmentation substantially enhanced CNN accuracy, yet tabulated‐feature machine learning models remained competitive and efficient. Threshold‐based rejection enabled robust open‐set classification, with random forest outperforming deeper models in rejecting unseen seeds. No single method was universally optimal; the best strategy depends on whether the task involves closed‐set or open‐set classification and on trade‐offs between accuracy and computational cost.
Abstract Premise Accurate species identification is crucial for ecological restoration and can be especially challenging for understudied non‐model species. Quercus garryana is the only native oak species in the Pacific Northwest and is an important component of the endangered oak savanna ecosystem. Quercus robur is an imported ornamental species from Europe and has been found to be mistakenly planted as Q. garryana in habitat restoration projects. Methods We measured leaf morphological traits sampled from herbarium collections in their native ranges using the digital morphometric tools MorphoLeaf and Tomato Analyzer. We then used Lasso logistic analysis to generate a predictive model and tested it on leaves from Portland, Oregon. To streamline this species detection process, we developed Garryanalyzer, an ImageJ plug‐in that automatically measures leaf traits and outputs species predictions. Results Garryanalyzer demonstrated 95% accuracy in predicting the species identity of herbarium specimens of oaks. Garryanalyzer correctly identified all Q. robur individuals sampled in Portland but showed lower accuracy for Q. garryana . Discussion Many existing morphometric software are not open source, which makes them unable to be customized to specific study systems. Garryanalyzer is built upon the widely used open‐source ImageJ platform. This study also demonstrates a viable workflow for developing similar tools for other ecologically important non‐model plant species.
Premise:DNA barcoding for timber species identification requires comprehensive reference datasets, informative DNA barcodes, and cost-effective protocols. We developed a workflow leveraging Hyb-Seq (target capture sequencing and genome skimming) to address these challenges, and we tested it on four genera from the mahogany family (Meliaceae). Methods:We sequenced up to 350 nuclear and 177 plastid loci from 132 herbarium specimens representing leaf samples of 22 species. We determined the DNA barcoding potential of each locus by looking at species recovery and monophyly in gene trees. We then selected 13 short regions (candidate barcodes) within high-potential loci and tested their PCR amplification and Sanger sequencing on wood DNA. Results:Three candidate barcodes emerged as the most reliably sequenced from wood DNA and as providing the most accurate species-level identifications, with species monophyly rates above 80%. Failure to obtain sequences from some wood DNA extracts was more often associated with potential DNA impurity (as inferred from DNA color) than with DNA degradation. Discussion:Our reference data and candidate barcodes provide a foundation to support the DNA barcoding of mahogany and its relatives. Our workflow illustrates how the wealth of Hyb-Seq data currently generated from global herbaria may be leveraged to monitor plant diversity.
Abstract Premise Conflicting phylogenetic signals are common in plant phylogenomics and often reflect evolutionary histories shaped by processes like hybridization, incomplete lineage sorting, and whole‐genome duplication (WGD). We aimed to identify and assess these complex processes in the hyper‐diverse family Asteraceae to offer insight into the underlying causes of phylogenetic discordance. Methods We used new and existing Hyb‐Seq and transcriptome data to explore phylogenetic discordance by testing for nuclear/plastid incongruences, WGD, and reticulation. We present a tutorial detailing the execution of complex bioinformatic analyses to increase transparency, facilitate reproducibility, and support advancements in the field of plant evolution (https://github.com/erika-r-moore/Ellestad_etal_2025_APPS_Hybridizations). Results We uncovered extensive discordance among nuclear gene trees and deep reticulation events, particularly among South American lineages. Signals of WGD were found across the family but were often difficult to interpret, likely due to variation in data completeness, the complexity of the events, and their ancient origins. Discussion Our study and tutorial, along with a growing body of phylogenomic research, emphasize the role of reticulation and WGD in the evolution of large, diverse clades, while also underscoring the challenges. We anticipate continued advancements in theoretical approaches that will further enhance empirical studies in reticulate evolution.
Premise:Advances in long-read sequencing offer new possibilities to investigate haplotype diversity across multiple genes in plants and other taxa through multi-locus, long-read amplicon sequencing (multi-locus LRAS). Despite this progress, there is a notable absence of dedicated bioinformatics pipelines for assembling diploid haplotypes of heterozygous individuals from such multi-locus LRAS datasets, which is required for highly polymorphic populations. Methods:We first evaluated various de novo and reference-based assembly methods, culminating in a custom pipeline (HapAsmbl) to assemble haplotypes from Oxford Nanopore Technologies (ONT) LRAS data of five flowering genes (FT3, FTL9, VRN1, VRN2A, and VRN2B) generated from perennial ryegrass, a highly heterozygous species. After verifying the efficacy using a simulated heterozygous dataset, the HapAsmbl pipeline was used to explore haplotype diversity of CO, FT3, and VRN1 across multiple ryegrass populations. Results:HapAsmbl outperformed existing tools by reliably reconstructing diploid haplotypes across multiple loci, enabling efficient haplotype characterization and novel allele discovery in genetically diverse populations. Discussion:HapAsmbl simplifies haplotype resolution from complex LRAS datasets from heterozygous individuals, allowing routine use of ONT long-read sequencing for scalable haplotype analysis. HapAsmbl will enable researchers to uncover novel alleles and relate these to phenotype, supporting plant-breeding efforts in non-model crops.
Premise:The recovery of non-target organism reads, especially when whole organisms are sampled, constitutes a great opportunity for studying microbial communities. The increase in whole genome sequencing feasibility and the development of new marker-based pipelines enable the use of short reads to study bacterial communities associated with organisms. Methods:We utilized population genomic data of the liverwort Calasterella californica obtained through the California Conservation Genomics Project to characterize the composition of its associated bacterial communities and explore its variation across the geographic space. Results:The bacterial communities associated with C. californica were dominated by the methanotroph Methylobacterium and other Hyphomicrobiales, a group that includes well-known plant symbionts. While diversity metrics of bacteria composition were similar across localities, we found significant differences in the relative abundance of a few taxa across California regions, likely driven by differences in precipitation and temperature seasonality. Discussion:Our results support previous observations that liverwort bacterial communities are not randomly assembled, suggesting a potential role of the plant in determining community composition, an emerging pattern that deserves more attention. The novel off-target metagenomics approach can be applied to any population-level resequencing where whole organisms are sequenced, opening the door to exciting avenues of microbiome research using repurposed data from landscape genomics.
Premise:Detecting historical introgression among populations or species from genomic data is a common goal in evolutionary genetics. Most current methods fall into two major categories: network inference and admixture inference. Network inference (e.g., SNaQ) is computationally challenging and typically requires first reducing large genomic datasets into a less informative collection of inferred gene trees. In contrast, admixture inference (e.g., ABBA-BABA tests) can accommodate enormous single-nucleotide polymorphism (SNP) datasets but is restricted to examining subsets of four to five samples at a time. Here, we demonstrate a new approach to evaluate SNP frequencies among quartet samples under a phylogenetic hypothesis (similar to ABBA-BABA tests), while examining all quartet information simultaneously (similar to the network inference methods). Methods and Results:To do this, our method simcat trains a neural network machine learning model on coalescent simulations to discriminate between introgression scenarios based on learned SNP frequency patterns. We demonstrate the accuracy of simcat to classify introgression events from simulations, evaluate its sensitivity to variation in species tree parameters, and demonstrate its application to an empirical dataset of oak trees (Quercus ser. Virentes). Conclusions:Our approach represents a first step towards leveraging machine learning to expand phylogenetic invariants-based methods beyond the scale of quartets to a larger phylogenetic context.
Premise:Phylogenetic trees depict evolutionary relationships among taxa. However, they are strictly bifurcating structures that do not take into account several types of evolutionary events such as horizontal gene transfer, hybridization, or introgression. Although the development of new methods in phylogenetic networks has recently increased, limited visualization software is available to plot the phylogenetic networks. Methods and results:Here, we present the R package tanggle, a visualization package for phylogenetic networks. Our package extends the widely used visualization package ggtree and allows a variety of input data from DNA sequences to extended Newick format; it also builds on the flexibility of ggplot2 to manipulate colors and other plot characteristics. In addition, our package allows for the inclusion of images and mapped morphological and geographical characteristics on the network. Conclusions:In response to growing demands for reproducible, open-source research, tanggle facilitates the production of script-based, publication-quality figures rather than graphics manually created with design software. By embedding figure code and metadata directly within analysis pipelines, tanggle improves transparency, traceability, and version control; enables automated regeneration of figures as data or methods change; and simplifies sharing and reuse of visualizations.
Premise:Hybrid capture sequencing (Hyb-Seq) is a widely used approach in phylogenomics, providing efficient access to targeted genomic regions. However, deriving high-quality phylogenetic trees from raw sequencing reads requires extensive bioinformatics processing, which increases complexity, the risk of errors, and challenges in file management, especially for users unfamiliar with bioinformatics workflows. Methods and Results:We developed HybSuite, a streamlined Bash-based bioinformatics pipeline built upon mainstream tools such as HybPiper 2, designed to simplify the Hyb-Seq phylogenomic analysis from raw reads to species trees. Compared to existing tools (e.g., HybPiper 2, CAPTUS), it offers a modular yet integrated workflow covering all key steps from downloading from the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA), adapter removal, data assembly, and paralog handling to species tree inference and extensive in-depth analysis. We validated HybSuite by reconstructing a robust phylogeny for the Elaeagnaceae family, using the Angiosperms353 probe set and a dataset of 100 single-copy nuclear loci from Arabidopsis. Conclusions:HybSuite provides a flexible and user-friendly pipeline for Hyb-Seq phylogenomic analyses, and its high accuracy and efficiency were demonstrated through benchmarking with two empirical datasets. HybSuite is freely available at https://github.com/Yuxuanliu-HZAU/HybSuite. The pipeline is compatible with both the Linux and MacOS platforms.
Premise There is a knowledge gap regarding how foliar injury and restricted water uptake can be detected by measuring root dielectric response. This pot study nondestructively evaluated the efficiency of real-time dielectric measurement to monitor the effects of glyphosate spraying.Methods Root dielectric properties were recorded on a minute scale in control and glyphosate-treated maize, cucumber, and pea. Chlorophyll, stomatal conductance, and biomass measurements were taken to interpret the dielectric changes.Results Electrical capacitance and conductance varied diurnally due to the circadian regulation of water uptake and hydraulic conductance. Glyphosate application reduced capacitance, indicating the impeded root growth and activity caused by impaired amino acid synthesis, foliar damage, and restricted transpiration. The dissipation factor decreased in response to glyphosate due to impeded apoplastic water flow, suppressed root lignification, and hampered water absorption. The enhanced leaf and root hydraulic resistance caused by glyphosate was manifested in sharply reduced electrical conductance. Changes in the species' dielectric response were consistent with physiological symptoms and biomass loss.Discussion Real-time dielectric measurement proved suitable for the nondestructive monitoring of plant responses to foliar stress through altered root traits. This method could be employed to evaluate herbicide tolerance in crops and to develop and determine dosage of herbicide ingredients.
Abstract Premise Third‐generation sequencing has revolutionized genomics, enabling in‐depth analysis of genome sequence, structure, and epigenetic features. Yet, extracting high‐quality DNA for long‐read sequencing remains a bottleneck—particularly in non‐model plants, such as mature trees growing in natural environments, which often contain abundant endogenous compounds that hinder extraction and downstream applications. Methods and Results We developed an optimized, robust, and cost‐effective DNA extraction protocol that yields high‐quality DNA suitable for Oxford Nanopore Technologies and PacBio sequencing. Validation across diverse taxa—including six Nothofagus species, gymnosperms endemic to Andean–Patagonian forests, exotic conifers of commercial value, and model plants—demonstrated consistently high DNA purity (A 260 /A 280 > 1.8, A 260 /A 230 > 2.0) and fragment sizes ≥30 kbp. Downstream sequencing confirmed suitability for applications requiring long, intact molecules and base modification detection. Conclusions Compared to commercial kits and standard protocols, this approach achieved superior DNA integrity and yield without specialized equipment, offering an accessible solution for researchers working with challenging plant species.
Premise:Applied ecology can significantly influence policy decisions on environmental issues. Therefore, research in this field should be as transparent and reproducible as possible. Existing expertise from a broad range of disciplines should also be integrated into ecological research to allow researchers to maximize understanding of complex systems. Methods:We illustrate how Pearl's causality can contribute to applied ecology. We demonstrate the implications of causal diagrams for assessing the effects of anthropogenic and abiotic factors in ecological systems, using Myricaria germanica in Italian river systems as an example. In particular, we showcase the interplay between explicit causal modeling and classical statistical techniques. Results:In our example, we find that river channel width and riverbank protections are the most important factors affecting the survival of M. germanica juveniles in northern Italy. Other factors such as altitude and other human activity also impact M. germanica survival through channel width as a mediator. Discussion:We demonstrate that causal diagrams can be an effective new language for ecological research. The causal diagram highlights that integrating hydrological, historical, or anthropological data could strengthen understanding of M. germanica populations in river systems. This framework facilitates interdisciplinary research and realizes the full potential of ecological datasets.
Premise:Plants are frequently exposed to combinations of abiotic and biotic stresses that pose a greater threat to yield and productivity than individual stresses. However, knowledge of the impact of many stress combinations in numerous plants is limited due to the lack of experimental data, which could take decades to generate. To overcome this limitation, we utilized existing literature data from various plant species and stress combinations to derive biological inferences, thereby gaining a comprehensive understanding of plant responses through a computational tool. Methods:Public databases were used to gather literature on the impact of various abiotic and biotic stress combinations. Then, a composite artificial neural network (ANN)-based multi-target classification and regression deep learning model was developed using machine learning algorithms. Results:The model predicted the impact of stress interactions in plants, including the morphological parameters affected and percentage changes in those parameters, with an overall accuracy of 76.33%. Predicted reductions in yield were validated in rice under combined drought and heat stress. Discussion:The ANN-based model developed in this study is a valuable resource for plant researchers seeking to understand the impact of stress combinations. The tool can make use of multivariate and complex combined stress datasets.
Abstract Premise Traditional methods to quantify mycelial growth rely on destructive sampling to quantify biomass. Moreover, these approaches limit continuous observation and require a sufficient mass to measure. Recent work examines hyphal network traits by reconstructing the hyphal network from spatial coordinates via images, providing information about branching patterns and spatial growth over time. Methods and Results We developed SkelPy, a Python‐based graphical user interface that skeletonizes images of hyphal networks and extracts biologically relevant structural parameters such as fractal dimension, a proxy for the complexity and branching structure of the hyphal network. Using a high‐throughput pipeline, we imaged three isolates of Botrytis cinerea grown in liquid culture for 72 h, generating a dataset of 180 time‐series images. Conclusions SkelPy enables efficient, non‐destructive, and scalable quantification of hyphal growth and complexity from time‐resolved image datasets, providing a powerful and user‐friendly tool for studying fungal network dynamics.