Earlham Institute (EI, formerly The Genome Analysis Centre (TGAC)) is a life science research institute located at the Norwich Research Park (NRP), Norwich, England. EI's research is focused on exploring living systems by applying computational science and biotechnology to answer ambitious biological questions and generate enabling resources.
A core goal of phylogenomics is to determine the evolutionary history of a set of species from biological sequence data. Phylogenetic networks are able to describe more complex evolutionary phenomena than phylogenetic trees but are more difficult to accurately reconstruct. Recently, there has been growing interest in developing methods to infer semi-directed phylogenetic networks. As computing such networks can be computationally intensive, one approach to building such networks is to puzzle together smaller networks. Thus, it is essential to have robust methods for inferring semi-directed phylogenetic networks on small numbers of taxa. In this paper, we investigate an algebraic method for performing phylogenetic network inference from nucleotide sequence data on 4-leaf semi-directed phylogenetic networks by analyzing the distribution of leaf-pattern probabilities. On simulated data, we found that we can correctly identify with high accuracy the undirected phylogenetic network for sequences of length at least 10 kbp. We found that identifying the semi-directed network is more challenging and requires sequences of length approaching 10 Mbp. We are also able to use our approach to identify treelike evolution and determine the underlying tree. Finally, we employ our method on a real data set from Xiphophorus species and use the results to build a phylogenetic network.
Gramene (gramene.org) is a comprehensive reference database for comparative plant genomics and pathway analysis, integrating functional annotations, evidence-based curated pathways and their projections, and multi-omics datasets. Since our last report, Gramene has added crop-specific pan-genome portals for maize, sorghum, rice, and grapevine. These pan-genome portals host population-scale datasets and multiple assembled genomes per species, all anchored by shared reference genomes. Importantly, these portals now adopt standardized rsIDs for genetic variants, advancing FAIR data principles and enabling cross-database interoperability. The main site is now Gramene Plants, emphasizing its broad genome coverage. Release 69 features 233 reference genomes, curated pathways for 139 species, expression data from 1026 studies across 27 species, and genetic variation data mapped to 27 genomes from 19 species. Key updates to the integrated search functionality include embedded expression viewers from the Bio-Analytic Resource for Plant Biology and EMBL-EBI Expression Atlas, a literature-curated catalog of gene functions, and a new Germplasm tab linking accessions with loss-of-function alleles to seed repositories. These advances reinforce Gramene as a comprehensive platform for exploring plant genomic diversity, gene function, and evolutionary conservation across the Green Tree of Life and within key agricultural species.
Third-generation long-read sequencing technologies, significantly improve metagenome assemblies. Highly accurate PacBio HiFi reads can yield hundreds of near-complete metagenome-assembled genomes (MAGs) from a single sample. Recently, the accuracy of the more cost-effective Oxford Nanopore Technologies (ONT) platform has increased to a per-base error rate of 1-2%. However, current metagenome assemblers are optimized for HiFi and do not scale to the large data sets that ONT enables. We present nanoMDBG, an evolution of metaMDBG, which supports the latest ONT reads through an error correction pre-processing step in minimizer-space. Across a range of ONT datasets, including a large 400 Gbp soil sample, nanoMDBG reconstructs up to twice as many high-quality MAGs as the next best ONT assembler, metaFlye, while requiring a third of the CPU time and memory. Critically, the latest ONT technology can now produce comparable MAG construction results as those obtained using PacBio HiFi at the same sequencing depth.
The diversity of plant inflorescence architecture is specified by gene expression patterns. In wheat (Triticum aestivum), the lanceolate-shaped inflorescence (spike) is defined by rudimentary spikelets at the base, which form as a result of delayed spikelet and floral development compared with central spikelets. While previous studies identified gene expression differences between central and basal inflorescence sections, gene expression patterns along the apical-basal axis remain poorly resolved due to bulk tissue-level techniques. Here, we optimize Multiplexed Error Robust Fluorescence In Situ Hybridization, a spatial transcriptomics technique, in wheat inflorescence tissue, enabling transcript localization for 200 genes to cellular resolution across 4 stages of development. Cell segmentation and clustering of 50,000 cells identified 18 expression domains and their enriched genes, revealing the spatio-temporal organization of spikelet and floral development, and characterizing tissue-level gene markers. Using these domain- and cell-level maps, we characterize expression patterns of genes differentially expressed across the apical-basal axis. We identify distinct, spatially coordinated expression patterns distinguishing axillary meristems and their subtending leaf ridges across the apical-basal axis before visible spikelet formation, highlighting factors patterning meristem identity and transition. To support the broader research community, all raw and processed data are publicly available, including through an interactive WebAtlas interface (www.wheat-spatial.com).
We present a genome assembly from an individual female Tachina fera (Arthropoda; Insecta; Diptera; Tachinidae). The genome sequence is 752 megabases in span. The majority of the assembly (99.98%) is scaffolded into 6 chromosomal pseudomolecules, with the X sex chromosome assembled. The complete mitochondrial genome was also assembled and is 17.4 kilobases in length.