
Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.
Single-cell RNA sequencing (scRNA-seq) enables high-resolution profiling of individual cells. However, inferring temporal relationships among cells remains a challenge. Here, we present Real-Time Course Analysis (RTCA), a simple direct phylogenetic signal framework that reconstructs cell lineages from somatic variants detected across cells using variants derived from nuclear-encoded transcripts in scRNA-seq data. Our simulations demonstrate that only a modest increase in informative variant sites is sufficient to maintain accurate tree reconstruction, even with a tenfold increase in the number of cells. This scalability makes RTCA applicable to datasets generated by recent single-cell sequencing technologies and is also suitable for reanalyzing existing datasets. We also performed a comparative analysis between our method and a representative genotype-mediated inference framework, PhylinSic, which shares certain conceptual similarities with RTCA. The simulation results show that RTCA is more robust to sparse mutation signals and dropout-induced missing data than PhylinSic. In an application to the datasets from two healthy human placental samples, RTCA successfully reconstructed bifurcating phylogenetic trees. By mapping expression-based cell type clusters onto the trees, we evaluated the degree of monophyly within lineages and found our results consistent with known placental differentiation pathways. The identified cell lineages also aligned with classifications based on gene expression and pseudotime analysis. Compared to gene expression-based and pseudotime analysis, RTCA provides a temporally-based model of cell trajectories, integrating lineage and expression information in a biologically meaningful manner. In summary, RTCA offers a scalable, cost-effective solution for reconstructing developmental processes in complex normal tissues.
Advances in mass spectrometry (MS)-based proteomics have enabled the large-scale characterization of posttranslational modifications (PTMs) through affinity-based enrichment. However, this technique introduces a bias towards selectively enrichable modifications, thus leaving oxidative modifications underexplored. Methionine oxidation (methionine sulfoxide) is an important indicator of cellular redox status, but its systematic analysis remains challenging because no enrichment method is available and artifactual oxidation can occur during sample preparation. Here, we developed an enrichment-free proteomic strategy for large-scale detection of methionine oxidation using a deep LC-MS platform. By optimizing acquisition conditions, we identified more than 260k precursors in a single-shot analysis. Under these conditions, methionine oxidation was efficiently detected, whereas many other PTMs remained poorly detected. To improve data reliability, we established a sample preparation workflow that minimized artifactual oxidation. Accordingly, we identified more than 3,500 methionine-oxidized proteins. Integration of methionine oxidation and expression proteomics across subcellular compartments revealed redox patterns under low-serum conditions, including increased mitochondrial oxidation and decreased endoplasmic reticulum oxidation. These changes are associated with metabolic reprogramming and altered antioxidant capacity. Overall, this study established an enrichment-free framework for the proteome-scale methionine oxidation analysis, and demonstrated that integrating oxidation and expression data enables the spatially resolved interpretation of cellular redox states.
The European anchovy is a keystone forage species in the Mediterranean and Northeast Atlantic, yet a chromosome-scale reference genome has been lacking. Here, we present a 1.4 Gb assembly anchored to 24 chromosomes (scaffold N50 = 56.4 Mb; BUSCO 96.6%) with 27,663 annotated protein-coding genes and 44.6% repetitive content dominated by DNA transposons, which likely account for the near doubling of genome size relative to Atlantic herring. Comparative analyses across 11,755 orthogroups revealed a net contractive trend in Clupeidae, superimposed on conspicuous expansions of antiviral and innate-immune gene families (TRIM16, GTPase IMAP member 8, IFI44, Toll-like receptors), with GO enrichment exceeding 400-fold for antiviral regulatory processes. Branch-site analyses identified 216 positively selected genes in E. encrasicolus, 81 of which map to innate signaling, epigenetic regulation, ubiquitin-mediated immunity, and selective autophagy. Multitissue RNA-Seq further demonstrated that expanded, positively selected, and epithelial barrier-associated genes, including 11 claudin paralogues, zonula-occludens scaffolds, and tricellular junction components, are coordinately enriched in the gill. Together, these findings highlight a gill-enriched immune landscape potentially shaped by filter-feeding ecology, providing a genomic framework for clupeid immunogenomics and aquaculture applications.
The giant triton snail (Charonia tritonis) is an ecologically critical predator that controls coral reef ecosystems by preying on crown-of-thorns starfish (COTS). However, overharvesting has driven severe population declines, threatening reef health. Despite its importance, the genomic underpinnings of its unique adaptations-including toxin resistance, prey detection, and environmental resilience-remain poorly understood. Here, we present the first chromosome-level genome assembly of C. tritonis (3.8 Gb), the largest sequenced gastropod genome, which encapsulates the large repetitive sequences comprising 70.1% of the genome and fueling massive genomic expansion. Macrosynteny and molecular dating analysis confirm a Jurassic-era (∼190 Mya) whole-genome duplication (WGD) that facilitated adaptive innovation. Comparative analyses identified gene family expansions critical to C. tritonis' ecological dominance: (i) toxin resistance via immune/detoxification systems (eg, galectins, Toll-like receptors, CYP450 enzymes); (ii) Sensory specialization (eg, rhodopsin expansions) enhancing prey detection in complex reef habitats; and (iii) DNA repair pathways (RADX proteins) supporting longevity and UV resistance. Tissue-specific expression profiling confirmed tentacle-exclusive localization of sensory receptors, directly linking genomic adaptations to foraging behavior. This study provides foundational insights into the evolution of a keystone marine predator, offering actionable genetic targets for conservation strategies and illuminating mechanisms that shape predator-prey coevolution in coral reefs.
Domestication alters animal behaviour, particularly tameness. We previously established 2 tamed mouse groups by selective breeding for active tameness-defined as the motivation to approach a human hand-from genetically heterogeneous wild-derived mouse stock, together with 2 nonselected control groups. Genetic analyses identified loci associated with active tameness, but their low heritability suggested contributions from nongenetic factors. We therefore hypothesized that the gut microbiota, which has been shown to influence brain function, contributes to behavioural changes associated with active tameness. To test this hypothesis, we conducted shotgun metagenomic analyses of faecal samples from 10 males and 10 females (80 individuals total) from the 2 tamed and 2 nonselected groups. Tamed mice exhibited markedly higher levels of active tameness, accompanied by elevated blood concentrations of oxytocin and pyruvate. While overall taxonomic and functional diversity of the gut microbiota was largely unchanged, the abundance of Limosilactobacillus reuteri was significantly increased in the tamed mice. Administration of a pyruvate-secreting L. reuteri strain to nonselected mice elevated blood oxytocin levels and enhanced active tameness, although plasma pyruvate levels were not increased. These findings suggest that L. reuteri is associated with behavioural modulation, potentially via oxytocin-related pathways, and provide mechanistic insight into microbial contributions to animal domestication.
Fusion genes play crucial roles in plant biological processes but remain far less explored than their human counterparts, largely due to limited validated datasets and the absence of plant-specific prediction tools. Existing approaches often produce high false-positive rates, restricting reliable discovery. To address this gap, we developed Plant Fusion Gene Predictor (PFGPred). This ensemble machine learning framework integrates Random Forest, XGBoost, and long short-term memory (LSTM) models into a meta-classifier for accurate identification of true and false fusion genes from RNA sequencing (RNA-Seq) data. PFGPred was trained on a high-confidence dataset of fusion genes validated by both RNA-Seq and whole-genome sequencing from Arabidopsis thaliana, Oryza sativa, Triticum aestivum, and Zea mays, to predict and rank candidate fusion genes for future functional validation. It outperformed individual baseline models, achieving accuracies of 0.97 on training data and 0.77 on independent test data. When evaluated on human datasets, it achieved 0.71 accuracy at the cost of lower sensitivity, reflecting biological differences between plant and human fusion events. Comparative analyses confirmed that PFGPred reliably identifies validated fusions, demonstrating its utility as a cost-effective, plant-specific prediction tool for high-throughput fusion gene screening and functional genomics research. It is freely available as a web server at http://www.nipgr.ac.in/PFGPred.
Allelic variation is a critical determinant of agronomic traits in heterozygous crops. Most existing approaches define variation as reference-anchored differences, such as SNPs or structural variants, confining allelic diversity to variant feature coordinates. Here, we present AlleleMiner, a Python-based pipeline that phases diploid gene sequences directly from PacBio HiFi reads. Rather than relying on reference-based coordinate systems for allele representation, AlleleMiner uses the reference genome solely to identify target gene region sequences and performs de novo assembly of read sets at each locus, minimizing reference dependence and reconstructing phased allele sequences. Across 18 citrus cultivars, the pipeline achieved an average phasing output of 91.5% of 1,409 single-copy genes, with coverage achieving. Coverage analyses using both real and simulated datasets indicated that a ∼30× HiFi depth is preferable for the stable recovery of heterozygous alleles, reducing potential allele dropout. Validation using pedigree information showed allele transmission patterns with known relationships. Using simulated haplotype data and the Citrus clementina assembly v1.0, AlleleMiner achieved complete-match reconstruction for both alleles at approximately 70% of loci. By enabling reference-minimized gene-level allele discovery, AlleleMiner provides a scalable framework for constructing allele databases and advancing marker-assisted and genomic selection in complex crops.
Cartilaginous fishes are divided into holocephalans and elasmobranchs, and they offer valuable systems for analysing the genetic basis of adaptation to diverse habitats and the evolution of chromosomal organization. Genomic studies on cartilaginous fishes were initiated early with holocephalans because of their compact genomes, but have concentrated primarily on the family Callorhinchidae. Here, we focused on the most species-rich holocephalan family Chimaeridae and characterized the genome of its member, silver chimaera (Chimaera phantasma), in pursuit of genomic traces of adaptation to deep-sea vision. The resulting genome assembly exhibited high continuity and completeness, enabling the first chromosome-level comparison among holocephalans. They displayed substantial intragenomic variation in chromosome length, correlated with intron size, alongside a high degree of one-to-one chromosomal homology. Our search for silver chimaera photoreceptor genes revealed a shrunken set of opsin genes, including rhodopsin exhibiting a sequence signature typical of deep-sea adaptation. We also performed whole-genome resequencing of multiple silver chimaera individuals of both sexes, which identified a putative X-chromosome fragment. This is the first evidence of a holocephalan sex chromosome and suggests male heterogametic sex determination. Our findings contribute to a deeper understanding of vertebrate genome diversity and lay the groundwork for future genetic studies on this species.
Chromosome-scale genome assemblies in gymnosperms have lagged behind those of angiosperms, likely due to their large genomes. Coniferous tree species, which belong to the gymnosperms, are important resources for wood production in the forestry industry. To elucidate the evolution and speciation of these species and establish genome resources for breeding, we integrated draft assemblies with optical and genetic mapping to construct chromosome-scale genomes for Japanese cypress (Chamaecyparis obtusa, 8.7 Gb), Japanese cedar (Cryptomeria japonica, 9.6 Gb), and Chinese fir (Cunninghamia lanceolata, 13.4 Gb). Additionally, we assembled and annotated their chloroplast and mitochondrial genomes. Comparative analysis of the nuclear genomes revealed that while synteny is largely conserved, distinct translocations and inversions occurred in chromosomes 2, 6, and 9. Notably, the significantly larger genome of Cu. lanceolata was associated with frequent tandem gene duplications rather than transposon expansion. These findings suggest that chromosomal rearrangements and segmental duplications played key roles in the divergence of these species. The genomic resources presented here including chromosome-scale sequences, gene annotations, and genetic maps will facilitate advanced conifer genetics and accelerate forest tree breeding programmes.
Nontuberculous mycobacteria occasionally harbour clustered tRNA genes, referred to as a tRNA array unit, which is considered a putative antidefense system within their genomes. However, the precise genomic location of these tRNA array units remains unclear. To address this, we sequenced the complete genomes of 5 Mycobacterium avium strains carrying a tRNA array unit using a hybrid assembly of long and short reads followed by manual curation. The assemblies indicated that each strain harbours 3 to 5 extrachromosomal elements. In all genomes, the tRNA array unit was found on a linear contig exceeding 300 kb. Pulse-field gel electrophoresis (PFGE) and sodium dodecyl sulphate-PFGE revealed that the strains harbour linear plasmids corresponding to these large contigs with protein-capped termini. These linear plasmids encode a hybrid type VII/type IV secretion system but lack relaxase genes, which are typically present in mycobacterial circular plasmids. Additionally, they contain approximately 415 bp inverted repeats at the termini. Sequences of related plasmids were identified exclusively in the genomes of M. avium isolates from Japan available in public databases, suggesting a possible Asian origin. This study provides the first experimental evidence that M. avium harbours giant invertron-type linear plasmids carrying a tRNA array unit.
Several species of toxic butterflies are known, including those from the Troidini tribe of the Papilionidae, which accumulate aristolochic acid from their host plants in Aristolochiae. However, the molecular mechanisms involved in utilizing aristolochic acid remain unknown. Toxic butterflies often exhibit warning colouration to signal their toxicity to predators, a complex adaptive trait with toxin utilization. In this study, we sequenced, assembled, and annotated the genomes of 2 toxic Troidini butterflies, Pachliopta aristolochiae (312.1 Mb, 13,497 genes) and Byasa alcinous (257.6 Mb, 14,669 genes), and conducted comparative genomics to identify genes involved in toxin utilization and warning colouration. Comparative analysis across 11 species revealed 31 gene families significantly expanded and 417 genes under positive selection. Additionally, 442 genes were highly expressed in the red spots on the hindwings of P. aristolochiae. The genes shared within these lists may be involved in the formation of the complex adaptive traits of toxin utilization and warning colouration. Functional analysis using RNAi confirmed the involvement of ebony, laccase2, and tyrosine hydroxylase (TH) in warning colouration. This research marks a significant starting point in understanding the genetic basis of aristolochic acid utilization and the formation of warning colouration, providing the first list of candidate genes.
The diatom Pleurosigma pacificum is a newly described tropical pelagic species from the Western Pacific Ocean with one of largest genome size among published diatom genomes, making it an ideal candidate for studying adaptation to tropical open ocean environments and diatom evolution. We employed HiFi long-read sequencing to construct a high-quality and contaminant-free genome. The assembled genome is 1.357 Gb in size and consists of 821 contigs with a contig N50 of 3.23 Mb. The GC content is 38.6%, which is much lower than that of other published diatom genomes. The genome contains 27,408 predicted genes, 540 of which were implicated in environmental adaptation. Gene features and gene family comparisons suggest that the primary driver of genome expansion and functional diversification is long terminal repeats (LTR) retrotransposons and tandem duplications. The phylogenetic analysis revealed that the clade of P. pacificum is closely associated with other members of Naviculales. The expansion of chlorophyll a/c proteins might facilitate the adaptation of P. pacificum to high-light conditions in pelagic environments. The percentage of approximately 3.2% horizontal gene transfer (HGT) events is observed in the P. pacificum genome. HGTs are a prevalent phenomenon in diatoms and serve as a common mechanism to enhance their adaptive capabilities. In conclusion, the P. pacificum genome provides important understanding into the development of large genome size and evolutionary adaptations of pelagic diatoms.
Genomic research is currently undergoing a paradigm shift from reliance on a single reference sequence to the use of breed-specific genomes. Chinese indicine cattle (Bos taurus indicus), characterized by their notable tick resistance and heat tolerance, display extensively genetic diversity than taurine. Here, we generated a chromosome level genome assembly of Chinese indicine cattle, achieving a contiguity N50 of 90.92 Mb and an overall size of 2.91 Gb, utilizing PacBio high-fidelity (HiFi) sequencing complemented by Hi-C sequencing technology. The assembly is characterized by near-complete chromosomes, telomeres, and less gaps. Utilizing this highly quality assembly, we explored the phylogenetic relationship and speciation time. The gene family and selection signatures analyses indicated that candidate genes and biosynthetic pathways potentially contributing to disease immunity and thermotolerance of indicine cattle. Altogether, this study enriches the bovine pangenome repository and advances our understanding of the complex evolutionary patterns and distinctive adaptation traits of Chinese indicine cattle.
Gerbera hybrida is one of the most popular ornamental plants and also serves as a valuable model plant within the Asteraceae family. Here, we report both the nuclear and organellar genome assemblies and annotations of G. hybrida, which was developed through hybridization of 2 wild species. Sequencing was performed using a combination of PacBio high fidelity (HiFi) reads and chromatin capture reads (Omni-C). The total span of the nuclear genome assembly is 2.32 gigabases, and 99.3% of the sequence assembled into 25 scaffolds, consistent with the known chromosome number. Genome annotation of the nuclear genome identified 36,160 protein-coding genes and 11,572 non-coding transcripts. The mitochondrial genome had 363,511 bp and contains 36 protein-coding genes, 3 rRNAs, and 21 tRNAs, while the chloroplast genome is 151,898 bp in length and includes 85 protein-coding genes, 8 rRNAs, and 37 tRNAs. A syntenic analysis of the G. hybrida genome and published Asterales genomes demonstrated that Gerbera has undergone a whole-genome triplication. This reference genome provides a foundational resource for future molecular breeding and genetic research in Gerbera and the broader Asteraceae family.
Malcolmia littorea, a member of the family Brassicaceae, is adapted to coastal and sandy environments and has become a model in studies of reproductive barriers. However, genomic resources for the species are limited. Here, with the aim of understanding the molecular mechanisms underlying key traits in M. littorea, we present a de novo genome assembly consisting of 10 chromosome-scale sequences. We employed a high-fidelity long-read sequencing technology for genome assembly. To anchor the sequences to chromosomes, we developed a single-pollen genotyping method to construct a genetic linkage map based on SNPs derived from transcriptomes of pollen grains, possessing recombinant haploid genomes. We built a genome assembly consisting of 10 chromosome-scale sequences (215 Mb in total) for M. littorea containing 30,266 predicted genes. A comparative genome analysis and gene prediction indicated that the genome of M. littorea is double the size of the Arabidopsis thaliana genome, consistent with a whole-genome duplication followed by gene subfunctionalization and/or neofunctionalization in M. littorea. This study provides a basis for research on M. littorea, an understudied species with ecological and evolutionary significance.
Chlorarachniophyte algae possess complex plastids derived from endosymbiosis between a cercozoan protist and green alga. As evidence of this event, remnant nucleus of the endosymbiont, nucleomorph, is present in the plastid intermembrane space. Chlorarachniophytes are excellent models to study genome evolution via endosymbiosis. Although the three organelle genomes of mitochondrion, plastid, and nucleomorph have been sequenced in several chlorarachniophyte species, nuclear genome information is currently limited to Bigelowiella natans. To gain insights into the genome diversity and evolution of chlorarachniophytes, we sequenced the nuclear genome of another chlorarachniophyte, Amorphochlora amoebiformis. Its size is approximately 214 Mb, which is more than twice that of B. natans. Remarkably, three-quarters of the nuclear genome encodes spliceosomal introns, indicating its highly intron-rich structure compared to other known eukaryotic genomes. Single nucleotide polymorphism analysis revealed that A. amoebiformis possessed a diploid nuclear genome, unlike the haploid genome of B. natans. Additionally, we identified organellar DNA fragments within the nuclear genome, suggesting recent DNA migration from the three organelles to the nucleus. Overall, our findings reveal that chlorarachniophyte nuclear genomes differ substantially in size, structure, and ploidy across species, and provide evidence of ongoing endosymbiotic gene transfer.
Oil palm (Elaeis guineensis Jacq.) is a globally important crop, and its genetic improvements benefit from comprehensive genome sequencing. Here, we report the whole-genome sequencing and annotation of two key genetic resources: the wild (Eg-DCM) and ancestral (Eg-DBG) Dura accessions, using a combination of short- and long-read sequencing technologies. De novo assembly followed by polishing, proximity ligation, and reference-guided scaffolding yielded high-quality assemblies with ungapped lengths of 1.71 Gb (Eg-DBG) and 1.48 Gb (Eg-DCM). Eg-DCM and Eg-DBG genomes exhibited high completeness, with over 97% of Benchmarking Universal Single-Copy Orthologs (BUSCOs) recovered across the Eukaryota, Viridiplantae, and Embryophyta datasets. Repetitive elements, particularly retrotransposons, dominated both genomes, accounting for 46.10% of Eg-DBG and 43.85% of Eg-DCM. Gene prediction initially identified 61,256 (Eg-DBG) and 53,985 (Eg-DCM) genes, which were refined into high-confidence gene sets of 39,263 and 35,298, respectively. Additionally, 1,760 and 1,684 putative resistance (R) genes were identified in Eg-DCM and Eg-DBG, with similar class distributions. The five major R gene classes comprise KIN, RLK, RLP, CNL, and CK. With further research, the assembled whole-genome sequences and the annotated genes of Eg-DBG and Eg-DCM offer valuable insights into the untapped genomic information of undomesticated accessions, with implications for future breeding and crop improvement efforts of oil palm.
Root exudates shape root-associated microbial communities that differ from those in soil. Notably, specific microorganisms colonize the root surface (rhizoplane) and strongly associate with plants. Although retrieving microbial genomes from soil and root-associated environments remains challenging, single amplified genomes (SAGs) and metagenome-assembled genomes (MAGs) are essential for studying these microbiomes. This study compared SAGs and MAGs constructed from short-read metagenomes of the same soil samples to clarify their advantages and limitations in soil and root-associated microbiomes, and to deepen insights into microbial dynamics in rhizoplane. We demonstrated that SAGs are better suited than MAGs for expanding the microbial tree of life in soil and rhizoplane environments, due to their greater gene content, broader taxonomic coverage, and higher sequence resolution of quality genomes. Metagenomic analysis provided sufficient coverage in the rhizoplane but was limited in soil. Additionally, integrating SAGs with metagenomic reads enabled strain-level analysis of microbial dynamics in the rhizoplane. Furthermore, SAGs provided insights into plasmid-host associations and dynamics, which MAGs failed to capture. Our study highlights the effectiveness of single-cell genomics in expanding microbial genome catalogues in soil and rhizosphere environments. Integrating high-resolution SAGs with comprehensive rhizoplane metagenomes offers a robust approach to elucidating microbial dynamics around plant roots.
Flowering cherries (genus Cerasus) are iconic trees in Japan, celebrated for their cultural and ecological significance. Despite their prominence, high-quality genomic resources for wild Cerasus species have been limited. Here, we report chromosome-level genome assemblies of two representative Japanese cherries: Cerasus itosakura, a progenitor of the widely cultivated C. ×yedoensis “Somei-yoshino,” and Cerasus jamasakura, a traditional popular wild species endemic to Japan. Using deep PacBio long-read and Illumina short-read sequencing, combined with reference-guided scaffolding based on near-complete C. speciosa genome, we generated assemblies of 259.1 Mbp (C. itosakura) and 312.6 Mbp (C. jamasakura), with both >98% BUSCO completeness. Consistent with their natural histories, C. itosakura showed low heterozygosity, while C. jamasakura displayed high genomic diversity. Comparative genomic analyses revealed structural variations, including large chromosomal inversions. Notably, the availability of both the previously published C. speciosa genome and our new C. itosakura genome enabled the reconstruction of proxy haplotypes for both parental lineages of “Somei-yoshino.” Comparison with the phased genome of “Somei-yoshino” revealed genomic discrepancies, suggesting that the cultivar may have arisen from genetically distinct or admixed individuals, and may also reflect intraspecific diversity. Our results offer genomic foundations for evolutionary and breeding studies in Cerasus and Prunus.