
Populus euphratica and Populus pruinosa represent core germplasm resources of desert riparian forests in northwest China. Accurate discrimination of putative interspecific hybrids between the two species is fundamental to population genetic research of desert poplars. Currently, publicly available and standardized genome-wide InDel marker resources specifically for desert poplars remain scarce. Traditional morphological identification exhibits low accuracy and reliability, and conventional markers (RAPD, AFLP, SSR) have limited repeatability and species specificity, failing to support standardized genotyping. Based on large-scale whole-genome resequencing data, this study developed and validated a set of highly specific and stable InDel markers to provide valuable genomic resources for subsequent genetic research on desert poplars. A total of 56,864 interspecific InDel loci with fragment length differences over 5 bp were screened from 439 P. euphratica and 146 P. pruinosa resequencing accessions, among which 3,417 loci had interspecific allele frequency differences > 0.9. After primer design and native polyacrylamide gel electrophoresis (native-PAGE) validation, 11 universal InDel primer pairs were obtained, and two robust diagnostic markers (P1599-3, P1599-4) were finally verified. Population validation involving 50 P. euphratica, 41 P. pruinosa and 93 morphologically selected putative-hybrid individuals revealed that this dual-marker set could reliably detect biparental banding patterns among tested candidates. The two verified co-dominant InDel markers provide reusable molecular resources for preliminary germplasm authentication and screening of putative hybrids within Populus sect. Turanga. The standardized screening and PAGE genotyping pipeline provided in this study can be transferred to other woody hybrid species, supporting marker-assisted breeding and germplasm conservation of arid poplars.
The gharial (Gavialis gangeticus) is a freshwater crocodylian endemic to the Indian subcontinent that has experienced ancient and recent catastrophic declines in population size. Conservation interventions initiated in the mid-1970s have enabled partial recovery of the species. Yet, how ancient and recent demographic declines have shaped contemporary genomic diversity in gharial remains unknown. Here, we present the first genome-wide analysis of the gharial using a single high-coverage whole-genome sequence and ddRAD data from 21 individuals sampled across five nesting sites along the Chambal River. Our results reveal extremely low genomic diversity in the largest gharial population, with Watterson’s θ = 0.0008 ± 0.0001 (mean ± SD) and Nucleotide diversity = 0.00036 ± 0.00005, both substantially lower than those reported for other extant crocodylians. Analyses of population genomic structure indicated weak, partially resolved structuring, with individuals from Radi forming a distinct group, whereas individuals from other sites showed limited spatial structuring. Demographic reconstructions revealed declines in effective population size across multiple timescales, including an ancient decline occurring 10,000 years ago and a more recent severe decline within the last 250 years. Consistent with this recent decline, long ROH suggest that recent close inbreeding ( 4 generations), rather than ancient demographic history alone, is the primary contributor to contemporary genome-wide homozygosity. Overall, these findings reflect both deep-time demographic constraints and ongoing genetic erosion in gharial, providing a critical foundation for genomics-informed conservation planning.
Staphylococcus aureus is an important cause of healthcare- and community-associated infections and can acquire antimicrobial resistance (AMR), virulence determinants, and mobile genetic elements that influence colonization, persistence, and disease potential. Whole-genome sequencing (WGS) enables integrated isolate-level characterization of these features, including in resource-constrained hospital settings where genomic surveillance data remain limited. We performed a descriptive WGS-based characterization of two S. aureus isolates collected from the orthopaedics ward at Bugando Medical Centre, Mwanza, Tanzania, in 2020. A55934 was recovered from a wheelchair handle on 16 January 2020 and is described as a fomite-associated isolate. A55935 was recovered from a rectal swab collected from a 43-year-old male patient 24 h after admission for a patella fracture on 6 April 2020 and is described as a patient-associated rectal-swab isolate. The two isolates were collected approximately 11 weeks apart and were not used to infer transmission or contemporaneous circulation. Genomes were analysed using the rMAP 2.0 bioinformatics pipeline for lineage assignment, AMR determinants, virulence-associated genes, and mobile genetic elements. The isolates were also placed within a contextual core-genome phylogeny of publicly available Tanzanian S. aureus genomes. The two isolates belonged to distinct lineages: A55934 was assigned to ST152, while A55935 was assigned to ST8. The patient-associated rectal-swab ST8 isolate carried mecA, consistent with an MRSA genotype, together with additional predicted AMR determinants including erm(C), tet(K), and dfrG, and multiple plasmid replicons. A blaOXA−60-like hit was also detected but is interpreted cautiously because OXA-type beta-lactamases are atypical in S. aureus and require further validation. In contrast, the fomite-associated ST152 isolate had no acquired AMR genes detected at the reporting thresholds and was consistent with an MSSA genotype. Both isolates carried genes associated with adhesion, biofilm formation, immune evasion, and other virulence-related functions, while leukocidin-associated genes were detected only in the ST8 isolate at the applied thresholds. In the contextual phylogeny, the two study isolates occupied distinct positions within broader Tanzanian S. aureus diversity. This two-isolate analysis provides a descriptive, proof-of-concept genomic characterization of source-distinct S. aureus isolates from a Tanzanian orthopaedics ward. The findings show that WGS can resolve lineage, AMR genotype, virulence-gene content, and mobile-element profiles from both fomite-associated and patient-associated isolates. However, because the isolates were few, non-contemporaneous, and not epidemiologically linked, these data cannot establish transmission, ward-level circulation, environmental reservoirs, or representativeness of the broader S. aureus population. Larger longitudinal studies integrating patient and environmental sampling, phenotypic antimicrobial susceptibility testing, and detailed epidemiological metadata are needed to define transmission dynamics and the contribution of high-touch surfaces to S. aureus persistence and AMR dissemination in this setting. Not applicable.
Klebsiella pneumoniae (K. pneumoniae) is a known pathogen that has been implicated in both hospital and community-acquired infections. Several strains are emerging as hypervirulent, raising public health concerns. This study aimed to gain genomic insight into multidrug-resistant (MDR) K. pneumoniae AA strain isolated from the leafy vegetable Corchorus olitorius (Ewedu), specifically characterizing its resistome, mobilome and virulence determinants. K. pneumoniae AA was isolated from the vegetable using standard microbiological methods. Antibiotic susceptibility testing was performed using the disc diffusion method, and isolates were screened for extended spectrum β-lactamase (ESBL) production using the double-disc diffusion method. Whole-genome sequencing was carried out using the PacBio Onso sequencing platform. K. pneumoniae AA was positive for ESBL production and was MDR, showing resistance to eight of the ten antibiotics tested. The genome of K. pneumoniae AA was 5.4 Mb in size, comprising 558 contigs and a G + C content of 57.5
Bacteria respond to iron limitation by activating distinct uptake and siderophore-biosynthesis pathways, and comparative transcriptomics under varying iron availability provide insights into these diverse adaptive mechanisms. Here we describe RNA-seq datasets generated from four freshwater bacterial isolates—Pseudomonas sp. FBCC-B13192, Herbaspirillum sp. B12834, Pantoea sp. FBCC-B5559, and Micrococcus sp. FBCC-B5738—cultured under FeCl3-treated and untreated conditions. For each strain–condition combination, three biological replicate cultures were prepared, their RNA was pooled, and a single sequencing library was constructed. : The dataset comprises eight paired-end libraries (16 FASTQ files) generated in 2024 and 2025, totaling 349,853,404 processed reads. Raw sequence data are openly available in the NCBI Sequence Read Archive under SRA study accession SRP695241 (BioProject PRJNA1456794; SRR38280322–SRR38280329; BioSamples SAMN57450602–SAMN57450609). Read alignment rates were 78.59
We present the de novo whole-genome assembly of a deepwater rice variety (Oryza sativa cv. Pin Gaew 56). Although the submergence escape response is a vital survival strategy in deepwater rice, current reference genomes lack sufficient genetic information to fully understand the molecular mechanisms behind this complex physiological process. Pin Gaew 56 (PG56) is a well-studied genetic resource with extensively documented physiological and molecular traits related to internode elongation under partial submergence. This study provides a cultivar-specific genomic resource for PG56 that can support further investigation of the genetic basis of the submergence escape response. A de novo whole-genome assembly of the deepwater rice cultivar PG56 was generated to support genomic studies of the submergence escape response. The final assembly has a total length of 407.7 Mb, comprising 138 scaffolds with an N50 of 32.5 Mb. Genome completeness assessed by BUSCO showed 94.5
Abstract Objectives Production of high-quality biological products depends on the careful selection and consistent management of seed strains. In this study, we generated complete genome sequences of bacterial strains employed in inactivated animal vaccines in South Korea and comprehensively examined their gene annotations. The genome data thus obtained provide a solid basis for confirming, at the strain level, whether manufacturers are using the same seed strains over time. In addition, these data offer a useful reference resource to support stable large-scale production and quality control of biological products. Data description The complete genomes of eleven bacterial strains employed for inactivated vaccine production were generated and annotated. A hybrid assembly workflow combining Illumina short reads (NovaSeq 6000) with Oxford Nanopore long reads (MinION) produced high-quality, gap-free genomes for all strains, with genome sizes ranging from 2.28 Mb to 5.35 Mb. The strains comprised Pasteurella multocida (D, 3A, A), Actinobacillus pleuropneumoniae (2 and 5), Glaesserella parasuis 4, Mannheimia haemolytica KO, Avibacterium paragallinarum C, and three Escherichia coli strains (F41, YC21-F17, K99S). All assemblies exhibited high completeness (>99%) and minimal contamination (<1%), ensuring reliable downstream genomic characterization.
Members of the genus Marinobacterium are widely distributed in marine environments and contribute to diverse ecological processes; however, genomic information for several species remains limited. In this study, we report the complete genome sequence of Marinobacterium marisflavi strain IMCC4074ᵀ, originally isolated from coastal seawater of the Yellow Sea, to provide insights into its genomic features, metabolic potential, and environmental adaptation. The genome of strain IMCC4074ᵀ was sequenced using a combination of short- and long-read sequencing approaches and assembled into a single circular chromosome of 3,091,487 bp with a G + C content of 52.45
Phylogenetic and functional genomic analyses were carried out to help explaining the endophytic lifestyle and the biocontrol activity of Streptomyces sp. strain DEF39, able to reduce Fusarium graminearum infection as well as deoxynivalenol production in wheat plants. Specifically, in this work, genes and biosynthetic pathways linked to secondary metabolite production, biocontrol activity, and plant interactions were identified. This approach supports the experimentally validated capability of DEF39 to protect plants from fungal diseases, highlighting the potential of genomics studies. Streptomyces sp. DEF39 was isolated as an endophyte from Secale cereale in Italy and has been shown to colonize wheat (Triticum spp.) and act as a biocontrol agent of Fusarium head blight. The complete genome (9.12 Mb) (GCA_978019405.1), assembled through a hybrid strategy based on long and short accurate reads combination, is composed of a linear chromosome of 8,887,323 bp with a GC content of 71.59
Microorganisms inhabiting cold and oligotrophic aquatic environments experience persistent physiological stress, necessitating genomic characterization to understand their survival strategies. Psychrotolerant strains have evolved diverse metabolic adaptations, including secondary metabolite biosynthesis, which may contribute to environmental fitness and offer potential for low-temperature biotechnological applications. However, the genus Lacisediminihabitans remains poorly represented at the genomic level, limiting our understanding of its ecological roles and metabolic potential. To address this gap, we generated a high-quality complete genome of a psychrotolerant Lacisediminihabitans strain isolated from Antarctic freshwater. Lacisediminihabitans sp. FW035 grew at 2 − 25 °C with an optimum at 20 °C. The genome of strain FW035 is 3,842,169 bp in size with a G+C content of 66.4
Crateva unilocularis is a widely consumed woody vegetable in the family Capparaceae. This study aimed to generate and comprehensively characterize the first complete mitochondrial genome of C. unilocularis using PacBio HiFi long-read sequencing, elucidating its structural organization and gene content. We also report the complete chloroplast genome for this accession. Together, these organelle genome resources provide a foundation for downstream studies on phylogeny, genome architecture, and adaptive evolution in C. unilocularis and related taxa. The chloroplast genome length of C. unilocularis is 156,525 bp, harboring 78 unique protein-coding genes (PCGs), 28 transfer RNAs (tRNAs), and four ribosomal RNAs (rRNAs). In addition, the assembly graph supports three circular mitochondrial molecules with a combined length of 566,203 bp, and junction-spanning HiFi reads validate the assembled configurations. Annotation identified 59 genes in total, including 37 PCGs, 19 tRNAs, and three rRNAs. The assembled and annotated organelle genomes represent key reference sequences for C. unilocularis within Capparaceae and will facilitate future comparative and functional genomic analyses.
Aspergillus niger is a renowned filamentous fungus with extensive industrial and biotechnological applications. A. niger is widely used to produce diverse organic acids, enzymes, and other value-added metabolites. Although strain-specific metabolic capabilities have been widely applied in industry, the genomic foundations of these capabilities have not yet been fully characterized. Here, we have described the isolation and provided the draft genome sequence of A. niger strain AN-L103_M1 to further study the genetic determinants of primary and secondary metabolism, given its metabolic versatility and great potential in biotechnology. This A. niger strain was developed through gamma radiation bombardment, and subsequently, all the mutants were screened through various steps for hyperproduction of citric acid. This genome sequence provided valuable information for further functional genomics and strain-improvement research. The A. niger strain was obtained through gamma radiation bombardment. After gamma radiation bombardment of the A. niger culture, it was subjected to various screening steps to achieve hyperproduction of citric acid. These experiments were performed in the Industrial Biotechnology Division, National Institute for Biotechnology and Genetic Engineering, Faisalabad, Pakistan. The mycotoxin production potential of this mutant strain was assessed using LC-MS analysis, which detected no mycotoxins.
Mushrooms are grossly under exploited and efforts to domesticate them are not yielding enough results as over 95
Lacticaseibacillus rhamnosus is a widely studied probiotic species with notable functional diversity among strains. To expand the genomic resources available for this species and to support future comparative and probiotic-related studies, we sequenced and analyzed the complete genome of L. rhamnosus B34-0-2, a strain isolated from Panax ginseng, with a focus on its genomic features potentially associated with probiotic-related traits. Genomic DNA was extracted and sequenced using a combination of PacBio long-read and Illumina short-read platforms. The assembled genome comprised 2,810 predicted coding sequences (CDSs), 59 tRNA genes, and 15 rRNA genes. Functional classification assigned 2,761 CDSs (98.25
Secreted proteins are known to be an important virulence factor in the successful invasion and colonization of the host by fungi, especially for biotrophic parasites/pathogens of plants. To predict protein secretion, a genome must be evaluated by software that interrogates the genome for conserved sequences characteristic of known secreted proteins. This work sought to identify putative secreted proteins of Microbotryum intermedium, a smut fungus found on plants of the Scabiosa family. To accomplish this, we used a pipeline of computational tools to serve as an efficient means of identifying potential targets as fungal effectors that can later be evaluated experimentally. So, this initial identification is ideally the first step in a larger investigation of the role of protein secretion in the development and progression of disease. A pipeline of computational tools was used to stringently predict secreted proteins for Microbotryum intermedium, The pipeline was used to predict the canonical secretome of this fungus. The annotated M. intermedium genome (mycocosm.jgi.doe.gov/Micin1/Micin1.home.html) has 8,148 predicted genes, of which, only 296 were classified as putative secreted proteins using the stringent pipeline for canonical secretion prediction. These data inform future functional analyses that test candidates for their roles in pathogenicity.
This Data Note reports the complete genome sequence and associated functional-genomic data of Methylomonas sp. strain 2F7, a methane-oxidizing bacterium isolated from rice paddy soil in South Korea. These data were generated to support comparative genomic analyses of methanotrophic bacteria and to document the genome-based taxonomic position and C1 metabolism gene repertoire of strain 2F7. The dataset includes Oxford Nanopore Technologies MinION raw reads, a two-replicon genome assembly, annotation summaries, functional classifications, a phylogenetic analysis, average nucleotide identity (ANI) and digital DNA–DNA hybridization (dDDH) comparisons, and a curated C1 metabolism gene table. The genome consists of a 4,979,984-bp chromosome and a 23,603-bp plasmid, with an overall G+C content of 51.4
Members of the genus Litorivicinus are marine heterotrophic bacteria with limited genomic representation, constraining our understanding of their taxonomic diversity, metabolic capabilities, and ecological roles. The genus currently comprises only two validly published species, with genomic data available exclusively for the type species, Litorivicinus lipolyticus. Here, we report the complete genome sequence of Litorivicinus marinus strain IMCC2782ᵀ, isolated from surface seawater of the Yellow Sea, Republic of Korea, to provide insights into its genomic features and metabolic potential. The genome of strain IMCC2782ᵀ was sequenced using a hybrid approach combining short-read (Illumina) and long-read (PacBio) technologies. The assembly yielded a single circular chromosome of 2,204,001 bp with a G + C content of 56.07
The FOXE1 gene is a transcription factor critical for thyroid development and function, and has been associated with thyroid disorders and other phenotypic traits such as cancer and orofacial clefts. This study explores the genetic variability in FOXE1, focusing on non-synonymous single nucleotide variants (nsSNVs) and untranslated region (UTR) variants, using in silico approaches to assess their functional and structural impacts. A total of 1,003 variants were retrieved from the Genome Browser, including 306 nsSNVs and 508 UTR variants. After analysis with the Variant Effect Predictor and eight additional prediction tools, 37 high-risk variants were identified. Of these, 23 variants led to smaller mutant amino acids, 13 to larger substitutions, and 1 had no size-related change. Structural predictions showed altered hydrophobicity and torsion angles in several variants, with a suggestion of significant functional consequences. Decreased stability was observed in 32 high-risk variants, indicating potential disruptions to protein function. Additionally, UTR variants rs41274262 and rs7043516, were predicted to increase regulatory activity slightly. Frequency analysis demonstrated considerable variability in the prevalence of these variants across different populations, with some UTR variants showing higher frequencies in African and East Asian populations. Furthermore, associations with clinical phenotypes, including thyroid hypoplasia, congenital hypothyroidism, and orofacial clefts, were noted for specific variants. Conservation analysis revealed that several UTR variants are highly conserved across species, suggesting potential functional roles in humans. This study provides comprehensive insights into the structural, functional, and regulatory impacts of genetic variation in the FOXE1 gene and its relevance to human health. Not applicable.
The objective of the study is to sequence the whole genome of multidrug resistant E. coli strain KAB-AI-497 that causes bacterial vaginosis and implicated in premature rupture of membrane in pregnant woman. The DNA of the E. coli strain KAB-AI-497 was extracted using the MagAttract HMW DNA Kit, and the extracted DNA was sequenced using an MGI DNBSEQ G99ARS platform. FastQC was used to perform quality control analysis and the reads were trimmed by Trimmomatic. De novo genome assembly was performed by SPAdes and it resulted to a draft assembled genome that has 5.1 Mb genome size, 153 contigs, and 50.5
Streptomyces californicus strain ADR1 is an endophytic actinobacterium isolated from Datura metel that produces secondary metabolites with potent antibacterial and anti-biofilm activities against WHO-listed high-priority Gram-positive pathogens. While anti-bacterial and antioxidant potential of the strain ADR1 has been extensively characterized, its complete genome sequence remains to be investigated for further insights into its biosynthetic potential. This study presents the complete genome sequence analysis of the strain ADR1 to provide a robust genomic foundation for understanding its metabolic versatility and biosynthesis of compounds with therapeutic significance. The ADR1 genome was sequenced using Illumina HiSeq. The assembly comprised 262 scaffolds with a total genome size of 8.4 Mb and G + C content of 72.5