
Drosophila nepalensis is a cold-adapted drosophilid endemic to the Himalayan region. Its ability to survive in harsh, cold conditions makes it a valuable Drosophila model for investigating how adaptation to thermal extremes may influence species persistence under future climate change. Here, we report the first de novo genome assembly of D. nepalensis, based on a hybrid sequencing strategy that combines Illumina short reads and Oxford Nanopore long reads. Illumina sequencing generated 49.88 million 150 bp paired-end reads (∼14.96 Gbp), while Nanopore sequencing produced 1.35 million long reads totaling ∼0.76 Gbp. The assembled genome spanned ∼178 Mb with an N50 of 83.6 kb and 98% BUSCO completeness, comparable to other well-annotated Drosophila genomes. Annotation identified 10,560 protein-coding genes, including transcription factor-rich and stress-related domains such as zinc fingers, WD40 repeats, and ankyrin motifs. Comparative orthology analysis across 6 Drosophila species identified 14,168 orthologous clusters, of which 9,173 were shared among all 6 species, indicating a conserved core genomic set across the sampled taxa. D. nepalensis showed 83 unique orthogroups and 50 singletons, suggesting some lineage-specific gene expansions associated with cold adaptation and endemicity, including families encoding caspase-family apoptotic regulators, chromatin remodeling proteins (HMGB/protamine-like), and SNARE-domain vesicle trafficking factors. Gene family evolution analysis revealed the highest expansions in the cold-tolerant Himalayan drosophilid, D. nepalensis, including significant expansions in serine protease, chaperone, and neurotransmitter transporter families, alongside dramatic contractions of core histone gene families, suggesting lineage-specific chromatin remodeling and ecological specialization.
Genetic perturbations are one of the great strengths of the model organism Drosophila melanogaster, with approaches such as classical mutagenesis and RNA interference enabling a wealth of biological discoveries. A more recent approach for altering gene expression is CRISPR/Cas9-based mutagenesis. As with any new tool, however, its use must be optimized. High expression of Cas9 has been shown to cause cytotoxicity in some cell types, including class IV dendritic arborization (da) neurons. In this study, we provide evidence that Cas9 expression causes cytotoxicity in class I da neurons, in addition to class IV da neurons, both of which are widely used to study neuronal development and regeneration. We then systematically evaluated available Cas9 transgenes designed to titrate Cas9 expression, called uCas9 transgenes. We show that the expression of these uCas9 transgenes results in little to no cytotoxicity in various classes of da neurons. Immunostaining revealed drastic reductions in Cas9 protein levels for da neurons expressing the uCas9(L) transgene. Lastly, we demonstrate that the uCas9(L) transgene effectively and specifically gene edits in both class I and class IV da neurons, lowering the expression of GFP-tagged proteins and producing loss-of-function morphological phenotypes when targeting endogenous loci. Thus, we refine the use of CRISPR mutagenesis in Drosophila da neurons through titration of Gal4/UAS-mediated Cas9 expression using existing uCas9 transgenes, a feasible and flexible approach that may be useful for other labs encountering Cas9 cytotoxicity in their own model systems.
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) continues to pose a threat to humans as well as domestic and wild animals. The variability in severity of clinical signs, the zoonotic potential, and the host-specific response to infection contribute to the persistence of circulation of disease. In wildlife species, white-tailed deer have been shown to be more permissive to infection than bovids. However, among bovids, American bison have shown a greater susceptibility than cattle. In this study, we investigate the transcriptomic response to experimental SARS-CoV-2 infection in bison over time. Substantial numbers of differentially expressed genes were identified between pre- and 2, 5, 7, 14, and 21 days post-inoculation. Kyoto Encyclopedia of Genes and Genomes and Gene Ontology term analysis identified associations with immune response, inflammatory response, and viral infection including COVID-19. Ingenuity Pathway Analysis of the coronavirus pathway highlighted differences in signaling at days 2 versus 21 post-inoculation. We additionally examined changes in the nasal microbiome of bison over the course of experimental infection, which suggested an increase in opportunity for secondary infection causing pathogens such as Mannheimia. Collectively, this study presents a profile of bison transcriptomic response to SARS-CoV-2 infection and continues to expand our understanding of variation in host response.
Metabolomics provides direct insight into physiological state, but for small organisms such as Drosophila melanogaster, it typically requires pooling individuals to obtain sufficient material. Pool sizes vary widely across studies with little justification, and the impact pooling and biological replication have on metabolomic characterization and signal detection remains poorly understood. We evaluated the effects of pool size and biological replication on metabolomic profiles and signal detection using two complementary designs in D. melanogaster. First, we tested how pooling (5, 50, or 100 individuals) affects metabolomic structure and reproducibility in inbred and outbred populations. Second, we tested how pool size interacts with replicate number to affect detection of diet-associated metabolite changes under a high-sugar perturbation. Pool size shaped metabolomic profiles: pools of five individuals consistently differed from larger pools, which improved reproducibility in a dataset-dependent manner. In the dietary experiment, smaller pools showed reduced sensitivity, detecting fewer true diet-associated metabolites without increasing false discoveries, and replicate downsampling showed that pool size and replication independently shape signal retention. Detection depended on effect size and variability: metabolites with larger, more stable effects were consistently retained, while smaller more variable effects were rapidly lost under reduced sampling. Beta-binomial modeling confirmed that detection probability reflects a balance between signal strength and measurement variability, with pool size and replicate number independently shaping this relationship. Together, these results show that metabolomic inference depends on the interplay of signal, noise, and sampling design, with pool size and replication jointly shaping the detectability, stability, and interpretation of biological signals.
Fruit texture is a key quality trait in cranberry (Vaccinium macrocarpon Ait.), with a direct impact on processing efficiency and overall profitability for the industry. To elucidate the genetic basis of cranberry texture, we conducted the first QTL mapping study in this crop implementing a fruit compression method with 10 mechanical texture traits in two biparental populations, CNJ02 (n=168) and CNJ04 (n=67), over two years. These 10 traits, associated with flesh hardness/stiffness and elasticity, showed similar phenotypic patterns of transgressive segregation and generally strong positive correlations. However, traits accounting for fruit size and shape effects (stress/strain-related traits and maximum contact pressure), together with slope-related parameters, exhibited consistently high heritability across populations, supporting their potential utility for breeding and genetic analyses. Genetic mapping revealed multiple consistently detected multi-trait QTL distributed across chromosomes 1,2,3,4,5,8,9,10,11 and 12, suggesting a complex genetic architecture. These genomic regions explained, on average, 4.67-18.6% of the phenotypic variance in CNJ02 and 10.97-27.75% in CNJ04. Candidate genes identified within these regions point to cell wall dynamics, including pectin modification, cellulose biosynthesis, and hemicellulose metabolism, as potentially relevant processes underlying cranberry fruit texture. Our findings provide loci in association with texture traits in cranberry and insights for the genetic improvement of this essential trait, informing the design of genotyping platforms and the implementation of genomic-assisted breeding approaches. Furthermore, we anticipate that these results will guide future studies aimed at optimizing postharvest fruit handling and management practices.
Pharmacogene missense variants can disrupt protein stability, catalytic competence, or substrate handling through distinct mechanisms. General-purpose predictors estimate clinical pathogenicity as a single scalar, whereas pharmacogene interpretation requires knowing which biochemical dimension a variant perturbs, since that determines whether reduced function is substrate-dependent. Five deep mutational scanning datasets comprising 26,198 missense variants across CYP2C9, CYP2C19, and NUDT15 were assembled from MaveDB. Paired assays showed that this dimensionality dominates the data: 28% of CYP2C9 variants (1,236 of 4,421) decoupled catalytic activity from abundance, and 48% of NUDT15 variants (1,364 of 2,844) decoupled thiopurine sensitivity from stability, with CYP2C9 discordance concentrating at substrate-channel residues. AlphaMissense, a representative general-purpose pathogenicity predictor, scored these classes in line with its clinical training objective rather than the assayed biochemistry, assigning likely-benign scores to 38 of 195 stable-but-dead CYP2C9 variants and likely-pathogenic scores to 140 of 222 destabilized but thiopurine-resistant NUDT15 variants. To test whether this dimensionality is recoverable, a supervised ESM-2 sequence baseline was benchmarked against the ESM1v zero-shot ensemble and AlphaMissense under position-based 5-fold cross-validation, together with three architectural extensions: AlphaFold structural features, multi-task learning across paired assays, and contact-graph neural networks. The baseline reached Pearson r of 0.54-0.72, matching or marginally exceeding both comparators, and no extension improved upon it. Trained directly on each assay, it nonetheless recovered the paired-assay difference at r = 0.28 for CYP2C9 and 0.43 for NUDT15, separating discordant variants at AUROC 0.60 and 0.51. Pharmacogene interpretation therefore requires assay-specific, substrate-aware functional measurements rather than a single generic score.
Tumor Suppressor Candidate 3 (TUSC3) is an integral component of the oligosaccharyltransferase complex and is required for the N-glycosylation of proteins. While TUSC3 has been implicated in cancer and autosomal recessive non-syndromic intellectual disability, some patients also exhibit distinct facial features. However, the role of Tusc3 in craniofacial development is unknown. Here, we report a patient presenting with cleft lip and palate who was found to carry a homozygous variant in the TUSC3 promoter region in addition to a maternally inherited chromosomal duplication. In order to evaluate a potential role for TUSC3 in craniofacial development, we analyzed Tusc3 deletion mouse mutants from the International Mouse Phenotyping Consortium (IMPC). Initial IMPC phenotyping data suggested that a subset of Tusc3 deletion mice could develop cleft palate, micrognathia, and aglossia. We further identified that Tusc3 is expressed in the craniofacial and brain regions during embryonic development. Upon characterizing the Tusc3 deletion line, we observed that most homozygous mutants exhibit preweaning lethality. While craniofacial defects occur at a very low frequency in the deletion mutants, we identified a modest reduction in cortical area. Furthermore, RNA-seq analysis surprisingly revealed that no other genes in the developing brain or face were significantly affected upon loss of Tusc3. These findings suggest that Tusc3 can contribute to congenital malformations in both craniofacial and cortical development.
The vine mealybug, Planococcus ficus, is a globally invasive pest of grapevine and a vector of leafroll viruses. Like other mealybugs, it reproduces through paternal genome elimination, a sex-determination system that operates without sex chromosomes and is associated with extreme sexual dimorphism. To characterize genome organization and sex-biased expression in this species, we generated a long-read reference genome spanning 369 Mb with 23,489 annotated genes and macrosynteny conserved with the citrus mealybug, Planococcus citri. Resequencing of four California field individuals provided the first genome-wide baseline of nucleotide diversity for P. ficus and 132 cross-validated microsatellite markers for population monitoring. Phloem feeders secrete proteins that mediate host interaction; of 2,129 predicted secreted proteins in P. ficus, a conserved core is shared with P. citri and each carries a lineage-specific set. Comparing adult male and female transcriptomes, we found sex-biased expression to be pervasive and skewed toward females: 41% of tested genes differed between the sexes, with female-biased genes both more numerous and showing larger fold changes. These female-biased genes were not randomly distributed but concentrated in discrete blocks of coordinately expressed, tandemly duplicated gene families, a pattern not previously described in a mealybug. Male- and female-biased secreted proteins also differed in origin, with male-biased proteins drawn from a conserved repertoire shared with P. citri and female-biased proteins spanning a more lineage-specific pool. Together, these results reveal a female-skewed, spatially clustered architecture of sex-biased expression in a mealybug that lacks sex chromosomes, and provide genomic resources for managing an invasive vineyard pest.
Sexual dimorphism provides a powerful framework for studying how regulatory mechanisms generate phenotypic divergence from a shared genome. Many sex-specific traits arise from differences in gene expression, which can be mediated by epigenetic mechanisms such as DNA methylation. These mechanisms contribute to the organisation of gene regulatory variation across tissues, genes, and chromosomes, with potential consequences for the emergence and maintenance of sexually dimorphic phenotypes. Using a comparative epigenomic approach we investigate sex-biased DNA methylation in gonad and muscle tissue across two closely related Poeciliid species, Poecilia reticulata and Poecilia wingei. We address three questions: (i) how conserved are sex-biased DNA methylation patterns across species and tissues; (ii) are conserved sex-biased methylation patterns associated with genes involved in developmental and transcriptional regulation; and (iii) are sex-biased epigenetic patterns enriched on the sex chromosome? We identify extensive conservation of sex-biased differentially methylated regions (DMRs) in gonadal tissue, with strong cross-species concordance, whereas muscle shows little conservation of sex specific differential methylation. Conserved male-hypomethylated regions are enriched near genes involved in developmental transcriptional regulation and gonadal differentiation. Methylation-expression coupling distinguishes sex-biased reproductive gene functions, with male-biased genes enriched for spermatogenesis and motility and female-biased genes enriched for sperm-egg interaction processes. Sex-biased DNA methylation is broadly distributed across the genome, with localized regulatory regions on the sex chromosome. Together, these results show that conserved sex-biased DNA methylation is highly tissue-specific and associated with reproductive regulatory function, supporting a model in which gonadal epigenetic architecture reflects conserved regulatory organisation underlying sexual dimorphism.
Texas wintergrass is a cool-season perennial bunchgrass native to North America with ecological and agronomic importance as a winter forage species, yet genomic resources for this species remain limited. Here, we present the first de novo genome assembly of N. leucotricha generated using PacBio HiFi sequencing. Assembly with hifiasm produced a 959.92 Mb genome with 8X coverage, comprising 2,394 contigs with an N50 of 530.6 kb. Assembly completeness was high, with 99.6% of Benchmarking Universal Single-Copy Orthologs (BUSCOs) identified, although a substantial proportion were duplicated. Genome size estimates based on flow cytometry and k-mer analysis, together with assembly metrics, indicate a repeat-rich and structurally complex genome. Repeat annotation revealed that 64.42% of the genome consists of repetitive elements, predominantly long terminal repeat (LTR) retrotransposons, including Ty1/Copia, and Gypsy/DIRS1. Gene prediction using the homology-based pipeline GeMoMa identified 54,132 high-confidence genes, of which 53,868 were functionally annotated and 26,346 were assigned to KEGG pathways. Mining for genome-wide simple sequence repeats (SSRs) identified over 38,000 markers, with a validated subset demonstrating use for genetic diversity analysis. Reference-guided alignment and scaffolding using Brachypodium distachyon and Oryza sativa provided comparative frameworks for evaluating conserved sequence relationships between Texas wintergrass and representative grass genomes. Our results establish a genomic resource for Texas wintergrass, supporting future studies in comparative genomics and molecular breeding.
The National Variety Trials (NVT) - the largest coordinated crop trial network in the world - is conducted each year to evaluate the yielding ability of commercial wheat cultivars in Australia. As the complexity of genotype-by-environment (G×E) interactions in these trials poses challenges for variety selection, envirotyping represents a promising way forward for more informed decision making. Here, we proposed to classify 800 NVTs into envirotypes and use those to model G×E as genotype-by-envirotype (G×ET) interaction. We compared approaches for such classification that varied in how phenotypic and environmental information are considered. The simplest approach was based on dimensionality reduction and clustering techniques using only static environmental indices - i.e., averaged by phenological period. The most complex method combines the inclusion of yield data for the selection of environmental traits and functional environmental indices to capture the dynamic nature of G×E over the season. Across methods, we found discrepancies in the defined envirotypes but consistency in the indices relevant for G×E variability. A substantial fraction of the G×E was captured by four defined envirotypes (18% across regions), offering a practical opportunity to leverage multi-environment yield and stress data to inform breeders, agronomists, and growers in variety recommendations. We discussed the need for an integration with multivariate Earth observations as way to anticipate envirotypes and support pre-season G×E predictions.
Species within the quinaria group of the genus Drosophila are models for ecological genetic studies on topics that include morphological diversity, color pattern development, and feeding behavior. Drosophila palustris and Drosophila subpalustris are closely related members within the quinaria species group that differ primarily in subtle morphological and behavioral traits. To support future evolutionary genomic studies in drosophilids, we completed new draft hybrid assemblies and annotations of the D. palustris and D. subpalustris genomes. We obtained a fragmented assembly for D. palustris (N50 = 874 Kb) and a near chromosome-scale D. subpalustris assembly (N50 = 27.5 Mbp). Assembly contiguity, repeat content, and gene annotation compared well to other published quinaria group genomes. To demonstrate the utility of our hybrid assemblies, we performed whole genome synteny and phylogenomic comparisons within the quinaria clade. When compared to the chromosome-level Drosophila innubila assembly, the median size of syntenic blocks among quinaria group genomes was approximately 1 Mbp, although this estimate was influenced by assembly contiguity. Phylogenomic analyses estimated that the two major branches of the quinaria group diverged ∼17.9 MYA, with D. palustris and D. subpalustris diverging ∼1-2 MYA. We expect that our D. palustris and D. subpalustris assemblies will provide important additions to studies of genome evolution within Drosophilidae.
Splicing of transcripts via the spliceosome machinery is a complex process involving a multitude of proteins and short noncoding RNAs. In addition to full-length transcripts with all exons in the genomically encoded order, alternatively spliced transcripts can also be produced through alternative splicing of pre-mRNA. We have attempted to identify which genes are most frequently targeted by differential splicing using both downloaded and original RNA-Seq data from several species and tissue types. In order to identify splicing variation on an individual-to-individual basis, random contrasting of samples within compatible sets of samples was done. Both when analyzing sequence data from similar source material (same tissue from individuals kept under similar conditions) and when contrasting samples that are more different, such as different tissue types, genes associated with the splicing process itself seem to be the most recurring targets.
Heat stress challenges embryo survival, but the molecular reasons for this are unclear. We investigated how heat stress alters the maternal-to-zygotic transition (MZT) during Drosophila melanogaster development. Using RNA sequencing, we characterized the MZT under nonstress, acute, and chronic heat stress conditions. MZT genes were defined as those differentially expressed between the minor and major waves of zygotic genome activation. MZT genes were associated with multiple processes, including transcription, splicing, translation, and development. Under acute stress, the MZT deviated little from nonstressed embryos except that heat shock protein (HSP) genes were activated, and maternal transcript clearance was minimally misregulated. Under chronic stress, a core MZT persisted; but stress-related gene expression was obvious, maternal transcripts were strongly misregulated, and embryos showed poor survival. Overall, the MZT is robust, but increasingly falters with more severe stress. We present our dataset as a resource to aid molecular understanding of embryo resilience and vulnerability to environmental stressors.
Maize grain yield is frequently constrained by water scarcity, particularly in tropical regions characterized by irregular rainfall patterns. Dissecting the genetic basis of drought-related traits remains challenging because their expression is strongly influenced by environmental conditions. In this study, we applied a multi-environment multi-locus genome-wide association study (MEML-GWAS) to identify genomic regions associated with drought-related traits in tropical maize. The association panel comprised 190 inbred lines from the Embrapa breeding program, which were genotyped with 500,108 GBS-derived SNPs, and crossed with two tester lines. Phenotypic data corresponded to the performance of the testcross hybrids, divided in Dent and Flint heterotic groups, evaluated across two years at two locations in Brazil under well-watered and water-stressed conditions. Traits analyzed included grain yield, anthesis-silking interval, female and male flowering time, and plant and ear height. Drought stress reduced grain yield by approximately 50% and increased the anthesis-silking interval by about two days. A total of 179 significant SNP-trait associations were detected, of which 166 showed significant SNP-by-environment interaction effects, while 13 displayed stable effects across environments. Several associations were detected specifically under water-stressed conditions, highlighting genomic regions potentially involved in drought adaptation. Functional annotation revealed candidate genes previously implicated in abiotic stress responses, including ZmTIP1, which encodes an S-acyltransferase regulating root hair development and drought tolerance. Among the novel candidate genes, GRMZM2G159125, encoding a phospholipase D, emerged as a particularly promising candidate due to its strong association with grain yield and its role in membrane lipid signaling pathways related to stress responses. Although a few associations overlapped genomic regions previously reported for drought tolerance in maize, most loci represent potentially novel genetic factors that may contribute to improving drought resilience in tropical maize breeding programs.
Animal traits develop through intricate patterns of gene expression that are regulated at multiple levels. This regulation includes interactions between transcription factors and cis-regulatory element (CRE) DNA sequences, and the dynamic accessibility of CREs due to chromatin modifications and remodeling. Polycomb Group (PcG) and Trithorax Group (TrxG) genes are evolutionarily conserved regulators of chromatin state. Although the PcG and TrxG genes have well understood roles in developmental gene regulation, the extent to which these complexes contribute to trait evolution remains unclear. Here, we performed a genetic screen to understand how PcG and TrxG genes shape the rapidly evolving gene regulatory network (GRN) responsible for the dimorphic abdomen tergite pigmentation of Drosophila (D.) melanogaster fruit flies. A near-comprehensive screen of TrxG and PcG genes was conducted using RNAi, and numerous genes were identified whose reduced expression caused alterations to tergite pigmentation. For the eight most impactful genes, their roles in regulating this GRN were explored. We assessed their effects on several key transcription factors, and downstream CREs that are responsible for the expression of the GRN's pigmentation enzyme genes. We show that multiple members of the PcG and TrxG complexes are required for distinct tiers in the pigmentation GRN, suggesting potential points where responsive elements for different factors may have evolved. The results set the stage for future studies to identify the direct GRN targets of PcG and TrxG complexes, and how they have participated in the evolution of this GRN.
Transcriptional programs regulated by the KDM5 family of chromatin-modifying proteins are dysregulated in cancer and intellectual disability (ID) disorders. To define the fundamental mechanisms by which KDM5 regulates disease-relevant gene expression, we use missense variants in the X-linked KDM5C gene associated with the ID disorder Claes-Jensen syndrome (also known as KDM5C-NDD). Here, we use Drosophila melanogaster to investigate the effects of KDM5A224T, equivalent to human KDM5CA77T, which affects a conserved residue outside the catalytic histone demethylase JmjC domain and alters both enzymatic and non-enzymatic activities. Quantifying levels of H3K4me3, the demethylase substrate of KDM5, in adult brains revealed that Kdm5A224T induced changes indistinguishable from those observed with a catalytically inactive allele. This effect was not due to reduced promoter recruitment of the variant KDM5A224T protein. Instead, TurboID studies demonstrate that KDM5A224T exhibits reduced proximity with proteins involved in promoter activity and chromatin remodeling. Together, these findings show that KDM5-dependent transcriptional regulation cannot be explained by demethylase activity alone and support altered chromatin regulatory interactions as a key mechanism underlying pathogenic KDM5 variants.
Reference-quality genomes remain scarce for true crocodiles (Crocodylus), limiting comparative analyses of genome evolution and demographic history. Here, we generated and analyzed 2 long-read genomes, 1 for Crocodylus intermedius and 1 for C. niloticus, to investigate genome architecture, coalescent effective population size (Ne), and patterns of molecular evolution across crocodilians. Comparative analyses revealed broadly similar repeat landscapes in both species and extensive macro-synteny with Alligator sinensis, indicating strong structural conservation across crocodilian genomes. Using phased diploid assemblies and MSMC2, we reconstructed historical Ne trajectories and found marked differences between species. Crocodylus intermedius exhibited persistently low Ne throughout most of the late Quaternary, with a pronounced decline during the Late Pleistocene-early Holocene transition. In contrast, C. niloticus showed substantially larger Ne over comparable time intervals. Genome-wide codon-based analyses identified significant heterogeneity in dN/dS (ω) among crocodilian lineages. Crocodylus niloticus showed the lowest genome-wide ω, whereas elevated values in C. intermedius and other lineages were consistent with reduced long-term efficacy of purifying selection under smaller historical population sizes. Branch-site tests identified candidate genes under positive selection in both focal species, with functional categories related to ion transport, endocrine regulation, and cellular signaling. Together, these results provide genomic resources for Crocodylus and support an association between long-term demographic history and genome-wide patterns of molecular evolution across crocodilians.
Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.