
Drosophila nepalensis is a cold-adapted drosophilid endemic to the Himalayan region. Its ability to survive in harsh, cold conditions makes it a valuable Drosophila model for investigating how adaptation to thermal extremes may influence species persistence under future climate change. Here, we report the first de novo genome assembly of D. nepalensis, based on a hybrid sequencing strategy that combines Illumina short reads and Oxford Nanopore long reads. Illumina sequencing generated 49.88 million 150 bp paired-end reads (∼14.96 Gbp), while Nanopore sequencing produced 1.35 million long reads totaling ∼0.76 Gbp. The assembled genome spanned ∼178 Mb with an N50 of 83.6 kb and 98% BUSCO completeness, comparable to other well-annotated Drosophila genomes. Annotation identified 10,560 protein-coding genes, including transcription factor-rich and stress-related domains such as zinc fingers, WD40 repeats, and ankyrin motifs. Comparative orthology analysis across 6 Drosophila species identified 14,168 orthologous clusters, of which 9,173 were shared among all 6 species, indicating a conserved core genomic set across the sampled taxa. D. nepalensis showed 83 unique orthogroups and 50 singletons, suggesting some lineage-specific gene expansions associated with cold adaptation and endemicity, including families encoding caspase-family apoptotic regulators, chromatin remodeling proteins (HMGB/protamine-like), and SNARE-domain vesicle trafficking factors. Gene family evolution analysis revealed the highest expansions in the cold-tolerant Himalayan drosophilid, D. nepalensis, including significant expansions in serine protease, chaperone, and neurotransmitter transporter families, alongside dramatic contractions of core histone gene families, suggesting lineage-specific chromatin remodeling and ecological specialization.
Genetic perturbations are one of the great strengths of the model organism Drosophila melanogaster, with approaches such as classical mutagenesis and RNA interference enabling a wealth of biological discoveries. A more recent approach for altering gene expression is CRISPR/Cas9-based mutagenesis. As with any new tool, however, its use must be optimized. High expression of Cas9 has been shown to cause cytotoxicity in some cell types, including class IV dendritic arborization (da) neurons. In this study, we provide evidence that Cas9 expression causes cytotoxicity in class I da neurons, in addition to class IV da neurons, both of which are widely used to study neuronal development and regeneration. We then systematically evaluated available Cas9 transgenes designed to titrate Cas9 expression, called uCas9 transgenes. We show that the expression of these uCas9 transgenes results in little to no cytotoxicity in various classes of da neurons. Immunostaining revealed drastic reductions in Cas9 protein levels for da neurons expressing the uCas9(L) transgene. Lastly, we demonstrate that the uCas9(L) transgene effectively and specifically gene edits in both class I and class IV da neurons, lowering the expression of GFP-tagged proteins and producing loss-of-function morphological phenotypes when targeting endogenous loci. Thus, we refine the use of CRISPR mutagenesis in Drosophila da neurons through titration of Gal4/UAS-mediated Cas9 expression using existing uCas9 transgenes, a feasible and flexible approach that may be useful for other labs encountering Cas9 cytotoxicity in their own model systems.
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) continues to pose a threat to humans as well as domestic and wild animals. The variability in severity of clinical signs, the zoonotic potential, and the host-specific response to infection contribute to the persistence of circulation of disease. In wildlife species, white-tailed deer have been shown to be more permissive to infection than bovids. However, among bovids, American bison have shown a greater susceptibility than cattle. In this study, we investigate the transcriptomic response to experimental SARS-CoV-2 infection in bison over time. Substantial numbers of differentially expressed genes were identified between pre- and 2, 5, 7, 14, and 21 days post-inoculation. Kyoto Encyclopedia of Genes and Genomes and Gene Ontology term analysis identified associations with immune response, inflammatory response, and viral infection including COVID-19. Ingenuity Pathway Analysis of the coronavirus pathway highlighted differences in signaling at days 2 versus 21 post-inoculation. We additionally examined changes in the nasal microbiome of bison over the course of experimental infection, which suggested an increase in opportunity for secondary infection causing pathogens such as Mannheimia. Collectively, this study presents a profile of bison transcriptomic response to SARS-CoV-2 infection and continues to expand our understanding of variation in host response.
Metabolomics provides direct insight into physiological state, but for small organisms such as Drosophila melanogaster, it typically requires pooling individuals to obtain sufficient material. Pool sizes vary widely across studies with little justification, and the impact pooling and biological replication have on metabolomic characterization and signal detection remains poorly understood. We evaluated the effects of pool size and biological replication on metabolomic profiles and signal detection using two complementary designs in D. melanogaster. First, we tested how pooling (5, 50, or 100 individuals) affects metabolomic structure and reproducibility in inbred and outbred populations. Second, we tested how pool size interacts with replicate number to affect detection of diet-associated metabolite changes under a high-sugar perturbation. Pool size shaped metabolomic profiles: pools of five individuals consistently differed from larger pools, which improved reproducibility in a dataset-dependent manner. In the dietary experiment, smaller pools showed reduced sensitivity, detecting fewer true diet-associated metabolites without increasing false discoveries, and replicate downsampling showed that pool size and replication independently shape signal retention. Detection depended on effect size and variability: metabolites with larger, more stable effects were consistently retained, while smaller more variable effects were rapidly lost under reduced sampling. Beta-binomial modeling confirmed that detection probability reflects a balance between signal strength and measurement variability, with pool size and replicate number independently shaping this relationship. Together, these results show that metabolomic inference depends on the interplay of signal, noise, and sampling design, with pool size and replication jointly shaping the detectability, stability, and interpretation of biological signals.
Fruit texture is a key quality trait in cranberry (Vaccinium macrocarpon Ait.), with a direct impact on processing efficiency and overall profitability for the industry. To elucidate the genetic basis of cranberry texture, we conducted the first QTL mapping study in this crop implementing a fruit compression method with 10 mechanical texture traits in two biparental populations, CNJ02 (n=168) and CNJ04 (n=67), over two years. These 10 traits, associated with flesh hardness/stiffness and elasticity, showed similar phenotypic patterns of transgressive segregation and generally strong positive correlations. However, traits accounting for fruit size and shape effects (stress/strain-related traits and maximum contact pressure), together with slope-related parameters, exhibited consistently high heritability across populations, supporting their potential utility for breeding and genetic analyses. Genetic mapping revealed multiple consistently detected multi-trait QTL distributed across chromosomes 1,2,3,4,5,8,9,10,11 and 12, suggesting a complex genetic architecture. These genomic regions explained, on average, 4.67-18.6% of the phenotypic variance in CNJ02 and 10.97-27.75% in CNJ04. Candidate genes identified within these regions point to cell wall dynamics, including pectin modification, cellulose biosynthesis, and hemicellulose metabolism, as potentially relevant processes underlying cranberry fruit texture. Our findings provide loci in association with texture traits in cranberry and insights for the genetic improvement of this essential trait, informing the design of genotyping platforms and the implementation of genomic-assisted breeding approaches. Furthermore, we anticipate that these results will guide future studies aimed at optimizing postharvest fruit handling and management practices.
Pharmacogene missense variants can disrupt protein stability, catalytic competence, or substrate handling through distinct mechanisms. General-purpose predictors estimate clinical pathogenicity as a single scalar, whereas pharmacogene interpretation requires knowing which biochemical dimension a variant perturbs, since that determines whether reduced function is substrate-dependent. Five deep mutational scanning datasets comprising 26,198 missense variants across CYP2C9, CYP2C19, and NUDT15 were assembled from MaveDB. Paired assays showed that this dimensionality dominates the data: 28% of CYP2C9 variants (1,236 of 4,421) decoupled catalytic activity from abundance, and 48% of NUDT15 variants (1,364 of 2,844) decoupled thiopurine sensitivity from stability, with CYP2C9 discordance concentrating at substrate-channel residues. AlphaMissense, a representative general-purpose pathogenicity predictor, scored these classes in line with its clinical training objective rather than the assayed biochemistry, assigning likely-benign scores to 38 of 195 stable-but-dead CYP2C9 variants and likely-pathogenic scores to 140 of 222 destabilized but thiopurine-resistant NUDT15 variants. To test whether this dimensionality is recoverable, a supervised ESM-2 sequence baseline was benchmarked against the ESM1v zero-shot ensemble and AlphaMissense under position-based 5-fold cross-validation, together with three architectural extensions: AlphaFold structural features, multi-task learning across paired assays, and contact-graph neural networks. The baseline reached Pearson r of 0.54-0.72, matching or marginally exceeding both comparators, and no extension improved upon it. Trained directly on each assay, it nonetheless recovered the paired-assay difference at r = 0.28 for CYP2C9 and 0.43 for NUDT15, separating discordant variants at AUROC 0.60 and 0.51. Pharmacogene interpretation therefore requires assay-specific, substrate-aware functional measurements rather than a single generic score.
Tumor Suppressor Candidate 3 (TUSC3) is an integral component of the oligosaccharyltransferase complex and is required for the N-glycosylation of proteins. While TUSC3 has been implicated in cancer and autosomal recessive non-syndromic intellectual disability, some patients also exhibit distinct facial features. However, the role of Tusc3 in craniofacial development is unknown. Here, we report a patient presenting with cleft lip and palate who was found to carry a homozygous variant in the TUSC3 promoter region in addition to a maternally inherited chromosomal duplication. In order to evaluate a potential role for TUSC3 in craniofacial development, we analyzed Tusc3 deletion mouse mutants from the International Mouse Phenotyping Consortium (IMPC). Initial IMPC phenotyping data suggested that a subset of Tusc3 deletion mice could develop cleft palate, micrognathia, and aglossia. We further identified that Tusc3 is expressed in the craniofacial and brain regions during embryonic development. Upon characterizing the Tusc3 deletion line, we observed that most homozygous mutants exhibit preweaning lethality. While craniofacial defects occur at a very low frequency in the deletion mutants, we identified a modest reduction in cortical area. Furthermore, RNA-seq analysis surprisingly revealed that no other genes in the developing brain or face were significantly affected upon loss of Tusc3. These findings suggest that Tusc3 can contribute to congenital malformations in both craniofacial and cortical development.
The vine mealybug, Planococcus ficus, is a globally invasive pest of grapevine and a vector of leafroll viruses. Like other mealybugs, it reproduces through paternal genome elimination, a sex-determination system that operates without sex chromosomes and is associated with extreme sexual dimorphism. To characterize genome organization and sex-biased expression in this species, we generated a long-read reference genome spanning 369 Mb with 23,489 annotated genes and macrosynteny conserved with the citrus mealybug, Planococcus citri. Resequencing of four California field individuals provided the first genome-wide baseline of nucleotide diversity for P. ficus and 132 cross-validated microsatellite markers for population monitoring. Phloem feeders secrete proteins that mediate host interaction; of 2,129 predicted secreted proteins in P. ficus, a conserved core is shared with P. citri and each carries a lineage-specific set. Comparing adult male and female transcriptomes, we found sex-biased expression to be pervasive and skewed toward females: 41% of tested genes differed between the sexes, with female-biased genes both more numerous and showing larger fold changes. These female-biased genes were not randomly distributed but concentrated in discrete blocks of coordinately expressed, tandemly duplicated gene families, a pattern not previously described in a mealybug. Male- and female-biased secreted proteins also differed in origin, with male-biased proteins drawn from a conserved repertoire shared with P. citri and female-biased proteins spanning a more lineage-specific pool. Together, these results reveal a female-skewed, spatially clustered architecture of sex-biased expression in a mealybug that lacks sex chromosomes, and provide genomic resources for managing an invasive vineyard pest.
Splicing of transcripts via the spliceosome machinery is a complex process involving a multitude of proteins and short noncoding RNAs. In addition to full-length transcripts with all exons in the genomically encoded order, alternatively spliced transcripts can also be produced through alternative splicing of pre-mRNA. We have attempted to identify which genes are most frequently targeted by differential splicing using both downloaded and original RNA-Seq data from several species and tissue types. In order to identify splicing variation on an individual-to-individual basis, random contrasting of samples within compatible sets of samples was done. Both when analyzing sequence data from similar source material (same tissue from individuals kept under similar conditions) and when contrasting samples that are more different, such as different tissue types, genes associated with the splicing process itself seem to be the most recurring targets.
Heat stress challenges embryo survival, but the molecular reasons for this are unclear. We investigated how heat stress alters the maternal-to-zygotic transition (MZT) during Drosophila melanogaster development. Using RNA sequencing, we characterized the MZT under nonstress, acute, and chronic heat stress conditions. MZT genes were defined as those differentially expressed between the minor and major waves of zygotic genome activation. MZT genes were associated with multiple processes, including transcription, splicing, translation, and development. Under acute stress, the MZT deviated little from nonstressed embryos except that heat shock protein (HSP) genes were activated, and maternal transcript clearance was minimally misregulated. Under chronic stress, a core MZT persisted; but stress-related gene expression was obvious, maternal transcripts were strongly misregulated, and embryos showed poor survival. Overall, the MZT is robust, but increasingly falters with more severe stress. We present our dataset as a resource to aid molecular understanding of embryo resilience and vulnerability to environmental stressors.
Maize grain yield is frequently constrained by water scarcity, particularly in tropical regions characterized by irregular rainfall patterns. Dissecting the genetic basis of drought-related traits remains challenging because their expression is strongly influenced by environmental conditions. In this study, we applied a multi-environment multi-locus genome-wide association study (MEML-GWAS) to identify genomic regions associated with drought-related traits in tropical maize. The association panel comprised 190 inbred lines from the Embrapa breeding program, which were genotyped with 500,108 GBS-derived SNPs, and crossed with two tester lines. Phenotypic data corresponded to the performance of the testcross hybrids, divided in Dent and Flint heterotic groups, evaluated across two years at two locations in Brazil under well-watered and water-stressed conditions. Traits analyzed included grain yield, anthesis-silking interval, female and male flowering time, and plant and ear height. Drought stress reduced grain yield by approximately 50% and increased the anthesis-silking interval by about two days. A total of 179 significant SNP-trait associations were detected, of which 166 showed significant SNP-by-environment interaction effects, while 13 displayed stable effects across environments. Several associations were detected specifically under water-stressed conditions, highlighting genomic regions potentially involved in drought adaptation. Functional annotation revealed candidate genes previously implicated in abiotic stress responses, including ZmTIP1, which encodes an S-acyltransferase regulating root hair development and drought tolerance. Among the novel candidate genes, GRMZM2G159125, encoding a phospholipase D, emerged as a particularly promising candidate due to its strong association with grain yield and its role in membrane lipid signaling pathways related to stress responses. Although a few associations overlapped genomic regions previously reported for drought tolerance in maize, most loci represent potentially novel genetic factors that may contribute to improving drought resilience in tropical maize breeding programs.
Animal traits develop through intricate patterns of gene expression that are regulated at multiple levels. This regulation includes interactions between transcription factors and cis-regulatory element (CRE) DNA sequences, and the dynamic accessibility of CREs due to chromatin modifications and remodeling. Polycomb Group (PcG) and Trithorax Group (TrxG) genes are evolutionarily conserved regulators of chromatin state. Although the PcG and TrxG genes have well understood roles in developmental gene regulation, the extent to which these complexes contribute to trait evolution remains unclear. Here, we performed a genetic screen to understand how PcG and TrxG genes shape the rapidly evolving gene regulatory network (GRN) responsible for the dimorphic abdomen tergite pigmentation of Drosophila (D.) melanogaster fruit flies. A near-comprehensive screen of TrxG and PcG genes was conducted using RNAi, and numerous genes were identified whose reduced expression caused alterations to tergite pigmentation. For the eight most impactful genes, their roles in regulating this GRN were explored. We assessed their effects on several key transcription factors, and downstream CREs that are responsible for the expression of the GRN's pigmentation enzyme genes. We show that multiple members of the PcG and TrxG complexes are required for distinct tiers in the pigmentation GRN, suggesting potential points where responsive elements for different factors may have evolved. The results set the stage for future studies to identify the direct GRN targets of PcG and TrxG complexes, and how they have participated in the evolution of this GRN.
Transcriptional programs regulated by the KDM5 family of chromatin-modifying proteins are dysregulated in cancer and intellectual disability (ID) disorders. To define the fundamental mechanisms by which KDM5 regulates disease-relevant gene expression, we use missense variants in the X-linked KDM5C gene associated with the ID disorder Claes-Jensen syndrome (also known as KDM5C-NDD). Here, we use Drosophila melanogaster to investigate the effects of KDM5A224T, equivalent to human KDM5CA77T, which affects a conserved residue outside the catalytic histone demethylase JmjC domain and alters both enzymatic and non-enzymatic activities. Quantifying levels of H3K4me3, the demethylase substrate of KDM5, in adult brains revealed that Kdm5A224T induced changes indistinguishable from those observed with a catalytically inactive allele. This effect was not due to reduced promoter recruitment of the variant KDM5A224T protein. Instead, TurboID studies demonstrate that KDM5A224T exhibits reduced proximity with proteins involved in promoter activity and chromatin remodeling. Together, these findings show that KDM5-dependent transcriptional regulation cannot be explained by demethylase activity alone and support altered chromatin regulatory interactions as a key mechanism underlying pathogenic KDM5 variants.
Reference-quality genomes remain scarce for true crocodiles (Crocodylus), limiting comparative analyses of genome evolution and demographic history. Here, we generated and analyzed 2 long-read genomes, 1 for Crocodylus intermedius and 1 for C. niloticus, to investigate genome architecture, coalescent effective population size (Ne), and patterns of molecular evolution across crocodilians. Comparative analyses revealed broadly similar repeat landscapes in both species and extensive macro-synteny with Alligator sinensis, indicating strong structural conservation across crocodilian genomes. Using phased diploid assemblies and MSMC2, we reconstructed historical Ne trajectories and found marked differences between species. Crocodylus intermedius exhibited persistently low Ne throughout most of the late Quaternary, with a pronounced decline during the Late Pleistocene-early Holocene transition. In contrast, C. niloticus showed substantially larger Ne over comparable time intervals. Genome-wide codon-based analyses identified significant heterogeneity in dN/dS (ω) among crocodilian lineages. Crocodylus niloticus showed the lowest genome-wide ω, whereas elevated values in C. intermedius and other lineages were consistent with reduced long-term efficacy of purifying selection under smaller historical population sizes. Branch-site tests identified candidate genes under positive selection in both focal species, with functional categories related to ion transport, endocrine regulation, and cellular signaling. Together, these results provide genomic resources for Crocodylus and support an association between long-term demographic history and genome-wide patterns of molecular evolution across crocodilians.
Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.
Dung beetles serve as cultivators of their natural habitats, improving soil health and functions in both natural and anthropogenic environments. Despite their ecological importance, whole genome sequences for Scarabaeinae are limited. Here, we present the draft annotated genome assemblies for 2 temperate species of North American dung beetles collected from eastern Tennessee: Canthon chalcites and Phanaeus vindex. Both genome assemblies were generated from PacBio long reads and have high completeness, with BUSCO scores of 98.1% and 98.6% for C. chalcites and P. vindex, respectively. For C. chalcites, the BRAKER3 pipeline predicted 12,799 genes, and the gene set was 93.7% complete. For P. vindex, the BRAKER3 predicted 12,252 genes, and the gene set was 94.9% complete. From the annotated gene sets, orthologous protein sequence analyses among C. chalcites, P. vindex, the dung beetle species Onthophagus taurus, and the more evolutionarily distant beetle Tribolium castaneum indicated that there are 260 unique protein clusters for C. chalcites and 210 unique protein clusters for P. vindex. These 2 draft genomes provide valuable data for comparative genomics, evolution, and phylogenic studies for dung beetle species.
Brook Trout (Salvelinus fontinalis) are experiencing genomic erosion and demographic declines across the southern portion of their native distribution. A regionally representative reference genome is necessary to support conservation genomic initiatives for this species. While Brook Trout exhibit substantial phylogenetic structure across their range, only a single reference genome from Northeastern Canada (ASM2944872v1) is currently used to guide range-wide analyses. Consequently, reference bias is expected to cause erroneous sequence alignment and spurious variant detection for Brook Trout from divergent lineages. To prevent reference bias and invalid inference when studying Mid-Atlantic Brook Trout, we assembled a de novo, chromosome-level genome using a wild individual from Virginia, USA and produced a gene annotation with publicly available RNA-seq data. The assembly combined three complementary sequencing types (PacBio HiFi, Oxford Nanopore Ultra-Long, and Dovetail Omni-C) and benefited from manual curation using the Pretext software suite. The 2.91 Gb genome, known as mSalvFont1.0, consists of 3,333 contigs (N50 = 3.3 Mb) and 1,299 scaffolds (N50 = 55 Mb), with 81% of the assembly contained within 42 chromosome-scale scaffolds. Indices of accuracy and completeness reveal a quality value of 59 (99.999% accuracy), k-mer completeness of 95%, and 99.3% complete BUSCOs from the Actinopterygii orthologous gene set. Additionally, we identified reference bias (26.94% heterozygosity inflation) when variant detection of a Mid-Atlantic-origin sample relied on alignment to ASM2944872v1, instead of mSalvFont1.0. We anticipate that mSalvFont1.0 will increase the accuracy and precision of regional conservation genomic initiatives and expedite the genesis of a Brook Trout pangenome.
OptOrch is a transparent and modular optimization toolkit for forest tree seed orchard (SO) deployment. It implements SO-specific optimal contribution models in AMPL (A Mathematical Programming Language), separating biological assumptions, input data, and numerical solvers. Unlike existing optimal-contribution or mate-allocation software, OptOrch exposes the objective function and constraints directly, allowing breeders to modify SO-specific requirements, including status-number thresholds, proportional contribution bounds, graft availability, pairwise coancestry restrictions, female and male gametic contributions, and external pollen flow. When the declared formulation is supported by the selected solver, solutions can be evaluated using solver-reported feasibility status, objective bounds, and optimality gaps. The algebraic framework also allows alternative formulations to be tested when biological constraints create mixed-integer, nonlinear, or nonconvex instances. We demonstrate OptOrch with simulation-based scenarios involving alternative status-number thresholds and pollen contamination. The framework quantifies gain--diversity trade-offs, evaluates external pollen effects, and compares feasible deployment strategies under realistic biological and operational constraints.
Drosophila pseudoobscura is an historically important organism in evolutionary genetics, serving as a model system in studies of chromosomal inversions, speciation, sex chromosome evolution, and sex-ratio drive. However, previous population genetics analysis of D. pseudoobscura focused on individual chromosomes or used fragmented genome assemblies as a reference. To address these shortcomings, we generated a D. pseudoobscura population genomics resource consisting of newly sequenced genomes from 60 inbred lines sampled across the species' geographic range in North America. Using these data and a chromosome-scale reference genome, we examined patterns of nucleotide diversity and population structure across the chromosomes. We found no strong evidence of population structure on most chromosomes, consistent with prior results. In contrast, we identified population structure on the third chromosome, which we attributed to a well-characterized inversion polymorphism. We assigned individual third chromosome haplotypes to inversion arrangements, demonstrating how tests for population structure can be used to identify polymorphic chromosomal rearrangements. Tajima's D was negative across most of the genome, consistent with a recent population expansion. However, the distribution of genetic variation differed across third chromosome inversion arrangements in ways that were consistent with their hypothesized evolutionary histories, and we identified inter-arrangement genetic differentiation that could be attributed to the inversions suppressing genetic exchange. The population genomic data we have collected is publicly available and will support future research on evolutionary genetics.