
Production livestock provide a natural system for studying gene regulation under physiologically demanding conditions shaped by rapid growth, environmental exposure, and immune challenges. Using farm pigs from the PigGTEx resource, we apply quantile regression to reveal context-dependent genetic effects on gene expression across tissues. We identify quantile-specific eQTLs missed by standard linear models that preferentially localize to distal regulatory elements and three-dimensional genome architecture, in contrast to the promoter-proximal bias of canonical eQTLs. Genes with quantile-dependent eQTLs are more intolerant to loss-of-function variants and exhibit distinct patterns of enrichment in GO functional categories, indicating their likely functional significance. Cross-species comparisons reveal substantial overlap between pig and human eGenes across tissues, indicating conservation of regulatory architecture. Notably, many quantile-specific eQTLs influence tail expression states and involve genes relevant to human disease. For example, we identify a cis-eQTL affecting the conserved transcriptional regulator BCL6B in pig blood that modulates enhancer activity and reduces expression at lower quantiles. In contrast, BCL6B is minimally expressed in resting human blood and lacks detectable cis-regulatory variation under baseline conditions, consistent with its reported induction during immune activation. These findings demonstrate that pig eQTL maps can reveal context-dependent regulatory variation at loci that remain silent or weakly variable in human cohorts.
Heritability is the foundation of quantitative genetics, yet the way it is defined often conflicts with how it is estimated. On one hand, heritability represents the proportion of phenotypic variance explained by genetic variance in the population from which data are sampled. On the other hand, heritability estimated with family and pedigree data represents the heritability in a hypothetical base population where all individuals are assumed to be independent and non-inbred. The discrepancy may not be obvious to many people in the quantitative genetics community. This study demonstrates the discrepancy and, more importantly, introduces a pedigree sparsity coefficient (PSC) to correct the genetic variance/heritability from the base population to the current population. For very dense pedigree data, the correction may allow breeders to better understand the genetic basis of the trait of the current population and predict genetic gain for selection within the current population. The theory and method have been validated with simulated data and data collected from a long-term selection experiment in house mice. The PSC also applies to genomic heritability, where the pedigree relationship is replaced by a genomic relationship matrix calculated from genome-wide markers.
Malaria remains a significant global health challenge with an estimated 282 million cases reported in 2024. CRISPR/Cas9-based gene-drive systems have emerged as promising tools to block Plasmodium transmission by mosquito vectors. The TP13 drive system targets the Anopheles gambiae cardinal (Agcd) gene and carries two engineered monoclonal antibodies to achieve rapid population modification to prevent parasite transmission. Previous cage trials demonstrated complete drive introduction in three to six generations and supported modeling predicting a potential >90% reduction in malaria incidence under optimal conditions. However, naturally-occurring genetic polymorphisms in wild mosquito populations, particularly single nucleotide polymorphisms (SNPs) within Cas9/guide RNA target sites, pose a potential barrier to drive efficiency. High genetic diversity in An. gambiae results in drive-system target-site variants, including an A→T transversion in the Agcd gene, which occurs at high frequencies in African populations and could affect TP13 drive dynamics. The impact of this and other SNPs on TP13 performance was assessed by establishing three An. gambiae Ndokayo lines, one with the wild-type Agcd and two with homozygous SNP haplotypes. We evaluated drive conversion rates in vivo, population dynamics in cage trials, fitness costs and parasite suppression efficacy. No negative effects on drive performance and parasite suppression were observed. The results provide insights into the influence of naturally-occurring polymorphisms on gene drive propagation, informing safety, efficacy and target product profile requirements for advancing gene-drive mosquitoes toward field trials.
Severe disease following infection with SARS-CoV or SARS-CoV-2 is driven in part by genetically regulated immune responses that promote lung injury. To study genetic mechanisms of immune-mediated coronavirus pathogenesis, we screened the Collaborative Cross (CC) panel for Mus musculus strains that are phenotypically divergent with respect to coronavirus susceptibility. We identified CC006/TauUnc and CC044/UncJ as susceptible and resistant, respectively, and crossed them in an intercross to produce a genetic mapping population. Previously, we mapped genetic loci associated with disease severity, including HrS43, which is both conserved across viruses and between mouse and human. Here, we incorporate immune cell data from flow cytometry and perform a systems genetics analysis to resolve immune pathways linking host genetic loci to disease outcomes. We identify: (1) immune predictors of disease severity, as revealed by Bayesian variable selection; (2) extensive genetic regulation of the immune system at homeostasis and in response to infection, as revealed by infection-stratified and combined genotype-by-treatment (GxT) QTL mapping; and (3) candidate causal pathways linking genetic loci, immune responses, and disease severity, as revealed by context-dependent Bayesian mediation analysis. Our analysis reveals distinct patterns of virus specificity across biological layers: immune effects on disease severity appear largely consistent across viruses; genetic architecture is both shared and virus-specific; and mediating immune traits are in most cases virus-dependent. Notably, HrS43 appears to influence disease severity through distinct immune mediators in SARS-CoV versus SARS-CoV-2, demonstrating that conserved genetic susceptibility can drive virus-specific immunopathology with translational relevance across species.
Predicting phenotypes from genetics and environmental inputs is a long-standing challenge in genetics and plant breeding. Deep neural networks are a promising approach due to their capacity to approximate nonlinear biological processes. Despite initial expectations, recent studies have found that deep neural networks under-perform compared to linear methods, even on continent-scale datasets. We attribute this to several failure modes of deep learning, including greedy learning, the tendency to over-emphasize a single type of input data. As a solution, we present the Structured Interaction Neural Network (SINN), which combines statistical decomposition of genetic, environmental and interaction effects with deep neural networks. SINN dissects phenotype prediction into isolated component modeling tasks, revealing poor generalization of learned representations to new environments as the main limitation for prediction of genotype-by-environment interactions and overall yield. We reach competitive performance on yield prediction in the next cycle of 2 maize multi-environment trial datasets, including new genotypes and environments. In the Genomes to Fields (United States) dataset, SINN achieved higher overall accuracy (0.63) than BLUP-based methods (0.43) and a neural network from previous literature (0.48), and surpassed previous top-performing models with a lower RMSE (2.40 vs. 2.46 Mg/ha, mean yield 9.51 Mg/ha). Similar gains were observed in a secondary dataset consisting of Chinese national maize variety trials (MaizeGEP). SINN achieved higher overall accuracy (0.79) than BLUP-based methods (0.49) and a neural network from previous literature (0.76). By combining statistical genetics and modern deep learning, SINN enables accurate, modular and scalable genomic prediction in new environments.
Reductive stress has remained underappreciated as a significant disrupter of redox homeostasis. Recent studies have begun to link the accumulation of NADH and NADPH to the development and progression of metabolic diseases such as cancer, cardiac disease, and diabetes. In this study we use the nematode Caenorhabditis elegans to examine the phenomenon of catastrophic reductive-death caused by combined biguanide treatment and fasn-1 deficiency. This process of synergistic biguanide-induced reductive stress correlates with the activation of hypodermal stress response genes and aberrant alternations in the nucleolar morphology of hypodermal cells. Interestingly, we find that loss-of-function and RNAi-based knockdown of the catalytic RNA exosome subunit crn-3 significantly protects against reductive death. RNAi knockdown of multiple other genes involved in rRNA synthesis recapitulate this phenotype. We postulate that this reversal of reductive death can be attributed to impaired ribosomal RNA biogenesis that promotes tolerance of accumulated reducing equivalents NADPH and NADH while also preventing the accumulation of GSH by potentially activating downstream signaling pathways. Notably, we identify a downstream nuclear RNAi pathway that is activated by phenformin in a fasn-1 dependent manner and is also activated upon disruption of rRNA processing. Overall, we identify a novel mechanism by which pathologic states of reductive stress-related diseases could be ameliorated.
Innate immune responses triggered by Drosophila larval hemocytes have been extensively characterized. However, the full extent of transcriptional and post-transcriptional regulation underlying these processes remains poorly understood. Here, we employed a hybrid sequencing strategy integrating Oxford Nanopore long-read and Illumina short-read sequencing to provide a more comprehensive transcriptome annotation. This enabled discovery of full-length transcripts with novel 5' and 3' boundaries and uncovered 349 previously unannotated long non-coding RNAs highly induced during late stages of wasp infestation. To ensure high confidence transcript models, we further eliminated potential intra-priming artifacts specific to long-read cDNA data. This high-confident full-length transcript models helped to reveal cell type-specific lncRNA markers in lamellocytes and crystal cells by single-cell analyses, which recapitulated hemocyte differentiation trajectories. Notably, RNAi-based depletion of two highly induced lncRNAs impaired lamellocyte formation under wasp infestation, highlighting their functional relevance. Collectively, our findings provide detailed insights into the Drosophila larval immune transcriptomes through the long-read sequencing and highlight the regulatory roles of non-coding RNAs in innate immunity.
The crystal-Stellate interaction is the first known example of piRNA-mediated regulation between heterochromatin (the crystal/Suppressor of Stellate [Su(Ste)] locus on the Y chromosome) and euchromatin (the Stellate locus on the X chromosome) in Drosophila melanogaster. Here, we comprehensively annotate the Y-linked crystal/Su(Ste) sequences using Release 6 and heterochromatin-enriched scaffolds and contigs. In Release 6, we mapped two distinct crystal/Su(Ste) repeat regions with different molecular organizations. Beyond the ordered, tandem "classical" region, we identified a novel, disorganized region containing sequence fragments in both orientations derived from both crystal/Su(Ste) repeats and transposable elements. This disorganized locus displays structural features of a piRNA cluster that regulates both Stellate repeats and transposons. Furthermore, we identified at least four additional loci with organized crystal/Su(Ste)-like repeats within heterochromatin-enriched sequences not included in Release 6. We also conducted a detailed examination of the complex relationship between Release 6 sequences and more recent scaffolds and contigs. Alongside the classic crystal/Su(Ste) repeat containing hoppel/1360 fragments, we discovered two new repeat types structurally organized like crystal/Su(Ste) but carrying portions of HetA and gypsy12 transposons. Overall, these findings provide structural insights into the crystal-Stellate system, refining the Drosophila melanogaster Y chromosome map.
Mitochondrial biogenesis requires the coordinated synthesis, targeting, and import of nuclear-encoded mitochondrial precursor proteins. Although ribosome-associated chaperones support co-translational protein folding, their genetic contributions to mitochondrial protein import and cellular homeostasis remain incompletely defined. Here, we investigate the roles of the nascent polypeptide-associated complex (NAC) and the ribosome-associated Hsp70 system Ssb1/2 in Saccharomyces cerevisiae. We show that NAC and Ssb1/2 have distinct yet partially overlapping functions in the handling of mitochondrial precursor proteins. Loss of NAC activates the mitochondrial retrograde pathway and enhances growth on ethanol as a non-fermentable carbon source without compromising respiratory competence, indicating metabolic adaptation rather than overt mitochondrial dysfunction. In contrast, Ssb1/2 deficiency disrupts cytosolic proteostasis, sensitizes cells to TORC1 inhibition, and impairs autophagy and mitophagy. Using a TEV protease-based import reporter, we show that Ssb1/2 promotes efficient co-translational distribution of precursor proteins, whereas NAC limits the accumulation of misfolded proteins at the mitochondrial surface. Biochemical analyses further reveal that Ssb1/2 supports the association of translating cytosolic ribosomes with the mitochondrial outer membrane, while NAC loss partially restores this interaction in the absence of Ssb1/2. Together, these findings establish NAC and Ssb1/2 as key components of an integrated network linking co-translational targeting, mitochondrial signaling, and cellular homeostasis.
Meiotic recombination is a key driver of evolution in sexually reproducing species, reshaping genetic diversity by generating novel allelic combinations. The rate of recombination varies substantially across living organisms depending on cis- or trans-acting genetic elements, as seen in many species, including the yeast Saccharomyces cerevisiae. Here, we report on an experimental evolution-based study to better understand the factors shaping this natural variation. Starting with a genetically diverse population of S. cerevisiae, we have carried out recurrent divergent selection on recombination rate using a fluorescence-based sorting approach in four independent lineages. After ten generations, we observed an average response of recombination rate of +28% after positive selection and -24% after negative selection, within the interval used for selection. In the adjacent region, however, we observed a weaker response in the opposite direction, and no response in four other unlinked genomic regions. Whole-genome sequencing of individuals selected for high recombination revealed mixed outcomes in the four independently evolved lineages for high genome-wide recombination rates. However, all four lineages showed selection for high recombination locally, with particular haplotypes heavily favored and sequence- or structural variation-based heterozygosity selected against within the selection interval. Overall, this experimental evolution approach provides original and useful insights into the evolvability of the meiotic recombination rate and the associated genetic determinants.
Optimum contribution selection (OCS) balances genetic gain and inbreeding by optimizing parental contributions to the next generation, but current implementations rely on point estimates of breeding values that discard the uncertainty inherent in genetic evaluations. We introduce CVaR-OCS, a novel formulation that incorporates the full posterior distribution of estimated breeding values (EBVs) directly into the OCS objective via Conditional Value at Risk (CVaR) (a coherent risk measure from financial portfolio theory) allowing a single optimization to simultaneously maximize expected genetic gain and protect against worst-case outcomes driven by EBV uncertainty. We evaluate CVaR-OCS on a simulated multi-generation genomic selection dataset with known true breeding values (QTL-MAS 2010; n = 3,226), enabling direct comparison to an oracle solution, and on Norway spruce (Picea abies n = 5,525) and Loblolly pine (Pinus taeda n = 926) forest tree breeding progeny trials, with consistent tail-gain protection observed across all three datasets. On the simulated dataset, point-estimate MAP-OCS recovered only 78% of the genetic gain achieved by an oracle solution with access to true breeding values, illustrating the cost of ignoring predictive uncertainty; CVaR-OCS attained comparable expected gain while improving tail-gain security (CVaR95 +0.69%) and recovering an additional oracle-optimal individual. In Norway spruce, the recommended CVaR-OCS operating point improved tail-gain security by 6.60% and broadened the selection base from 145 to 159 individuals at a genetic gain cost of only 0.70%. Complementary MCMC-based robustness scores revealed that 25 MAP-OCS selections in Norway spruce were unstable across the posterior distribution; post-hoc exclusion of these individuals failed to improve tail-gain security, motivating the principled CVaR-OCS approach. CVaR-OCS provides breeders with a principled, computationally efficient tool for uncertainty-aware selection decisions, and multi-generation simulation studies are needed to fully characterize its long-term effects on genetic gain trajectories and inbreeding accumulation.
We study one of the simplest scenarios of polygenic selection that can be imagined: a subdivided population of diploid individuals expressing an additive trait under spatially homogeneous stabilizing selection. We are interested in the amounts of variation that can be maintained at mutation-selection-migration-drift equilibrium, at individual loci and at the level of the trait, within and among subpopulations. We derive analytical approximations for variance components and summary statistics such as FST and QST under the assumptions of the infinite-island model and compare these with individual-based simulations. We find that: (i) There is a critical migration threshold (which depends on effect sizes of trait loci) below which population structure strongly inflates genic variance in the subdivided population to levels well above those in a panmictic population. Variation within each subpopulation is maximized close to the critical migration rate. (ii) The genetic basis of trait variation across subpopulations is most similar close to this migration threshold and (counter-intuitively) becomes less similar for higher migration rates. This has consequences for the `portability' of Genome-Wide Association Studies (GWAS) between subpopulations, i.e, the extent to which loci with large contributions to variance in one subpopulation explain variance in other subpopulations. (iii) An analytical mean-field approach based on the single-locus diffusion approximation, together with effective migration and selection parameters (to account for associations between loci), very accurately predicts various quantities.
Neuropeptides modulate brain function by acting as ligands for a wide range of receptors. Systematic pharmacological screening efforts revealed extensive promiscuity in peptide-receptor pairs. However, the endogenous receptors and functions of promiscuous neuropeptides remain poorly understood. Here, we investigated how the Caenorhabditis elegans promiscuous RFamide-peptide FLP-14 and its receptors modulate a circuit for O2-evoked locomotory arousal. We show that FLP-14 peptides released from O2-sensing neurons and interneurons attenuate O2-evoked arousal by acting on three phylogenetically diverse G protein-coupled receptors (DMSR-3, FRPR-8, and NPR-40). Loss of each receptor individually confers strong hyperarousal by O2 comparable to the loss of flp-14. Distinct FLP-14 receptors have non-redundant roles in different target cells, including RID and RIM interneurons, that correlate antagonistically with locomotory arousal. Our results show that FLP-14 signaling requires coordinated modulation of multiple circuit nodes through distinct receptors in attenuating O2-evoked arousal. This promiscuous RFamide peptide network is reminiscent of monoaminergic neuromodulatory networks, which shape behavior through multiple receptors that differentially regulate distributed target neurons.
Gene expression data have been proposed as a natural single-cell lineage marker. Here, we critically examine the feasibility of reconstructing lineage trees from single-cell transcriptomic data using both modeling and empirical data. We first introduce a notion of neutrality for transcriptomic data, and then, under a model for neutral gene expression, establish theoretical bounds for accurate lineage tree reconstruction. Our findings indicate that reconstruction guarantees for even small trees or sub-trees require thousands of independent, neutral traits-a condition that is likely rarely met in practice due to the dominance of non-neutral developmental signals. Furthermore, errors introduced by measurement sampling have the potential to destroy any existing lineage signal. We conclude that gene expression data have limited potential as a natural lineage recorder and should not be used for phylogenetic lineage tree inference without further, rigorous validation.
In containing the first six-to-seven glycolytic enzymes within the organelle matrix, the glycosomes of trypanosomatid protists and their nearest relatives represent extreme forms of peroxisome specialisation. How or why such extreme peroxisomes evolved is not known. Many proteins are peroxisome targeted following recognition of a C-terminal type-1 peroxisome targeting signal (PTS1). Here, we identified in Naegleria gruberi, an amoeboflagellate distantly related to trypanosomatids, cryptic PTS1 motifs in a variety of metabolic enzymes, including in some glycolytic enzymes. These signals arise from stop-codon read-through or alternative splicing and are conserved in opportunistic pathogen N. fowleri. We show selected cryptic N. gruberi PTS1 motifs function in protein import into glycosomes of trypanosomatid Crithidia fasciculata. Further analysis revealed similar cryptic PTS1 motifs in evolutionarily diverse protists, including some from Discoba - the broad eukaryotic group to which Naegleria and trypanosomatids belong. Fragmentary data allowed only a cursory, equivocal glimpse of peroxisome biochemistry within Euglenida, the protists most closely related to trypanosomatids. Collectively, however, our data indicate Naegleria displays more versatility within its unusual metabolism than previously appreciated but moreover suggest dual protein localisation may have been used to diversify peroxisome function early in eukaryotic evolution and point towards at least two stages in glycosome evolution.
Notch mediates cell fate decisions in development and tissue homeostasis and can act as an oncogene or a tumor suppressor depending on the cell context. In my first essay for the GENETICS Perspectives series, "Notch and the Awesome Power of Genetics," I gave my perspective on developments in the field in the 20th century, from the isolation of the eponymous Notch mutations in Drosophila through the elucidation of the mechanism of signal transduction. That essay had a bildungsroman quality about it because the burgeoning of the Notch field and the increasing impact of Caenorhabditis elegans as a model organism coincided with my development as a scientist. But the "awesome power of genetics" was the guiding star of the essay because genetic approaches were largely responsible for identifying the core components of the signaling system and for deducing the novel cleavage mechanism by which Notch transduces signals. In this sequel, I offer my perspective as a geneticist on how the awesome power of genetics continued to have a starring role in the mechanistic discoveries that followed during the first quarter of the 21st century. I aim to show how genetics remains essential both in posing questions and in answering them even as the molecular events associated with activation of Notch signal transduction continue to be elucidated in exquisite structural detail.
Medical genetic studies focus on high-risk groups to enrich for targeted phenotypes. Although genetic effects may vary across the lifespan, age-related enrichment of target phenotypes can also arise from cumulative environmental exposures, without requiring changes in genetic architecture. Accumulating environmental effects over the lifetime can inflate nongenetic variance, creating the illusion of greater genetic signal while in fact masking underlying genetic effects. We propose investigating the genetic architecture of complex phenotypes in lower-risk, younger individuals, which may reduce residual variance and improve power to isolate the genetic underpinnings of complex disease-related traits. To illustrate this, we analyzed 9 complex traits across age-stratified White European cohorts in the UK Biobank. Genome-wide association analyses reveal a greater number of genome-wide significant SNPs in younger cohorts. Complementing these results, heritability analyses show a higher relative genetic contribution in younger cohorts. Together, these patterns demonstrate that genetic signals are more readily detected in lower-risk, younger individuals. This perspective suggests that temporal dynamics in genetic architecture should guide participant selection, challenging the traditional focus on high-risk older groups and highlighting the value of targeting higher-heritability groups for genetic discovery.
The impact of selection versus genetic drift on the evolution of mutation patterns is unclear. In Saccharomyces cerevisiae, which is predominantly diploid in nature, there is evidence that haploid cells have a higher mutation rate than diploids, suggesting that a haploid-specific mutator phenotype may have evolved due to the limited opportunity for selection to act on this rare cell type. Mutation in haploids was primarily elevated in late-replicating regions of the genome, implicating error-prone translesion synthesis (TLS) repair. Additional research has demonstrated that removing REV1, a gene responsible for initiating TLS, causes a reduction in haploid mutation rate. To assess whether the preferential use of this error-prone repair pathway by haploids explains the difference in genome-wide mutation patterns between cell types, we deleted REV1 in both diploid and haploid S. cerevisiae and estimated their mutation rates using a mutation accumulation experiment. Consistent with a previous study, we found a 50% higher single nucleotide mutation rate in REV1+ haploids than in REV1+ diploids. Deleting the REV1 gene caused this difference to vanish, with mutation rates in haploid and diploid rev1Δ lines converging on 2.4 × 10-10. Our results suggest that the mutagenic effect of translesion synthesis is much stronger in haploids, reflecting a limited opportunity for selection to act on mutation rates in rarer cells or smaller populations. We also find evidence that REV1 plays an important role in mitochondrial genome maintenance in both cell types.