
Pangenome graphs capture extensive structural diversity, but resolving complex loci from shallow sequencing remains challenging, particularly when samples are of low quality such as in ancient DNA. We introduce COSIGT (COsine SImilarity-based GenoTyper), which assigns diploid genotypes by matching read-depth distributions to haplotype paths via cosine similarity. Because this metric evaluates relative coverage profiles rather than absolute read counts, COSIGT substantially outperforms existing likelihood-based tools at low coverage (1-2X). We demonstrate scalability to thousands of modern and ancient genomes, enabling robust, population-scale analyses of complex variation directly from low-coverage datasets.
Genome annotation is an important step in deriving functional meaning from prokaryotic sequencing data, yet systematic evaluations guiding tool selection are lacking. We present the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes. This includes Escherichia coli strains for baseline performance, thousands of archaea and bacteria genomes, as well as frameshifted and metagenome-assembled genomes. Bakta excels in annotating high-quality bacterial genomes, while PGAP was better for archaeal genomes and challenging bacterial assemblies, including metagenome-assembled, fragmented, or contaminated samples. For Gene Ontology annotation, PGAP consistently provides broader term coverage, whereas EggNOG-mapper offers more terms per feature. Our findings highlight tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin (e.g. MAGs). This study provides an evidence-based guide for users and informs future tool development.
Embryonic stem cells (ESCs) exhibit transcriptional heterogeneity in key pluripotency regulators, such as Nanog, suggesting functional roles for variability in cell state plasticity and regulation. While the importance of this variability is well-recognized, the specific molecular regulators that stabilize or destabilize these fluctuations remain poorly understood. We developed an integrative pipeline combining single-cell RNA sequencing (scRNA-seq) with ChIP-seq datasets to identify candidate regulators of gene expression variability in ESCs. This approach identified Nucleosome Assembly Protein 1-Like 1 (NAP1L1), a member of the NAP1 histone chaperone family, as a potential transcriptional stabilizer. Targeted recruitment of NAP1L1 to the Nanog promoter using a dCas9 fusion protein stabilized NANOG expression in ESCs. Genome-wide single-cell transcriptomic analysis of wild-type and Nap1l1-knockout ESCs revealed increased transcriptional variability upon NAP1L1 loss, supporting a global role in transcriptional stabilization. Co-immunoprecipitation followed by mass spectrometry identified Developmental Pluripotency Associated 3 (DPPA3) as a key interacting partner of NAP1L1, suggesting a mechanistic basis for its function. Our study establishes a systematic framework for discovering regulators of transcriptional variability and identifies NAP1L1 as a transcriptional stabilizer that modulates gene expression heterogeneity in ESCs.
Prioritizing T-cell receptor (TCR) candidates for defined peptide-HLA targets is an important step in TCR-based immunotherapy development, but it still relies heavily on laborious and expensive experimental screening. Recent advancements in generative artificial intelligence have demonstrated promising power in protein design and engineering. In this regard, we propose a pre-trained transformer model, termed Epitope-Receptor-Transformer (ERTransformer), for the epitope-conditioned generation of candidate TCR β-chain CDR3 sequences. ERTransformer is built on EpitopeBERT and ReceptorBERT, which are trained using 1.9 million epitope sequences and 33.1 million TCR sequences, respectively. To demonstrate the model capability, we generate 1,000 candidate TCR β-chain CDR3 sequences for each of the five epitopes with known natural TCRs. The generated candidates show low sequence similarity to natural TCR β-chains while retaining plausible CDR3 length, amino-acid composition, and conservative substitution patterns. We further conduct wet-lab experiments using flow cytometry in defined TCR/pMHC contexts and find that the level of T cell activation induced by selected artificial TCRs is either comparable to or even surpasses that of natural ones. Our work suggests that ERTransformer can expand and prioritize candidate TCR β-chain CDR3 sequences for downstream experimental screening in defined peptide-HLA and TCR-chain contexts.
Characterizing the extensive isoform diversity revealed by long-read RNA-sequencing remains challenging. After removal of technical artifacts, existing pipelines apply arbitrary expression thresholds that filter out bona fide transcript structures, obscuring diversity and hindering reproducibility. Instead of discarding isoforms, we propose a fundamentally distinct approach to quantifying isoform diversity using perplexity–the effective number of isoforms for a gene, derived from Shannon entropy–wherein every isoform, including low-abundance ones, contributes proportionally to a gene’s diversity. Analyzing 124 ENCODE4 PacBio datasets spanning 55 human cell types, we show that perplexity provides interpretable and reproducible isoform diversity measurements across genes, regulatory levels, and tissues.
Epistasis plays a crucial role in explaining missing heritability in complex traits, yet most detection methods fail to effectively capture local and global interaction patterns and long-range dependencies critical to complex trait architecture. We propose Epiformer that leverages genome language model Evo 2 to capture long-range dependencies from genetic background, and a dual-channel network to jointly model local and global epistasis and additive effects. Epiformer can identify key SNPs and their interactions from complex genomic data to empower phenotype prediction with interpretability, which reinforces epistasis detection. Epiformer performs robustly across species, revealing biologically meaningful patterns and offering new insights into genetic architecture.
Wild rice species harbour extensive but largely uncharacterised transcriptional diversity that underpins key agronomic traits and represents a valuable resource for crop improvement. Here we integrate single-molecule long-read sequencing (Iso-Seq) with short-read RNA-seq to resolve transcriptomes of Australian Oryza rufipogon-like (O. rufipogon Aus) and O. meridionalis, the closest and most phylogenetically distant AA-genome relatives of Asian domesticated rice, respectively. Iso-Seq reveals substantially greater transcriptome complexity than represented in current reference annotations, identifying 48,876 transcript isoforms in O. rufipogon Aus and 69,180 in O. meridionalis. More than 60
High-throughput chromatin assays require flexible workflows and context-aware parameter choices. However, unconstrained large language model-based analysis can suffer from inconsistent tool selection, parameterization, and execution. We present ChromSkills, a curated library of domain-specific analytical Skills for agentic chromatin data analysis on coding-agent platforms that support Skills. ChromSkills encodes expert decision logic and parameter-selection rules as modular, human-readable Skills linked to structured tool interfaces, enabling interpretable workflow composition and consistent execution from natural-language tasks. Across representative analyses, ChromSkills improved tool and parameter consistency, execution stability, and token efficiency, providing a transparent and domain-guided framework for AI-assisted chromatin data analysis.
Population-scale proteomics is driving precision medicine by enabling systematic drug target identification and robust biomarker discovery. Comparable, well-powered studies across diverse global populations are essential to elucidate shared and population-specific disease pathways. Mass-spectrometry-based and multiplexed affinity-based assays have emerged as complementary, leading strategies for quantifying proteomic variation in large-scale population-based studies. With the objective of performing a comprehensive comparative evaluation of the latest assays for each of these technologies, we compare three leading affinity-based and mass-spectrometry-based proteomic platforms (SomaScan11K, Olink Explore HT, Orbitrap Astral with Seer Proteograph [MS-Seer]) in a multi-ethnic Asian cohort, to inform biomarker discovery and functional genomic studies of global populations. We find limited overlap of 1,740 proteins out of 12,825 total proteins quantified across the three platforms, with modest correlations (0.34–0.10). SomaScan had lower missingness (< 1
Alzheimer’s Disease gene expression studies often rely on the reproducibility of known disease associations to build confidence in novel discoveries. This approach assumes that true biological variation associated with Alzheimer’s Disease is robust to a range of technical and biological confounders that likely differ across experiments. Here, we systematically assess the impact of an important yet frequently overlooked technical confounder in Alzheimer’s Disease differential expression analysis: changes in gene quantification resulting from ongoing updates to the human genome reference and transcriptome annotations. To investigate this, gene expression in a large brain transcriptomic dataset from Alzheimer’s Disease patients and controls (ROSMAP) was quantified using five transcriptome annotations mapped across a broad range of human genome references, from NCBI36 to CHM13v2.0. Although nearly all genes exhibit significant differences in expression between reference pairs (i.e., were differentially quantified), genome-wide Alzheimer’s Disease differential expression signatures remain highly concordant across references. Notably, while a core set of significantly differentially expressed Alzheimer’s Disease genes is largely conserved across references, reference-specific Alzheimer’s Disease-associated genes are enriched for pathways known to be involved in the disease. Genes whose direction of association with Alzheimer’s Disease significantly differs across references are also identified, highlighting the potential influence of reference choice on Alzheimer’s Disease-relevant genes and pathways. All results were replicated in an independent transcriptomics dataset (MSBB) derived from the same brain region. These findings demonstrate that reference genome choice impacts the identification and interpretation of specific Alzheimer’s Disease-associated genes, despite overall reproducibility of genome-wide differential expression signatures, underscoring the critical importance of reference selection in Alzheimer’s Disease transcriptomic studies.
Cytarabine (Ara-C) remains the cornerstone of induction therapy for acute myeloid leukemia (AML). However, its overall clinical efficacy is suboptimal, and the regimen is frequently accompanied by severe disruption of intestinal homeostasis. Consequently, identifying novel adjunctive therapies to improve treatment outcomes represents a critical clinical need. Gut microbiome profiling in AML mouse models reveals that Ara-C treatment significantly depletes Lactobacillus abundance. Through a probiotic intervention model, we demonstrate that administration of Lacticaseibacillus paracasei NCU-21 reduces leukemic burden, as well as size and weight of the spleen (p < 0.05). Mechanistically, L. paracasei NCU-21 restores microbial homeostasis, upregulates tight junction proteins (p < 0.05), and suppresses the lipopolysaccharide (LPS)-Toll-like receptor 4 (TLR4) pathway (p < 0.01), thereby attenuating systemic pro-inflammatory cytokine levels. Subsequent metabolomics confirms that indole-3-lactic acid (ILA) is the primary functional metabolite of L. paracasei NCU-21. In vivo administration of ILA recapitulates the anti-AML effects of the probiotic, and in vitro assays confirm that ILA directly induces AML cell apoptosis (p < 0.001) by suppressing the PI3K-AKT pathway. This work demonstrates that L. paracasei NCU-21 mitigates AML progression via the production of ILA, supporting its potential as an adjunct therapy.
Breast cancer risk is shaped by the vast heterogeneity of mammary epithelial cells, comprising basal, luminal progenitor, and mature luminal populations. While transcriptional variation among these lineages has been extensively studied, protein-level features—particularly in high-risk women—remain underexplored, limiting insight into early cellular and molecular determinants of susceptibility. Moreover, little is known about how clinical covariates influence clonogenic capacity, proteomic states, and epithelial proportions. We combine low-input proteomics with functional clonogenic assays to profile mammary epithelial cell subpopulations from a cohort of 22 breast tissues encompassing different germline mutation backgrounds, parity status and age. We quantify 5,555 proteins and observed marked inter-donor variation in epithelial composition, proteomic programs, and colony-forming capacity. Multivariable modeling reveals that clinical covariates—including age, parity, and germline mutation status—modulate both global proteomic architecture and lineage-specific pathway activity. Parity is associated with reduced basal cell abundance, altered luminal progenitor and mature luminal proteomes, and changes in clonogenicity. Pathway analyses identify both conserved and lineage-restricted responses to shared risk factors. Projection of clonogenic signatures onto METABRIC and TCGA tumors further links functional programs to tumor subtypes and molecular phenotypes. This study provides the most comprehensive proteomic atlas of cell-type resolved diversity in the high-risk breast to date. By defining how clinical covariates shape epithelial composition and molecular state, it clarifies key sources of biological variability that challenge controlled study design and offers a resource for improving mechanistic insight, risk assessment, and prevention strategies.
Recent development of mitochondrial DNA (mtDNA) editors has made it possible to generate experimental models carrying patient-relevant mtDNA mutations, characterize disease-associated phenotypes in cells and animals, and directly correct pathogenic variants in vivo. Here, we summarize the current landscape of mtDNA base editors and discuss key considerations for their application to mitochondrial disease model generation, including target selection, cross-species conservation, editing window constraints, strand bias, sequence-context dependence, and the balance between editing efficiency and safety. We further review strategies for validating mtDNA-edited models at functional, molecular, and organismal levels, and highlight recent studies demonstrating therapeutic rescue in animal models.
Cyclin-dependent kinase 12 (CDK12) has been identified as a susceptibility locus for kidney function, but its role in chronic kidney disease (CKD) remains unclear. We generated tubule-specific CDK12 knockdown and overexpression mice and establish CKD models via adenine-induced and unilateral ureteral obstruction. We assessed renal injury, lipid metabolism, and transcriptional alterations using histology, functional assays, full-length transcriptome sequencing, and mechanistic rescue experiments. We detected significant reduction of CDK12 expression in renal tubular epithelial cells in human patients and experimental chronic kidney disease models. We find tubule-specific CDK12 knockdown exacerbates renal dysfunction, fibrosis, and lipid accumulation, whereas CDK12 overexpression confers protection. Mechanistically, CDK12 deficiency induces intronic polyadenylation of NCEH1 (neutral cholesterol ester hydrolase 1), resulting in reduced NCEH1 expression and cholesteryl ester accumulation. Restoring NCEH1 partially rescues lipid dysregulation and renal injury, identifying it as a key downstream effector. This study reveals that CDK12 protects against CKD progression by suppressing NCEH1 intronic polyadenylation and maintaining lipid homeostasis. The CDK12-NCEH1 axis represents a previously unrecognised mechanism linking transcriptional regulation to renal lipotoxicity and fibrosis, and may provide a potential therapeutic target.
Large language models can learn new tasks through in-context learning (ICL), yet this ability remains underexplored for biological sequence classification. We evaluate ICL across 20 large language models on three antibody tasks: species-origin, antibody specificity, and isotype class classification. Few-shot prompting improves over zero-shot performance, but matching the performance of protein language model classifiers requires sequence-similar demonstrations. Building on this observation, we introduce a sequence similarity-based strategy for ICL in antibody sequence classification, Sim-ICL. Using 32-shot prompting, Sim-ICL achieves competitive performance on two of three tasks. Its simplicity makes few-shot ICL promising for antibody characterization, especially for researchers with limited coding expertise.
The protospacer adjacent motif (PAM) requirement limits CRISPR-Cas9 targetability. Here, we develop REPAM, an ancestral sequence reconstruction-based strategy that re-engineers PAM-interacting domains to expand PAM recognition. REPAM broadens the PAM specificities of Nme1Cas9 and Nme2Cas9 from N4GATT/N4CC to N4CNH, while additional mutations further extend Nme2Cas9 recognition to N4VHH. Applied to SpaCas9, REPAM expands PAM compatibility from NNGYRA to NNHH, increasing theoretical target coverage by approximately 36-fold. The engineered SpaCas9 achieves up to 76.6
Structural variants (SVs) in the human genome play an important role in health and disease. Identification of SVs is commonly performed using short-read sequencing. However, because sequenced fragments are typically shorter than the variants themselves, accurate SV calling remains a challenging problem. Previous benchmarking studies consistently find substantial variability in performance among SV calling tools, with disagreements largely driven by differences in underlying algorithms. We evaluate 14 high-performing SV calling tools using four samples from a newly published, high-confidence pedigree-based truth set, together with an in-house dataset. Its recent release minimizes the likelihood that it was used for training by tools, thereby reducing the risk of over-fitting. We implemented frequency filtering that reduced the downstream variants interpretation load by nearly half. We also assess alignments from a graph genome assembly, but found only small effect on performance, compared to a linear reference. DRAGEN achieves the highest overall performance, with a mean recall of 0.36 and mean F1 score of 0.51. Among the open-source tools, Manta and Dysgu achieve the highest recall, both with a mean recall of 0.26, while Manta achieves the highest mean F1 score of 0.41. We also evaluate several multi-caller ensemble strategies, which in some settings achieved higher recall than any individual tool. Across both ensemble strategies, Dysgu, Octopus, Manta, and Tardis are part of the top-performing sets. Our findings highlight the importance of designing SV calling strategies using one or more tools according to the intended application, while balancing accuracy, recall, computational resources, reference choice, and clinical interpretation burden.
Introgression between closely related species can profoundly influence evolutionary trajectories, yet how genomic divergence interacts with introgression, especially in invasive hybridization, remains insufficiently understood. Structural variants (SVs), particularly chromosomal inversions, are key contributors to divergence and may modulate patterns of gene flow. The Asian corn borer (ACB) and European corn borer (ECB), two globally important agricultural pests, provide an excellent system to investigate these processes. We construct a graph-based pangenome from 23 high-quality genome assemblies to characterize genome-wide SVs and introgression between ACB and ECB. We identify over 216,000 SVs, most of which are associated with transposable elements and contribute to interspecific divergence. Population genomic analyses reveal widespread yet asymmetric introgression, predominantly from ACB into ECB populations in China. Introgressed regions are more prevalent in autosomes than in sex chromosomes and are enriched for genes involved in adaptive pathways. In contrast, large inversion regions on the Z chromosome exhibit strong genetic differentiation and reduced introgression, suggesting a role in maintaining species barriers. Notably, a highly divergent region containing the circadian clock gene period (per) shows signatures of selection. Functional validation using CRISPR/Cas9 demonstrates that per significantly influences diapause regulation, linking genomic divergence to adaptive phenotypic variation. Our results demonstrate that introgression and structural variation jointly shape genomic divergence in these species. While introgression facilitates the spread of adaptive variation, chromosomal inversions restrict gene flow and maintain species integrity. This study provides new insights into the evolutionary mechanisms underlying divergence, adaptation, and invasion in major agricultural pests.
Petals are a key evolutionary innovation of flowers that reshape plant-pollinator interactions and contribute to the dominance of angiosperms in terrestrial ecosystems. However, their evolutionary origin remains debated as previous studies have largely relied on morphological inferences with limited spatially resolved molecular evidence during floral organogenesis. Here we employ spatial transcriptome sequencing on developing floral buds of the basal angiosperm Nymphaea colorata to dissect the potential gene regulatory landscape underlying petal formation. Our analyses reveal an unexpected hierarchical structure of four meristematic cell groups that persist at the floral base until carpel formation, displaying a gradient of decreasing meristematic potential along the proximal–distal axis. Among them, one group gives rise to stamens and petals (inner tepals). Developmental trajectory reconstruction further uncovers successive cell-fate transitions from stamen primordium to developing petals. These transitions are driven by quantitative variation in MADS-box tetramers composition that gradually reduces the reproductive identity of outer-whorl floral organs. Spatial mapping of developmental regulatory modules involving thousands of tissue-preferentially expressed genes demonstrate that petal morphogenesis integrates additional leaf-like genetic programs, ultimately distinguishing petals from stamens. Together, these results uncover the evolutionary derivation of petals from reproductive organs through recruitment of ancestral regulatory networks established before seed plants, providing new insights into the developmental basis and evolutionary innovation of petaloid organs in flowering plants.