The interaction between brain structure and genetic influences is key to understanding neuropsychiatric disorders. However, most large-scale datasets are unimodal, providing either neuroimaging or genetics data. We propose CALM, a framework that learns interpretable associations between brain ROIs and genetic pathways from completely disjoint populations. CALM aligns the two modalities in a shared latent space via linear projections that simultaneously match the class-conditional latent distributions and ensure group separability. These projections provide interpretable pathway–ROI associations. When trained on unimodal imaging and genetics datasets, CALM generalizes to an unseen paired dataset, outperforming several state-of-the-art methods and ablation baselines. We also demonstrate stability of the learned associations against a paired baseline. Our experiments on autism spectrum disorder reveal immune and metabolic pathways linked to specific cortical regions and are consistent with established literature. Thus, CALM opens the door to leveraging large unimodal repositories for studying cross-modal interactions in brain disorders across disparate datasets.
Long-read sequencing (LRS) and diploid genome assembly have enabled nearly complete structural variant (SV) discovery. Using 293 nearly complete genomes, we characterize the full spectrum of genetic variation and show that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions. We identify 24 gene-rich regions subject to megabase-scale variation, 2,293 potentially unstable tandem repeats, and 890 novel expression quantitative trait loci associated with SVs in humans. Expanding to 1,218 LRS samples from the 1000 Genomes Project and applying a newly developed cross-platform breakpoint evaluation tool, BoostSV, we construct a nonredundant callset comprising 614,522 SVs. We demonstrate the utility of this population-level SV reference callset by filtering >99% of the common variation from 44 unsolved LRS probands from the Undiagnosed Diseases Network to discover likely disease-causing SVs. Second, we genotype 1,053 high-impact biallelic SVs from the pangenome callset in 232,090 samples from All of Us and discover 105 SVs with significant associations, including 26% where the SV is the lead variant. This publicly available pangenome SV resource will drive new disease associations and further our understanding of the missing heritability of human genetic disease.
The BioDIGS project is a nationwide initiative involving students, researchers and educators across more than 40 research and teaching institutions. Participants lead sample collection, computational analysis and results interpretation to understand the relationships between the soil microbiome, environment and health.
Motivation:Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. Results:In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. Availability:AniAnn's is open source software and available at github.com/marbl/anianns.
Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.
Motivation:Long-read metagenomic sequencing improves assembly contiguity and enables genome-resolved analysis of complex microbial communities, but accurate taxonomic classification of long reads and assembled contigs remains challenging. Highly scalable k-mer-based classifiers such as Kraken2 frequently over-assign fine-rank taxonomic labels when applied to long-read data, producing high false positive classification rates driven by sparse or localized k-mer matches, particularly in microbiomes with extensive taxonomic novelty. Results:We present Perseus, a lineage-aware confidence estimation framework for taxonomic classification that models the spatial distribution and hierarchical consistency of k-mer evidence along sequences. This formulation reframes taxonomic classification as a hierarchical confidence estimation problem rather than a single-rank prediction task. Perseus refines k-mer-level taxonomic signals from Kraken2 using a multi-headed convolutional neural network that estimates calibrated confidence scores for taxonomic correctness at each canonical rank. Using these estimates, Perseus confirms assignments, backs off to higher taxonomic ranks, or abstains when evidence is insufficient, prioritizing correctness and lineage consistency over overly specific assignments. Across simulations of taxonomic novelty and real-world metagenomic datasets, Perseus consistently and substantially reduces the false assignment rate while improving precision and lineage-consistent accuracy. These improvements are most pronounced for long reads and assembled contigs, where spatial context enables reliable discrimination between consistent taxonomic signal and spurious matches. Availability and implementation:Perseus integrates with existing Kraken2 workflows and is available at https://github.com/matnguyen/perseus.
Summary Genotype imputation remains essential for large-scale human genetics studies, but its performance is limited by the size and ancestral diversity of available reference panels, reducing accuracy for rare variants and underrepresented populations. Here, we present a cloud-based imputation service built on a multi-ancestry reference panel derived from 515,579 jointly phased genomes from the All of Us (N=414,830) and National Human Genome Research Institute’s Analysis, Visualization, and Informatics Lab-space (AnVIL, N=100,749) datasets. The All of Us + AnVIL reference panel is highly diverse and includes 261,163 participants most genetically similar to non-European reference populations, spanning 665,398,839 high-quality autosomal sites, representing a nearly 50% increase over TOPMed, the previous largest imputation service. Across multiple ancestry groups, the panel enables accurate imputation (empirical R² > 0.8) for variants with allele frequencies as low as 0.2%, extending reliable imputation into the rare-variant frequency spectrum, including allele frequencies down to 0.002% and 0.006% for samples with European ancestry and African ancestry in the United States, respectively. Compared with TOPMed, the panel improves imputation accuracy across all ancestry groups except Africans, and recovers additional trait-associated variants not represented in existing reference panels. To facilitate broad community access while preserving participant privacy, we deploy the panel through a secure cloud-based imputation platform using privacy-preserving recombined haplotypes. This resource establishes a new foundation for genome-wide association studies (GWAS) and fine-mapping, especially in previously underrepresented populations.
Conventional genome mapping-based approaches systematically overlook genetic variation, particularly in regions that substantially differ from the reference. To explore this hidden variation, here we examine unmapped and poorly mapped reads from the genomes of 640 human individuals from South Asian populations in the 1000 Genomes Project and the Simons Genome Diversity Project. We assemble tens of megabases of non-redundant sequence in tens of thousands of large contigs, a significant portion of which is present in both South Asian and other populations. We demonstrate that much of this sequence is not discovered by traditional variant discovery approaches even when using complete genomes and pangenomes. Across 20,000 placed contigs, we find 8215 intersections with 106 protein coding genes and over 15,000 placements within 1 kbp of a known GWAS hit. We use long-read data from a subset of samples to validate the majority of their assembled sequences, align RNA-seq data to identify hundreds of unplaced contigs with transcriptional potential, and query existing nucleotide databases to infer the origins of the remaining unplaced sequences. Our results highlight the limitations of even the most complete reference genomes and provide a model for understanding the distribution of hidden variation in any human population.
Neofunctionalization is a rare fate of gene duplication, classically defined as the acquisition of novel functions that potentiate the emergence of new traits. Rather than evolving to function autonomously, neofunctionalized genes may also remain embedded within their ancestral regulatory networks, potentially reshaping the genetic trajectories through which phenotypic change occurs. Testing this hypothesis, we leveraged a pan-genetic platform comprising ten Solanaceae species and show that a paralog of the flowering hormone florigen neofunctionalized into a flowering antagonist and was repeatedly selected during crop domestication and adaptation of wild plants across 50 million years of evolution. Independent selection of cis-regulatory and coding mutations in SELF-PRUNING 5G (SP5G) enabled rapid flowering in the wild ancestor of domesticated tomato from Central America as well as major and indigenous eggplant crop lineages domesticated in Asia and Africa. We further found that cis-regulatory sequence changes reduced SP5G expression and flowering time in wild species native to distinct environments in the Americas and Australia, relationships that we validated by genome editing. Together with similar patterns observed across diverse species and developmental networks, we propose that antagonistic neofunctionalized paralogs create evolutionary contingencies that channel adaptive trajectories across plant lineages.
Tumour prevalence varies dramatically throughout the animal kingdom despite broadly conserved cellular and developmental processes, raising the question of how evolution has shaped susceptibility 1,2. Here, we link macroevolutionary variation in tumour prevalence to gene-level selection by integrating comparative genomics data from 109 species of birds and mammals using a Bayesian phylogenetic framework to estimate pangenome-wide rates of genetic evolution across >150 million years of evolutionary change. We identify 3,206 genes in which natural selection is associated with shifts in tumour prevalence, with more than 80% of which are linked to reduced prevalence, suggesting pervasive selection for cancer suppression. Using causal phylogenetic inference, we show that genes associated with reduced tumour prevalence act predominantly through indirect effects on body size, revealing growth as a key mediator of cancer risk across species. In contrast, genes associated with increased tumour prevalence exert direct effects independent of body size. Finally, at the species-level, we demonstrate that exceptionally low rates of benign tumours do not necessarily coincide with reduced malignancy, revealing that benign and malignant tumour processes are evolutionarily decoupled. Together, these results reveal how natural selection has fine-tuned the link between genotype, phenotype, and cancer risk across species.
In-context learning (ICL) – the capacity of a model to infer and apply abstract patterns from examples provided within its input – has been extensively studied in large language models trained for next-token prediction on human text. In fact, prior work often attributes this emergent behavior to distinctive statistical properties in human language. This raises a fundamental question: can ICL arise organically in other sequence domains purely through large-scale predictive training? To explore this, we turn to genomic sequences, an alternative symbolic domain rich in statistical structure. Specifically, we study the Evo2 genomic model, trained predominantly on next-nucleotide (A/T/C/G) prediction, at a scale comparable to mid-sized LLMs. We develop a controlled experimental framework comprising symbolic reasoning tasks instantiated in both linguistic and genomic forms, enabling direct comparison of ICL across genomic and linguistic models. Our results show that genomic models, like their linguistic counterparts, exhibit log-linear gains in pattern induction as the number of in-context demonstrations increases. To the best of our knowledge, this is the first evidence of organically emergent ICL in genomic sequences, supporting the hypothesis that ICL arises as a consequence of large-scale predictive modeling over rich data. These findings extend emergent meta-learning beyond language, pointing toward a unified, modality-agnostic view of in-context learning.
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
Abstract Piwi-interacting RNAs (piRNAs) are small non-coding RNAs essential for transposon silencing and germline integrity across metazoans. In many species, piRNA expression is sexually dimorphic, yet the molecular mechanisms underlying this sex specificity remain poorly understood. In Caenorhabditis elegans, sexually dimorphic piRNA expression is regulated at the transcriptional level. We previously identified SNPC-1.3, a paralog of the small nuclear RNA (snRNA) activating protein complex (SNAPc/SNPC) subunit SNAPC1, as a male-specific piRNA transcription factor. However, the factors governing female piRNA expression remained elusive. Here, we identify SNPC-1.2, a second SNPC-1 paralog, as a female-specific piRNA transcription factor. SNPC-1.2 interacts with the core piRNA transcriptional machinery, binds female piRNA loci, is required for female piRNA expression, and promotes hermaphrodite fertility. In contrast, a third paralog, SNPC-1.1, retains the ancestral SNAPc function in snRNA transcription and is dispensable for piRNA biogenesis. Together, these findings reveal how gene duplication and functional specialization within the snpc-1 gene family generate specificity factors that direct the core SNAP complex to distinct genomic targets, providing a molecular mechanism for sexually dimorphic piRNA expression while maintaining canonical snRNA transcription.
Autism Spectrum Disorder standardized behavioral assessments provide quantitative measures of symptoms, yet their reliability and consistency have not been systematically evaluated. We present the first large-scale comparative analysis of four widely used assessments. We analyzed behavioral assessments across three autism cohorts using correlations, clustering, and diagnostic agreement analyses. We related behavioral variation to genetic and imaging data to evaluate biomarker associations. Sentence-level embeddings generated by large language models reveal substantial semantic overlap across instruments. Nonetheless, behavioral scores are weakly correlated (0.26 ± 0.21), and diagnostic classification shows only 65-80% agreement between tests. These patterns hold across three datasets comprising N = 1 954. None of the assessments show consistent associations with widely studied MRI or genetic biomarkers. These findings expose critical inconsistencies among widely used autism assessments and underscore the need for more reliable tools to support precision phenotyping, biomarker discovery, and individualized care. Rather than diminishing the utility of behavioral assessment in autism, the inconsistencies identified here highlight a critical opportunity to refine how behavioral phenotypes are defined and operationalized.
The Vertebrate Genomes Project (VGP) aims to produce complete and near-error-free reference genomes for all ~70,000 extant vertebrate species1. Organized in four phases, it progressively targets all vertebrate orders, families, genera, and eventually all species. Here we present the completion of VGP Phase I, delivering reference genomes for ~95% of vertebrate orders, along with additional lineages within those orders, totaling 816 species and 1.6 trillion base pairs of main haplotype sequence. These genomes were assembled and annotated over an 8-year period (2018-2026) of rapid advances in genome sequencing, assembly, and annotation methods2-4, alongside the growth of associated consortium initiatives and international collaborations5-9. They represent some of the highest-quality vertebrate genomes currently available, and most have become the primary reference for their respective species in public databases. Comparative analyses across a subset of 579 species when we reached a threshold of 85% of orders allowed us to reconstruct the genome of the last common ancestor of all vertebrates 500 million years ago, identify diverse modes of sex chromosome evolution, reveal clade-specific three-dimensional genome architecture, discover methylated epigenetic landscapes across vertebrates, and provide a framework for studying gene and pseudogene evolution, immune loci, cancer-associated genes, and other trait-associated loci. Approximately a quarter of this subset are listed as Vulnerable to Critically Endangered by the IUCN Red List of Threatened Species, and have enabled more advanced genomic investigations of extinction risk. VGP Phase I delivers a reference backbone for vertebrate genomics, enabling discoveries that would otherwise remain out of reach across evolution, conservation, and medicine.
Nanopore signal analysis enables detection of nucleotide modifications from native DNA and RNA sequencing, providing both accurate genetic or transcriptomic and epigenetic information without additional library preparation. At present, only a limited set of modifications can be directly basecalled (for example, 5-methylcytosine), while most others require exploratory methods that often begin with alignment of nanopore signal to a nucleotide reference. We present Uncalled4, a toolkit for nanopore signal alignment, analysis and visualization. Uncalled4 features an efficient banded signal alignment algorithm, BAM signal alignment file format, statistics for comparing signal alignment methods and a reproducible de novo training method for k-mer-based pore models, revealing potential errors in Oxford Nanopore Technologies' state-of-the-art DNA model. We apply Uncalled4 to RNA 6-methyladenine (m6A) detection in seven human cell lines, identifying 26% more modifications than Nanopolish using m6Anet, including in several genes where m6A has known implications in cancer. Uncalled4 is available open source at github.com/skovaka/uncalled4 .
The NHGRI Genomic Data Science Analysis, Visualization, and Informatics Lab-space (AnVIL) provides a secure cloud-based environment where research and education communities can analyze genomic and biomedical data. The platform supports a wide range of data analysis as well as the ability to safely store and access data in compliance with NIH policies. Work on the AnVIL platform can be easily shared to promote reproducible science and collaboration. The purpose of this study is to better understand the current user base of the AnVIL platform. The AnVIL Community Poll aimed to collect baseline information, identify development opportunities, guide the prioritization of user support strategies, and succinctly but comprehensively describe the current AnVIL Community. The AnVIL Team disseminated the inaugural AnVIL Community Poll by sharing it broadly on social media and through AnVIL and related consortia mailing lists. We categorized respondents as either returning or potential users of the AnVIL platform (based on their provided usage description) and examined user experiences: specifically user backgrounds, technological comfort, research interests, computational needs, and preferences for training and support. Our sample of the AnVIL community found opportunities for platform adoption beyond the current user base and identified areas where training should be enhanced, training preferences, and user computational needs. Specifically, while most respondents were involved in human genomics research, there may be potential for growth in adoption of the platform by prioritizing materials to support clinical researchers. All respondents felt availability of specific tools or datasets was a key feature of the platform. The broader community may also benefit from further development or showcasing of resources to facilitate cost management, finding and incorporating analysis tools, and data import. Our sample greatly preferred virtual training opportunities and returning users of the platform foresaw needing large amounts of storage. This poll provided an insightful snapshot of the current state of the AnVIL and demonstrated areas where the AnVIL Team can take specific steps to address barriers related to platform adoption and further support the existing and varied AnVIL Community. This work can be built upon through user interviews, community discussion, and coordinating a recurring poll.
Genetic mutation and drift, coupled with natural and human-mediated selection and migration, have produced a wide variety of genotypes and phenotypes in farmed animals. We here introduce the Farm Animal Genotype-Tissue Expression (FarmGTEx) Project, which aims to elucidate the genetic determinants of gene expression across 16 terrestrial and aquatic domestic species under diverse biological and environmental contexts. For each species, we aim to collect multiomics data, particularly genomics and transcriptomics, from 50 tissues of 1,000 healthy adults and 200 additional animals representing a specific context. This Perspective provides an overview of the priorities of FarmGTEx and advocates for coordinated strategies of data analysis and resource-sharing initiatives. FarmGTEx aims to serve as a platform for investigating context-specific regulatory effects, which will deepen our understanding of molecular mechanisms underlying complex phenotypes. The knowledge and insights provided by FarmGTEx will contribute to improving sustainable agriculture-based food systems, comparative biology and eventual human biomedicine.
Cancer is fundamentally a disease of the genome, characterized by extensive genomic, transcriptomic, and epigenomic alterations. Most current studies predominantly use short-read sequencing, gene panels, or microarrays to explore these alterations; however, these technologies can systematically miss or misrepresent certain types of alterations, especially structural variants, complex rearrangements, and alterations within repetitive regions. Long-read sequencing is rapidly emerging as a transformative technology for cancer research by providing a comprehensive view across the genome, transcriptome, and epigenome, including the ability to detect alterations that previous technologies have overlooked. In this Perspective, we explore the current applications of long-read sequencing for both germline and somatic cancer analysis. We provide an overview of the computational methodologies tailored to long-read data and highlight key discoveries and resources within cancer genomics that were previously inaccessible with prior technologies. We also address future opportunities and persistent challenges, including the experimental and computational requirements needed to scale to larger sample sizes, the hurdles in sequencing and analyzing complex cancer genomes, and opportunities for leveraging machine learning and artificial intelligence technologies for cancer informatics. We further discuss how the telomere-to-telomere genome and the emerging human pangenome could enhance the resolution of cancer genome analysis, potentially revolutionizing early detection and disease monitoring in patients. Finally, we outline strategies for transitioning long-read sequencing from research applications to routine clinical practice.
Pangenomes are growing in number and size, thanks to the prevalence of high-quality long-read assemblies. However, current methods for studying sequence composition and conservation within pangenomes have limitations. Methods based on graph pangenomes require a computationally expensive multiple-alignment step, which can leave out some variation. Indexes based on k-mers and de Bruijn graphs are limited to answering questions at a specific substring length k. We present Maximal Exact Match Ordered (MEMO), a pangenome indexing method based on maximal exact matches (MEMs) between sequences. A single MEMO index can handle arbitrary-length queries over pangenomic windows. MEMO enables both queries that test k-mer presence/absence (membership queries) and that count the number of genomes containing k-mers in a window (conservation queries). MEMO's index for a pangenome of 89 human autosomal haplotypes fits in 2.04 GB, 8.8x smaller than a comparable KMC3 index and 11.4x smaller than a PanKmer index. MEMO indexes can be made smaller by sacrificing some counting resolution, with our decile-resolution HPRC index reaching 0.67 GB. MEMO can conduct a conservation query for 31-mers over the human leukocyte antigen locus in 13.89 seconds, 2.5x faster than other approaches. MEMO's small index size, lack of k-mer length dependence, and efficient queries make it a flexible tool for studying and visualizing substring conservation in pangenomes.
Arthur L. Delcher合作论文数Department of Computer Science, Loyola University of Maryland14