Systematic mapping of protein-protein interaction (PPI) networks and determining how causal mutations rewire them in autism spectrum disorder (ASD) provide a powerful framework for uncovering disease mechanisms and therapeutic opportunities. Using affinity purification-mass spectrometry, we systematically mapped PPIs for 100 high-confidence ASD genes, uncovering more than 1800 interactions. By assessing the impact of pathogenic missense mutations, leveraging AlphaFold, and validating key findings in human-derived model systems, we identified marked convergence onto shared protein complexes in the wild-type state and convergent PPI rewiring driven by independent mutations. For example, distinct patient-derived variants in FOXP1 disrupt its interactions with FOXP4, leading to changes in cortical neurogenesis and neural activity in brain organoids. Overall, these findings link genetic variation to protein networks and convergent neurodevelopmental dysfunction in ASD.
The PEX1/PEX6 AAA-ATPase is required for the biogenesis and maintenance of peroxisomes. Mutations in HsPEX1 and HsPEX6 disrupt peroxisomal matrix protein import and are the leading cause of peroxisome biogenesis disorders. The most common disease-causing mutation in PEX1 is the HsPEX1G843D allele, which results in a reduction of peroxisomal protein import. Here, we demonstrate that the homologous yeast mutant, ScPex1G700D, reduces the stability of Pex1's active D2 ATPase domain and impairs assembly with Pex6 in vitro, but can still form an active AAA-ATPase motor. In vivo, ScPex1G700D exhibits only a slight defect in peroxisome import. We generated model human HsPEX1G843D cell lines and show that PEX1G843D is rapidly degraded by the proteasome, but that induced overexpression of PEX1G843D can restore peroxisome import. Additionally, we found that the G843D mutation reduces PEX1's affinity for PEX6, and that impaired assembly is sufficient to induce degradation of PEX1WT. Lastly, we found that fusing a deubiquitinase to PEX1G843D significantly hinders its degradation in mammalian cells. Altogether, our findings suggest a novel regulatory mechanism for PEX1/PEX6 hexamer assembly and highlight the potential of protein stabilization as a therapeutic strategy for peroxisome biogenesis disorders arising from the G843D mutation and other PEX1 hypomorphs.
Chromosome analysis is essential for diagnosing genetic disorders. For hematologic malignancies, identification of somatic clonal aberrations by karyotype analysis remains the standard of care. However, karyotyping is costly and time-consuming because of the largely manual process and the expertise required in identifying and annotating aberrations. Efforts to automate karyotype analysis to date fell short in aberration detection. Using a training set of ~10k patient specimens and ~50k karyograms from over 5 years from the Fred Hutchinson Cancer Center, we created a labeled set of images representing individual chromosomes. These individual chromosomes were used to train and assess deep learning models for classifying the 24 human chromosomes and identifying chromosomal aberrations. The top-accuracy models utilized the recently introduced Topological Vision Transformers (TopViTs) with 2-level-block-Toeplitz masking, to incorporate structural inductive bias. TopViT outperformed CNN (Inception) models with >99.3% accuracy for chromosome identification, and exhibited accuracies >99% for aberration detection in most aberrations. Notably, we were able to show high-quality performance even in "few shot" learning scenarios. Incorporating the definition of clonality substantially improved both precision and recall (sensitivity). When applied to "zero shot" scenarios, the model captured aberrations without training, with perfect precision at >50% recall. Together these results show that modern deep learning models can approach expert-level performance for chromosome aberration detection. To our knowledge, this is the first study demonstrating the downstream effectiveness of TopViTs. These results open up exciting opportunities for not only expediting patient results but providing a scalable technology for early screening of low-abundance chromosomal lesions.
Translating high-confidence (hc) autism spectrum disorder (ASD) genes into viable treatment targets remains elusive. We constructed a foundational protein-protein interaction (PPI) network in HEK293T cells involving 100 hcASD risk genes, revealing over 1,800 PPIs (87% novel). Interactors, expressed in the human brain and enriched for ASD but not schizophrenia genetic risk, converged on protein complexes involved in neurogenesis, tubulin biology, transcriptional regulation, and chromatin modification. A PPI map of 54 patient-derived missense variants identified differential physical interactions, and we leveraged AlphaFold-Multimer predictions to prioritize direct PPIs and specific variants for interrogation in Xenopus tropicalis and human forebrain organoids. A mutation in the transcription factor FOXP1 led to reconfiguration of DNA binding sites and altered development of deep cortical layer neurons in forebrain organoids. This work offers new insights into molecular mechanisms underlying ASD and describes a powerful platform to develop and test therapeutic strategies for many genetically-defined conditions.
The heterohexameric ATPases associated with diverse cellular activities (AAA)-ATPase Pex1/Pex6 is essential for the formation and maintenance of peroxisomes. Pex1/Pex6, similar to other AAA-ATPases, uses the energy from ATP hydrolysis to mechanically thread substrate proteins through its central pore, thereby unfolding them. In related AAA-ATPase motors, substrates are recruited through binding to the motor's N-terminal domains or N terminally bound cofactors. Here, we use structural and biochemical techniques to characterize the function of the N1 domain in Pex6 from budding yeast, Saccharomyces cerevisiae. We found that although Pex1/Delta N1Pex6 is an active ATPase in vitro, it does not support Pex1/ Pex6 function at the peroxisome in vivo. An X-ray crystal structure of the isolated Pex6 N1 domain shows that the Pex6 N1 domain shares the same fold as the N-terminal domains of PEX1, CDC48, and NSF, despite poor sequence conservation. Integrating this structure with a cryo-EM reconstruction of Pex1/Pex6, AlphaFold2 predictions, and biochemical assays shows that Pex6 N1 mediates binding to both the peroxisomal membrane tether Pex15 and an extended loop from the D2 ATPase domain of Pex1 that influences Pex1/Pex6 heterohexamer stability. Given the direct interactions with both Pex15 and the D2 ATPase domains, the Pex6 N1 domain is poised to coordinate binding of cofactors and substrates with
Motivation Structural variants (SVs) play a causal role in numerous diseases but can be difficult to detect and accurately genotype (determine zygosity) with short-read genome sequencing data (SRS). Improving SV genotyping accuracy in SRS data, particularly for the many SVs first detected with long-read sequencing, will improve our understanding of genetic variation.Results NPSV-deep is a deep learning-based approach for genotyping previously reported insertion and deletion SVs that recasts this task as an image similarity problem. NPSV-deep predicts the SV genotype based on the similarity between pileup images generated from the actual SRS data and matching SRS simulations. We show that NPSV-deep consistently matches or improves upon the state-of-the-art for SV genotyping accuracy across different SV call sets, samples and variant types, including a 25% reduction in genotyping errors for the Genome-in-a-Bottle (GIAB) high-confidence SVs. NPSV-deep is not limited to the SVs as described; it improves deletion genotyping concordance a further 1.5 percentage points for GIAB SVs (92%) by automatically correcting imprecise/incorrectly described SVs.Availability and implementation Python/C++ source code and pre-trained models freely available at https://github.com/mlinderm/npsv2.
We present a machine learning method capable of accurately detecting chromosome abnormalities that cause blood cancers directly from microscope images of the metaphase stage of cell division. The pipeline is built on a series of fine-tuned Vision Transformers. Current state of the art (and standard clinical practice) requires expensive, manual expert analysis, whereas our pipeline takes only 15 seconds per metaphase image. Using a novel pretraining-finetuning strategy to mitigate the challenge of data scarcity, we achieve a high precision-recall score of 94 del(5q) and t(9;22) anomalies. Our method also unlocks zero-shot detection of rare aberrations based on model latent embeddings. The ability to quickly, accurately, and scalably diagnose genetic abnormalities directly from metaphase images could transform karyotyping practice and improve patient outcomes. We will make code publicly available.
The heterohexameric AAA-ATPase Pex1/Pex6 is essential for the formation and maintenance of peroxisomes. Pex1/Pex6, similar to other AAA-ATPases, uses the energy from ATP hydrolysis to mechanically thread substrate proteins through its central pore, thereby unfolding them. In related AAA-ATPase motors, substrates are recruited through binding to the motor’s N-terminal domains or N-terminally bound co-factors. Here we use structural and biochemical techniques to characterize the function of the N1 domain in Pex6 from budding yeast, S. cerevisiae . We found that although Pex1/ΛN1-Pex6 is an active ATPase in vitro , it does not support Pex1/Pex6 function at the peroxisome in vivo . An X-ray crystal structure of the isolated Pex6 N1 domain shows that the Pex6 N1 domain shares the same fold as the N terminal domains of PEX1, CDC48, or NSF, despite poor sequence conservation. Integrating this structure with a cryo-EM reconstruction of Pex1/Pex6, AlphaFold2 predictions, and biochemical assays shows that Pex6 N1 mediates binding to both the peroxisomal membrane tether Pex15 and an extended loop from the D2 ATPase domain of Pex1 that influences Pex1/Pex6 heterohexamer stability. Given the direct interactions with both Pex15 and the D2 ATPase domains, the Pex6 N1 domain is poised to coordinate binding of co-factors and substrates with Pex1/Pex6 ATPase activity.
Biological age, distinct from an individual's chronological age, has been studied extensively through predictive aging clocks. However, these clocks have limited accuracy in short time-scales. Here we trained deep learning models on fundus images from the EyePACS dataset to predict individuals' chronological age. Our retinal aging clocking, 'eyeAge', predicted chronological age more accurately than other aging clocks (mean absolute error of 2.86 and 3.30 years on quality-filtered data from EyePACS and UK Biobank, respectively). Additionally, eyeAge was independent of blood marker-based measures of biological age, maintaining an all-cause mortality hazard ratio of 1.026 even when adjusted for phenotypic age. The individual-specific nature of eyeAge was reinforced via multiple GWAS hits in the UK Biobank cohort. The top GWAS locus was further validated via knockdown of the fly homolog, Alk, which slowed age-related decline in vision in flies. This study demonstrates the potential utility of a retinal aging clock for studying aging and age-related diseases and quantitatively measuring aging on very short time-scales, opening avenues for quick and actionable evaluation of gero-protective therapeutics.
Supplementary Table 1. Crystallographic Data and Refinement statistics. Supplementary Table 2. Comparing Target Inhibition by Quizartinib and PLX3397 in Engineered Ba/F3 cells Expressing FLT3-ITD, FLT3-ITD/F691L and FLT3-ITD/D835Y. Supplementary Table 3. Inhibitory Concentration (IC50) for Proliferation of Human Leukemia Cell Lines in PLX3397. Supplementary Table 4. 48 Hour Inhibitory Concentration (IC50) for Proliferation of Ba/F3 Cells Expressing Quizartinib and PLX3397-Resistant FLT3-ITD Mutant Isoforms. Supplementary Table 5. Patient Characteristics. Supplementary Table 6. Low Frequency FLT3 Kinase Domain Mutations of Uncertain Significance Observed at the Time of Resistance in Patient 1.14. Supplementary Figure 1. Composite omit map contoured at 1sigma level for the bound quizartinib and a structural water molecule. Supplementary Figure 2. Comparing the actual and previously predicted binding modes of quizartinib. Supplementary Figure 3. Structural model showing that D835 induces the DFG-in conformation of the activation loop, resulting in an orientation of F830 that precludes quizartinib binding. Supplementary Figure 4. Structural superposition of quizartinib (green) and PLX3397 (orange) highlighting the key structural difference responsible for their different susceptibilities to L691. Supplementary Figure 5. PLX3397 Inhibits FLT3 Signaling In Vitro. Supplementary Figure 6. Molm14 F691L Cells Demonstrate Resistance to Quizartinib. Supplementary Figure 7. Activity of PLX3397Against Quizartinib Resistance-Causing FLT3-ITD Kinase Domain Mutations.
Immunoglobulins (IGs), crucial components of the adaptive immune system, are encoded by three genomic loci. However, the complexity of the IG loci severely limits the effective use of short read sequencing, limiting our knowledge of population diversity in these loci. We leveraged existing long read whole-genome sequencing (WGS) data, fosmid technology, and IG targeted single-molecule, real-time (SMRT) long-read sequencing (IG-Cap) to create haplotype-resolved assemblies of the IG Lambda (IGL) locus from 6 ethnically diverse individuals. In addition, we generated 10 diploid assemblies of IGL from a diverse cohort of individuals utilizing IG-Cap. From these 16 individuals, we identified significant allelic diversity, including 36 novel IGLV alleles. In addition, we observed highly elevated single nucleotide variation (SNV) in IGLV genes relative to IGL intergenic and genomic background SNV density. By comparing SNV calls between our high quality assemblies and existing short read datasets from the same individuals, we show a high propensity for false-positives in the short read datasets. Finally, for the first time, we nucleotide-resolved common 5-10 Kb duplications in the IGLC region that contain functional IGLJ and IGLC genes. Together these data represent a significant advancement in our understanding of genetic variation and population diversity in the IGL locus.
We demonstrate early progress toward constructing a high-throughput, single-molecule protein sequencing technology utilizing barcoded DNA aptamers (binders) to recognize terminal amino acids of peptides (targets) tethered on a next-generation sequencing chip. DNA binders deposit unique, amino acid-identifying barcodes on the chip. The end goal is that, over multiple binding cycles, a sequential chain of DNA barcodes will identify the amino acid sequence of a peptide. Toward this, we demonstrate successful target identification with two sets of target-binder pairs: DNA-DNA and Peptide-Protein. For DNA-DNA binding, we show assembly and sequencing of DNA barcodes over six consecutive binding cycles. Intriguingly, our computational simulation predicts that a small set of semi-selective DNA binders offers significant coverage of the human proteome. Toward this end, we introduce a binder discovery pipeline that ultimately could merge with the chip assay into a technology called ProtSeq, for future high-throughput, single-molecule protein sequencing.
Background: Structural variants (SVs) play a causal role in numerous diseases but are difficult to detect and accurately genotype (determine zygosity) in whole-genome next-generation sequencing data. SV genotypers that assume that the aligned sequencing data uniformly reflect the underlying SV or use existing SV call sets as training data can only partially account for variant and sample-specific biases. Results: We introduce NPSV, a machine learning-based approach for genotyping previously discovered SVs that uses next-generation sequencing simulation to model the combined effects of the genomic region, sequencer, and alignment pipeline on the observed SV evidence. We evaluate NPSV alongside existing SV genotypers on multiple benchmark call sets. We show that NPSV consistently achieves or exceeds state-of-the-art genotyping accuracy across SV call sets, samples, and variant types. NPSV can specifically identify putative de novo SVs in a trio context and is robust to offset SV breakpoints. Conclusions: Growing SV databases and the increasing availability of SV calls from long-read sequencing make stand-alone genotyping of previously identified SVs an increasingly important component of genome analyses. By treating potential biases as a "black box" that can be simulated, NPSV provides a framework for accurately genotyping a broad range of SVs in both targeted and genome-scale applications.
Aptamers are single-stranded nucleic acid ligands that bind to target molecules with high affinity and specificity. They are typically discovered by searching large libraries for sequences with desirable binding properties. These libraries, however, are practically constrained to a fraction of the theoretical sequence space. Machine learning provides an opportunity to intelligently navigate this space to identify high-performing aptamers. Here, we propose an approach that employs particle display (PD) to partition a library of aptamers by affinity, and uses such data to train machine learning models to predict affinity in silico. Our model predicted high-affinity DNA aptamers from experimental candidates at a rate 11-fold higher than random perturbation and generated novel, high-affinity aptamers at a greater rate than observed by PD alone. Our approach also facilitated the design of truncated aptamers 70% shorter and with higher binding affinity (1.5 nM) than the best experimental candidate. This work demonstrates how combining machine learning and physical approaches can be used to expedite the discovery of better diagnostic and therapeutic agents.
Modern experimental technologies can assay large numbers of biological sequences, but engineered protein libraries rarely exceed the sequence diversity of natural protein families. Machine learning (ML) models trained directly on experimental data without biophysical modeling provide one route to accessing the full potential diversity of engineered proteins. Here we apply deep learning to design highly diverse adeno-associated virus 2 (AAV2) capsid protein variants that remain viable for packaging of a DNA payload. Focusing on a 28-amino acid segment, we generated 201,426 variants of the AAV2 wild-type (WT) sequence yielding 110,689 viable engineered capsids, 57,348 of which surpass the average diversity of natural AAV serotype sequences, with 12–29 mutations across this region. Even when trained on limited data, deep neural network models accurately predict capsid viability across diverse variants. This approach unlocks vast areas of functional but previously unreachable sequence space, with many potential applications for the generation of improved viral vectors and protein therapeutics. Viable AAV capsids are designed with a machine learning approach.
Despite a need for high-throughput protein sequencing, existing methods are not capable of sequencing the entire human proteome. We demonstrate early progress toward constructing the components of a protein sequencing technology that requires sequential capture and recording of binding over many cycles. The Barcode Cycle Sequencing (BCS) assay uses DNA-barcoded binders to construct a chain of barcodes directly onto a DNA sequencing chip that are used to identify a target. We demonstrated successful barcode capture and DNA-barcode chain sequencing for DNA-DNA and nanobody-protein binding pairs. For DNA-DNA binding, we showed assembly and sequencing of barcodes over 6 consecutive binding cycles. We additionally introduce a binder discovery pipeline called Target-Switch SELEX to discover aptamer binders. Ultimately, we envision merging BCS and Target-Switch SELEX into an end-to-end technology called ProtSeq, a high-throughput, single-molecule protein sequencing method in which DNA-barcoded aptamers read a protein’s amino acid sequence over multiple cycles of terminal amino acid degradation.
Qin Yang Aptitude Medical Systems Inc Ali Bashir Google Research Jinpeng Wang Stephan Hoyer Google Research Wenchuan Chou Cory McLean Google Research Geoff Davis Google Research Qiang Gong Zan Armstrong Google Research Junghoon Jang Hui Kang Annalisa Pawlosky Google Research Alexander Scott George E. Dahl Google Research Marc Berndl Google Research Michelle Dimon ( mdimon@google.com ) Google Research B. Scott Ferguson ( scott.ferguson@aptitudemedical.com )
Summary While next-generation sequencing (NGS) has dramatically increased the availability of genomic data, phased genome assembly and structural variant (SV) analyses are limited by NGS read lengths. Long-read sequencing from Pacific Biosciences and NGS barcoding from 10x Genomics hold the potential for far more comprehensive views of individual genomes. Here, we present MsPAC, a tool that combines both technologies to partition reads, assemble haplotypes (via existing software) and convert assemblies into high-quality, phased SV predictions. MsPAC represents a framework for haplotype-resolved SV calls that moves one step closer to fully resolved, diploid genomes. Availability and implementation https://github.com/oscarlr/MsPAC. Supplementary information Supplementary data are available at Bioinformatics online.
An incomplete ascertainment of genetic variation within the highly polymorphic immunoglobulin heavy chain locus (IGH) has hindered our ability to define genetic factors that influence antibody-mediated processes. Due to locus complexity, standard high-throughput approaches have failed to accurately and comprehensively capture IGH polymorphism. As a result, the locus has only been fully characterized two times, severely limiting our knowledge of human IGH diversity. Here, we combine targeted long-read sequencing with a novel bioinformatics tool, IGenotyper, to fully characterize IGH variation in a haplotype-specific manner. We apply this approach to eight human samples, including a haploid cell line and two mother-father-child trios, and demonstrate the ability to generate high-quality assemblies (>98% complete and >99% accurate), genotypes, and gene annotations, identifying 2 novel structural variants and 15 novel IGH alleles. We show multiplexing allows for scaling of the approach without impacting data quality, and that our genotype call sets are more accurate than short-read (>35% increase in true positives and >97% decrease in false-positives) and array/imputation-based datasets. This framework establishes a desperately needed foundation for leveraging IG genomic data to study population-level variation in antibody-mediated immunity, critical for bettering our understanding of disease risk, and responses to vaccines and therapeutics.