Abstract The p150 isoform of the double-stranded RNA editing enzyme ADAR1 binds Z-DNA and Z-RNA through the conserved winged helix-turn-helix Zα domain. Here, we describe an inverse computational design strategy to map protein interactors of Zα. We used RFdiffusion and ProteinMPNN to generate ∼10,000 synthetic binders optimized for the Zα recognition surface, then used their sequences as structural templates for BLASTp searches against the human proteome. Multi-stage screening of ∼1,200 candidate regions from 298 proteins via ColabFold pDockQ identified 79 candidates for high-resolution AlphaFold3 modeling, which revealed the m6A reader YTHDC1 as the top-ranked interactor. AlphaFold3 predicts that a glutamate-rich poly-E disordered region of YTHDC1 (residues 199–254) docks into the basic recognition pocket of Zα through a charge-complementary mechanism that mimics the phosphate backbone of Z-RNA. Microsecond molecular dynamics simulations confirmed stability of the binary ADAR1p150–YTHDC1 complex, with the Zα–poly-E interface maintaining RMSD below 3 Å throughout. Ternary complex simulations showed that dsRNA acts as a co-anchoring scaffold stabilizing simultaneous engagement of both proteins in a catalytically dormant conformation. YTHDC1 localizes to transcription-associated YT bodies where nascent RNAs undergo m6A modification and negative supercoiling promotes Z-DNA formation, suggesting that YTHDC1 recruits ADAR1p150 to promote editing of intron-containing substrates prior to splicing. Short Statement ADAR1 can edit RNAs after they are made, changing the message they carry and removing double-stranded RNAs that can activate inflammatory responses. Using AI-driven protein design, we discovered that ADAR1 physically interacts with YTHDC1, a protein that recognizes newly made self-RNAs. The interaction allows ADAR to edit RNAs as they are made. This unexpected connection between two fundamental RNA modification systems opens new avenues for understanding and potentially targeting autoimmune disorders and cancers driven by the mis-editing of RNA transcripts.
While it is well established that the Zα domains of ADAR1 and ZBP1 proteins bind Z-form-prone nucleic acids (Z-NAs), it has also been shown that the Zα domain of ADAR1 binds DNA G-quadruplexes (GQs). However, no binding partner of the structurally homologous Zβ domain of ADAR1 has been identified to date. Based on AlphaFold and molecular dynamics simulations, it has recently been suggested that Zβ of ADAR1 targets its substrate by recognizing GQs. Here, we provide the first experimental evidence for Zβ binding to select GQ RNA and DNA in vitro, with structural specificity and nanomolar affinity. We also demonstrate that the Zα domains of ZBP1 and ADAR1 bind to both DNA and RNA GQs with similar affinity. These findings extend the range of potential functional roles for these proteins and open new hypotheses for testing in cells.
Unusual DNA structures and misfolded proteins Alan Herbert, Founder and President of InsideOutBio, discusses the evolution of biological concepts related to DNA and protein structures, highlighting how traditional models have evolved due to new data and discoveries. Much of modern life depends on precisely engineered, highly reliable machines. Such concepts originally shaped our view of biology, beginning with the famous ‘watchmaker’, whom William Paley anointed in 1802 as the creator of life. This concept of ideals was continued by the lock-and-key model of antibody-bound toxins proposed by Ehrlich in 1900, and by the modern view that proteins control the readout of genetic information through the highly specific recognition of nucleotide sequences. All these proposals assume that the world as we see it today has always been this way. They ignore the past worlds that existed before these innovations arose. Each model was based on limited data or on the selective use of that data. They failed when the volume of contradictory information generated by technological advances was overwhelming. These transitions occurred as the Linnean catalog of life forms in the 1700s expanded, as new chemical elements were discovered in the 1800s, and as the zoo of exotic subatomic particles was discovered in the early 1900s. A similar transformation is now occurring in biology, where the old concepts of just a few years ago no longer align with the data.
Triplexes (TRX) are a class of flipons that can form due to the interaction of RNA with B-DNA. While many proteins have been proposed to bind triplexes, structural models of these interactions do not exist. Here, I present AlphaFold V3 (AF3) models that reveal interactions between the high-mobility group protein B1 (HMGB1), HNRNPU (SAF-A), TP53, ARGONAUTE (AGO), and REL domain proteins. The TRXs result from the sequence-specific docking of RNAs to DNA via Hoogsteen base pairing. The RNA and DNA strands in apolar TRX are oriented in the opposite 5' to 3' direction, while copolar TRX have RNA and DNA strands pointing in the same 5' to 3' direction. TRXs can incorporate different RNA classes, including long noncoding RNAs (lncRNAs), short RNAs, such as miRNAs, piRNAs, and tRNAs, nascent RNA fragments, and non-canonical base triplets. Many pathways regulated by TRX formation have evolved to constrain retroelements (EREs), which are both an existential threat to the host and a source of genotypic variation. TRXs help set the boundaries of active chromatin, repressing the expression of most EREs, while depending on other flipons to modulate cellular programs. The TRXs help nucleate folding of intrinsically disordered proteins.
Deep learning methods have become methods of choice in the analysis of genomic data. The performance of deep learning models depends on the information available for training. A growing trend in deep learning applications involves leveraging multi-omics data - spanning genomics, transcriptomics, epigenomics, proteomics, metabolomics, and other domains. When a deep learning model trained on omics data achieves high performance, the important question is to define factors that contribute to model's predictive power. Explainable AI (XAI) methods can be categorized as model-aware and model-agnostic. Model-agnostic approaches, which rely on combinatorial feature perturbations to assess impact, are often computationally prohibitive for deep learning models. To address this, we developed OmiXAI, a pipeline integrating ensemble model-aware XAI methods for deep learning models trained on explicit omics feature matrices. Our framework incorporates gradient-based techniques - including Integrated Gradients, InputXGradients, Guided Backpropagation, and Deconvolution (for CNNs and GNNs) - as well as Saliency Maps and GNNExplainer (specifically for GNNs). We evaluated OmiXAI on a case study of functional genomic element prediction using interval-aligned epigenomic features, demonstrating its efficacy through feature importance analysis and benchmarking of XAI methods. Notably, OmiXAI enabled feature engineering, reducing the critical feature set from almost 2,000 to just 50. Its modular design allows seamless integration of additional attribution methods, ensuring adaptability beyond omics to diverse problem domains. While testing the ensemble approach we benchmarked individual XAI methods and discuss their drawbacks and limitations. OmiXAI is freely available at https://github.com/aameliig/OmiXAI.
Glioblastomas are the most prevalent primary brain tumors and are associated with a dramatically poor prognosis. Despite an intensive treatment approach, including maximal surgical tumor removal followed by radio- and chemotherapy, the median survival for glioblastoma patients has remained around 18 months for decades. Glioblastoma is distinguished by its highly complex mechanisms of immune evasion and pronounced heterogeneity. This variability is apparent both within the tumor itself, which can exhibit multiple phenotypes simultaneously, and in its surrounding microenvironment. Another key feature of glioblastoma is its “cold” microenvironment, characterized by robust immunosuppression. Recent advances in single-cell RNA sequencing have uncovered new promising insights, revealing previously unrecognized aspects of this tumor. In this review, we consolidate current knowledge on glioblastoma cells and its microenvironment, with an emphasis on their biological properties and unique patterns of molecular communication through signaling pathways. The evidence underscores the critical need for personalized poly-immunotherapy and other approaches to overcome the plasticity of glioblastoma stem cells. Analyzing the tumor microenvironment of individual patients using single-cell transcriptomics and implementing a customized immunotherapeutic strategy could potentially improve survival outcomes for those facing this formidable disease.
Genomic sequences that form three-stranded triplexes (TPXs) under physiological conditions (called T-flipons) play an important role in defining DNA nucleosome-free regions (NFRs). Within these NFRs, other flipon types can cycle conformations to actuate gene expression. The transcripts read from the NFR form condensates that engage proteins and small RNAs. The helicases bound then trigger RNA polymerase release by dissociating the 7SK ribonucleoprotein. The TPXs formed usually incorporate RNA as the third strand. TPXs made only from DNA arise mostly during DNA replication. Many small RNA types (sRNAs) and long noncoding (lncRNA) can direct TPX formation. TPXs made with circular RNAs have greater stability and specificity than those formed with linear RNAs. LncRNAs can affect local gene expression through TPX formation and transcriptional interference. The condensates seeded by lncRNAs are updated by feedback loops involving proteins and noncoding RNAs from the genes they regulate. Some lncRNAs also target distant loci in a sequence-specific manner. Overall, lncRNAs can rapidly evolve by adding or subtracting sequence motifs that modify the condensates they nucleate. LncRNAs show less sequence conservation than protein-coding sequences. TPXs formed by lncRNAs and sRNAs help place nucleosomes to restrict endogenous retroelement (ERE) expression. The silencing of EREs starts early in embryogenesis and is essential for bootstrapping development. Once the system is set, EREs play a different role, with a notable enrichment of Short Interspersed Nuclear Repeats (SINEs) in Enhancer–Promoter condensates. The highly programmable TPX-dependent processes create a chromaverse capable of many complexities.
Recent advances in Natural Language Processing (NLP) have spurred the application of Large Language Models (LLMs) to bioinformatics, enabling innovative approaches to DNA sequence encoding. However, genomic function is not solely determined by the primary DNA sequence—it also depends on complex, multi-layered omics data. This auxiliary data is inherently sparse, structured across multiple tracks, and poses a challenge for traditional unimodal approaches. Simultaneously, many bioinformatics tasks demand a unified signal source that encapsulates this information, rather than requiring researchers to input each omics feature individually. To overcome these limitations, we introduce a novel framework for the integrated representation of DNA sequences and their associated omics data. Central to our approach is a large-scale, batch-structured omics dataset optimized for deep learning at scale. Our framework is built around four key novelties: (1) a new model architecture that jointly processes DNA and omics signals; (2) an extension of Masked Language Modeling (MLM) to omics tracks for effective data reconstruction; (3) single-nucleotide embeddings that fuse all input modalities; and (4) interval-level embeddings that summarize broader genomic regions. We release a collection of pretrained models capable of reconstructing, embedding, and generating unified representations of DNA enriched with functional omics context. Code [https://github.com/aaai-2025-submission/anomymous\_submission\_aaai2025][1] Model and Datasets ### Competing Interest Statement The authors have declared no competing interest. [1]: https://github.com/aaai-2025-submission/anomymous_submission_aaai2025
This paper is focused on the origins of the contemporary genetic code. A novel explanation is proposed for how the mapping of nucleotides in DNA to amino acids in proteins arose that derives from repeat nucleotide sequences able to form alternative nucleic acid structures (ANS), such as the unusual left-handed Z-DNA, triplex, G-quadruplex and I-motif conformations. The scheme identifies sequence-specific contacts that map ANS repeats to dipeptide polymers (DPS). The stereochemistry required naturally evolves into a non-overlapping, triplet code for mapping nucleotides to amino acids. The ANS/DPS complexes form a simple, genetically transmitted, self-templating, autonomously replicating collection of ‘tinkers’ for Nature to evolve. Tinkers have agency and promote their own synthesis by forming catalytic scaffolds with metals, further enhancing their capabilities. Initial support for the model is provided by computational models built with AlphaFold3. The predictions made are properly falsifiable with the currently available methodology.
The double-stranded RNA editing enzyme ADAR1 connects two forms of genetic programming, one based on codons and the other on flipons. ADAR1 recodes codons in pre-mRNA by deaminating adenosine to form inosine, which is translated as guanosine. ADAR1 also plays essential roles in the immune defense against viruses and cancers by recognizing left-handed Z-DNA and Z-RNA (collectively called ZNA). Here, we review various aspects of ADAR1 biology, starting with codons and progressing to flipons. ADAR1 has two major isoforms, with the p110 protein lacking the p150 Zα domain that binds ZNAs with high affinity. The p150 isoform is induced by interferon and targets ALU inverted repeats, a class of endogenous retroelement that promotes their transcription and retrotransposition by incorporating Z-flipons that encode ZNAs and G-flipons that form G-quadruplexes (GQ). Both p150 and p110 include the Zβ domain that is related to Zα but does not bind ZNAs. Here we report strong evidence that Zβ binds the GQ that are formed co-transcriptionally by ALU repeats and within R-loops. By binding GQ, ADAR1 suppresses ALU-mediated alternative splicing, generates most of the reported nonsynonymous edits and promotes R-loop resolution. The recognition of the various alternative nucleic acid conformations by ADAR1 connects genetic programming by flipons with the encoding of information by codons. The findings suggest that incorporating G-flipons into editmers might improve the therapeutic editing efficacy of ADAR1.
1 While it is well established that the Zα domains of ADAR1 and ZBP1 proteins bind Z-form-prone nucleic acids (Z-NAs), it has also been shown that the Zα domain of ADAR1 binds DNA G-Quadruplexes (GQ). However, no binding partner of the structurally homologous Zβ domain of ADAR1 has been identified to date. Based on AlphaFold and molecular dynamics simulations, it has recently been suggested that the Zβ domain of ADAR1 targets its substrate by recognizing GQs. Here, we provide the first experimental evidence for Zβ domain binding to select G-quadruplex RNA and DNA in vitro , with structural specificity and low micromolar affinity. We also demonstrate that the Zα domains of ZBP1 bind to both DNA and RNA GQs with similar affinity. These findings extend the range of potential functional roles for these proteins and open new hypotheses for testing in cells.
Large language models (LLMs) in genomics have successfully predicted various functional genomic elements. While their performance is typically evaluated using genomic benchmark datasets, it remains unclear which LLM is best suited for specific downstream tasks, particularly for generating whole-genome annotations. Current LLMs in genomics fall into three main categories: transformer-based models, long convolution-based models, and state-space models (SSMs). In this study, we benchmarked three different types of LLM architectures for generating whole-genome maps of G-quadruplexes (GQ), a type of flipons, or non-B DNA structures, characterized by distinctive patterns and functional roles in diverse regulatory contexts. Although GQ forms from folding guanosine residues into tetrads, the computational task is challenging as the bases involved may be on different strands, separated by a large number of nucleotides, or made from RNA rather than DNA. All LLMs performed comparably well, with DNABERT-2 and HyenaDNA achieving superior results based on F1 and MCC. Analysis of whole-genome annotations revealed that HyenaDNA recovered more quadruplexes in distal enhancers and intronic regions. The models were better suited to detecting large GQ arrays that likely contribute to the nuclear condensates involved in gene transcription and chromosomal scaffolds. HyenaDNA and Caduceus formed a separate grouping in the generated de novo quadruplexes, while transformer-based models clustered together. Overall, our findings suggest that different types of LLMs complement each other. Genomic architectures with varying context lengths can detect distinct functional regulatory elements, underscoring the importance of selecting the appropriate model based on the specific genomic task. The code and data underlying this article are available at https://github.com/powidla/G4s-FMs.
Sequences called flipons can adopt discrete, alternative nucleic acid conformations, such as the left-handed Z-DNA and Z-RNA double helices (referred to collectively as ZNA), and the four-stranded RNA and DNA G-quadruplexes. Each flipon conformation encodes different information. For example, the base-specific interactions of proteins with B-DNA enable sequence-specific recognition. In contrast, the higher energy Z-DNA and G-quadruplexes facilitate the speedy scanning of chromosomes to locate active regions of the genome. Results synthesized from small-scale benchside and large-scale computational experimental approaches provide compelling evidence that zinc-finger protein domains (ZFDs) not only engage in base-specific recognition of B-DNA, but also bind directly to Z-DNA and G-quadruplexes. The findings address the long-standing speed-stability paradox of how high-affinity ZFPs with multiple zinc fingers can rapidly localize to a specific binding site. The energy gap between different DNA interaction modes enables fast off-rates during the scanning of Z-DNA for cognate binding sites, and a slow off-rate following engagement of the B-DNA conformer. ZFPs represent the most prominent human transcription factor family with 804 annotated members. The coevolution of flipons and ZFP enhances suppression of retroelements and enables rapid, context-specific responses. ZNA and GQ binding proteins are consequently more frequent in the proteome than currently conceded.
The role of alternative nucleic acid structures (ANS) in biology is an area of increasing interest. These non-canonical structures include the Z-DNA and Z-RNA duplexes (ZNA), the three-stranded triplex, the four-stranded G-quadruplex (GQ), and i-motifs. Previously, the biological relevance of ANS was dismissed. Their formation in vitro often required non-physiological conditions, and there was no genetic evidence for their function. Further, structural studies confirmed that sequence-specific transcription factors (TFs) bound B-DNA. In contrast, ANS are formed dynamically by a subset of repeat sequences, called flipons. The flip requires energy, but not strand cleavage. Flipons are enriched in promoters where they modulate transcription. Here, computational modeling based on AlphaFold V3 (AF3), under optimized conditions, reveals that known B-DNA-binding TFs also dock to ANS, such as ZNA and GQ. The binding of HLH and bZIP homodimers to Z-DNA is promoted by methylarginine modifications. Heterodimers only bind preformed Z-DNA. The interactions of TFs with ANS likely enhance genome scanning to identify cognate B-DNA-binding sites in active genes. Docking of TF homodimers to Z-DNA potentially facilitates the assembly of heterodimers that dissociate and are stabilized by binding to a cognate B-DNA motif. The process enables rapid discovery of the optimal heterodimer combinations required to regulate a nearby promoter.
G-quadruplexes (GQs) are non-canonical DNA structures encoded by G-flipons with potential roles in gene regulation and chromatin structure. Here, we explore the role of G-flipons in tissue specification. We present a deep learning-based framework for the genome-wide G-flipon predictions across 14 human tissue types. The model was trained using high-confidence experimental maps of GQ-forming sequences and ATAC-seq peaks, conjoined with the location of RNA polymerase, histone marks, and transcription factor binding sites. The training dataset for the DeepGQ model was derived from EndoQuad level 4-6 GQs. Model predictions were subsequently validated against the comprehensive EndoQuad dataset (levels 1-6) to optimize the whole-genome prediction threshold. To identify tissue-specific regulatory patterns, we classified GQ promoter predictions as either 'core' or 'tissue-specific'. We identified a notable overlap between predicted unique tissue-specific GQ sites and master regulatory genes (MRGs), tissue-specific DNase-hypersensitivity sites, and proteins that modulate R-loop formation. Collectively, the findings highlight the transactions between MRG and G-flipons intermediated by RNA: DNA hybrids associated with tissue specification.
Alan Herbert, Founder and President of InsideOutBio, discusses the significant advancements in RNA therapeutics, highlighting their role in supporting public health and their transformative potential in modern medicine, particularly for addressing genetic conditions and cancer. Currently, much attention is focused on healthy life-span extension. We forget that in 1900, the life expectancy for many countries was around 40, not the 70 years or more common today. As part of our cultural amnesia, we overlook the threat that pathogens have historically posed to killing individuals at a young age. The advances in life expectancy have come from improvements in water and sewage management, the elimination of insect vectors, such as fleas, lice, and mosquitoes, that transmit infectious agents, and from universal vaccination. The SARS-CoV pandemic was a grim reminder of that earlier time when 10-50% of the population died. In the current era, increasingly dense urbanization and rapid worldwide travel allow the rapid spread of similar human pathogens.
Recent findings have confirmed the long-held belief that alternative DNA conformations encoded by genetic elements called flipons have important biological roles. Many of these alternative structures are formed by sequences originally spread throughout the human genome by endogenous retroelements (ERE) that captured 50% of the territory before being disarmed. Only 2.6% of the remaining DNA codes for proteins. Other organisms have instead streamlined their genomes by eliminating invasive retroelements and other repeat elements. The question arises, why retain any ERE at all? A new synthesis suggests that flipons enable genomes to learn and programme the context-specific readout of information by altering the transcripts produced. The exchange of energy for information is mediated through changes in DNA topology. Here I provide a formulation for how genomes learn and describe the underlying p-bit algorithm through which flipons are tuned. The framework suggests new strategies for the therapeutic reprogramming of cells.
Herpes simplex virus 1 (HSV-1) and influenza A viruses (IAV) induce Z-form-nucleic-acid-binding protein 1 (ZBP1)-initiated cell death1-8. ZBP1 is activated by Z-RNA1,7,9, and the Z-RNAs that trigger ZBP1 during HSV-1 and IAV infections were assumed to be of viral origin1. Here, however, we show that host cell-encoded Z-RNAs are major and sufficient ZBP1-activating ligands after infection by these two human pathogens. The majority of cellular Z-RNAs mapped to intergenic endogenous retroelements embedded within abnormally long 3' extensions of host cell mRNAs. These aberrant host cell transcripts arose as a consequence of disruption of transcription termination (DoTT)-a virus-driven phenomenon that disables cleavage and polyadenylation specificity factor (CPSF)-mediated 3' processing of nascent pre-mRNAs10-15. Mutant viruses lacking ICP27 or NS1-the virus-encoded proteins responsible for inhibiting CPSF and triggering DoTT13,15-did not induce host cell Z-RNA accrual and were attenuated in their ability to stimulate ZBP1. Ectopic expression of HSV-1 ICP27 or IAV NS1 or pharmacological blockade of CPSF activity induced accumulation of host cell Z-RNAs and activated ZBP1. These results demonstrate that DoTT-generated cellular Z-RNAs are bona fide ZBP1 ligands, and position ZBP1-activated cell death as a host response to counter viral disruption of the cellular transcriptional machinery.