Proteins function through dynamic interactions with other proteins in cells, forming complex networks fundamental to cellular processes. While high-resolution and high-throughput methods have significantly advanced our understanding of how proteins interact with each other, the molecular details of many important protein-protein interactions are still poorly characterized, especially in non-mammalian species, including Drosophila. Recent advancements in deep learning techniques have enabled the prediction of molecular details in various cellular pathways at the network level. In this study, we used AlphaFold2 Multimer to examine and predict protein-protein interactions from both physical and functional datasets in Drosophila. We found that functional associations contribute significantly to high-confidence predictions. Through detailed structural analysis, we also found the importance of intrinsically disordered regions in the predicted high-confidence interactions. Our study highlights the importance of disordered regions in protein-protein interactions and demonstrates the importance of incorporating functional interactions in predicting physical interactions between proteins. We further compiled an interactive web interface to present these predictions, facilitating functional exploration, comparative analysis, and the generation of mechanistic hypotheses for future studies.
Abstract Innate immunity has been dissected in exquisite detail in Drosophila melanogaster , yet how immune systems diversify between species remains largely unknown. Here we compare responses to bacterial infection across five drosophilid species spanning 60 million years of divergence, from D. melanogaster to Scaptodrosophila lebanonensis . Seven-day survival after Gram-negative infection ranges from 19% to 99% and tracks inversely with bacterial load in a phylogenetically corrected model, indicating that clearance rather than tolerance drives these differences. RNA sequencing of all five species after sterile wounding or infection reveals strongly species-biased transcriptional responses to the same pathogen, together with numerous uncharacterized lineage-restricted genes, including predicted antimicrobial peptides. We chemically synthesized candidate peptides and confirmed their activity in vitro: Athelas (CG43920), Mtkl and the S. lebanonensis -specific Athelas-like are active against bacteria and fungi, while Daisho2, previously described as antifungal, also kills Gram-positive and Gram-negative bacteria. De novo assembly and machine-learning prediction recover further candidate peptides from intronic and intergenic regions missed by current annotation. Rewiring of conserved genes and turnover of young, often unannotated effectors therefore act together to diversify antibacterial defense within a single insect family. One-sentence summary Five drosophilids over 60 million years defend against bacteria differently using lineage-specific young effectors and rewired conserved genes.
The regulation of gene expression is crucial for the functional integration of evolutionarily young genes, particularly those that emerge de novo. However, the regulatory programmes governing the expression of de novo genes remain unknown. To address this, we applied computational methods to single-cell RNA sequencing data, identifying key transcription factors probably instrumental in regulating de novo genes. We found that transcription factors do not have the same propensity for regulating de novo genes; some transcription factors regulate more de novo genes than others. Leveraging genetic and genomic tools in Drosophila, we further examined the role of two key transcription factors, achintya and vismay, and the regulatory architecture of new genes. Our findings identify key transcription factors associated with the expression of de novo genes and highlight how transcription factors, and possibly their duplications, are linked to the expressional regulation of de novo genes.
Recent studies reveal that de novo gene origination from previously non-genic sequences is a common mechanism for gene innovation. These young genes provide an opportunity to study the structural and functional origins of proteins. Here, we combine high-quality base-level whole-genome alignments and computational structural modeling to study the origination, evolution, and protein structures of lineage-specific de novo genes. We identify 555 de novo gene candidates in D. melanogaster that originated within the Drosophilinae lineage. Sequence composition, evolutionary rates, and expression patterns indicate possible gradual functional or adaptive shifts with their gene ages. Surprisingly, we find little overall protein structural changes in candidates from the Drosophilinae lineage. We identify several candidates with potentially well-folded protein structures. Ancestral sequence reconstruction analysis reveals that most potentially well-folded candidates are often born well-folded. Single-cell RNA-seq analysis in testis shows that although most de novo gene candidates are enriched in spermatocytes, several young candidates are biased towards the early spermatogenesis stage, indicating potentially important but less emphasized roles of early germline cells in the de novo gene origination in testis. This study provides a systematic overview of the origin, evolution, and protein structural changes of Drosophilinae -specific de novo genes.
LIM homeodomain transcription factor 1‐alpha (LMX1a) is a neuronal lineage‐specific transcription activator that plays an essential role during the development of midbrain dopaminergic (mDA) neurons. LMX1a induces the expression of multiple key genes, which ultimately determine the morphology, physiology, and functional identity of mDA neurons. This function of LMX1a is dependent on its homeobox domain. Here, we determined the structures of the LMX1a homeobox domain in complex with the promoter sequences of the Wnt family member 1 (WNT1) or paired like homeodomain 3 (Pitx3) gene, respectively. The complex structures revealed that the LMX1a homeobox domain employed its α3 helix and an N‐terminal loop to achieve specific target recognition. The N‐terminal loop (loop1) interacted with the minor groove of the double‐stranded DNA (dsDNA), whereas the third α‐helix (α3) was tightly packed into the major groove of the dsDNA. Structure‐based mutations in the α3 helix of the homeobox domain significantly reduced the binding affinity of LMX1a to dsDNA. Moreover, we identified a nonsyndromic hearing loss (NSHL)‐related mutation, R199, which yielded a more flexible loop and disturbed the recognition in the minor groove of dsDNA, consistent with the molecular dynamics (MD) simulations. Furthermore, overexpression of Lmx1a promoted the differentiation of SH‐SY5Y cells and upregulated the transcription of WNT1 and PITX3 genes. Hence, our work provides a detailed elucidation of the specific recognition between the LMX1a homeobox domain and its specific dsDNA targets, which represents valuable information for future investigations of the functional pathways that are controlled by LMX1a during mDA neuron development.
Mutations of the SNF2 family ATPase HELLS and its activator CDCA7 cause immunodeficiency, centromeric instability, and facial anomalies syndrome, characterized by DNA hypomethylation at heterochromatin. It remains unclear why CDCA7-HELLS is the sole nucleosome remodeling complex whose deficiency abrogates the maintenance of DNA methylation. We here identify the unique zinc-finger domain of CDCA7 as an evolutionarily conserved hemimethylation-sensing zinc finger (HMZF) domain. Cryo–electron microscopy structural analysis of the CDCA7-nucleosome complex reveals that the HMZF domain can recognize hemimethylated CpG in the outward-facing DNA major groove within the nucleosome core particle, whereas UHRF1, the critical activator of the maintenance methyltransferase DNMT1, cannot. CDCA7 recruits HELLS to hemimethylated chromatin and facilitates UHRF1-mediated H3 ubiquitylation associated with replication-uncoupled maintenance DNA methylation. We propose that the CDCA7-HELLS nucleosome remodeling complex assists the maintenance of DNA methylation on chromatin by sensing hemimethylated CpG that is otherwise inaccessible to UHRF1 and DNMT1.
Post-mating responses play a vital role in successful reproduction across diverse species. In fruit flies, sex peptide (SP) binds to the sex peptide receptor (SPR), triggering a series of post-mating responses. However, the origin of SPR predates the emergence of SP. The evolutionary origins of the interactions between SP and SPR and the mechanisms by which they interact remain enigmatic. In this study, we used ancestral sequence reconstruction, AlphaFold2 predictions, and molecular dynamics simulations to study SP-SPR interactions and their origination. Using AlphaFold2 and long-time molecular dynamics (MD) simulations, we predicted the structure and dynamics of SP-SPR interactions. We show that SP potentially binds to the ancestral states of Diptera SPR. Notably, we found that only a few amino acid changes in SPR are sufficient for the formation of SP-SPR interactions. Ancestral sequence reconstruction and MD simulations further reveal that SPR interacts with SP through residues that are mostly involved in the interaction interface of an ancestral ligand, myoinhibitory peptides (MIPs). We propose a potential mechanism whereby SP-SPR interactions arise from the pre-existing MIP-SPR interface as well as early chance events both inside and outside the pre-existing interface that created novel SP-specific SP-SPR interactions. Our findings provide new insights into the origin and evolution of SP-SPR interactions and their relationship with MIP-SPR interactions.
AbstractAlthough previously thought to be unlikely, recent studies have shown thatde novogene origination from previously non-genic sequences is a relatively common mechanism for gene innovation in many species and taxa. These young genes provide a unique set of candidates to study the structural and functional origination of proteins. However, our understanding of their protein structures and how these structures originate and evolve are still limited, due to a lack of systematic studies. Here, we combined high-quality base-level whole genome alignments, bioinformatic analysis, and computational structure modeling to study the origination, evolution, and protein structure of lineage-specificde novogenes. We identified 555de novogene candidates inD. melanogasterthat originated within theDrosophilinaelineage. We found a gradual shift in sequence composition, evolutionary rates, and expression patterns with their gene ages, which indicates possible gradual shifts or adaptations of their functions. Surprisingly, we found little overall protein structural changes forde novogenes in theDrosophilinaelineage. Using Alphafold2, ESMFold, and molecular dynamics, we identified a number ofde novogene candidates with protein products that are potentially well-folded, many of which are more likely to contain transmembrane and signal proteins compared to other annotated protein-coding genes. Using ancestral sequence reconstruction, we found that most potentially well-folded proteins are often born folded. Interestingly, we observed one case where disordered ancestral proteins become ordered within a relatively short evolutionary time. Single-cell RNA-seq analysis in testis showed that although mostde novogenes are enriched in spermatocytes, several youngde novogenes are biased in the early spermatogenesis stage, indicating potentially important but less emphasized roles of early germline cells in thede novogene origination in testis. This study provides a systematic overview of the origin, evolution, and structural changes ofDrosophilinae-specificde novogenes.
Mitochondria must maintain adequate amounts of metabolites for protective and biosynthetic functions. However, how mitochondria sense the abundance of metabolites and regulate metabolic homeostasis is not well understood. In this work, we focused on glutathione (GSH), a critical redox metabolite in mitochondria, and identified a feedback mechanism that controls its abundance through the mitochondrial GSH transporter, SLC25A39. Under physiological conditions, SLC25A39 is rapidly degraded by mitochondrial protease AFG3L2. Depletion of GSH dissociates AFG3L2 from SLC25A39, causing a compensatory increase in mitochondrial GSH uptake. Genetic and proteomic analyses identified a putative iron-sulfur cluster in the matrix-facing loop of SLC25A39 as essential for this regulation, coupling mitochondrial iron homeostasis to GSH import. Altogether, our work revealed a paradigm for the autoregulatory control of metabolic homeostasis in organelles.
Mutations of the SNF2 family ATPase HELLS and its activator CDCA7 cause immunodeficiency-centromeric instability-facial anomalies (ICF) syndrome, characterized by hypomethylation at heterochromatin. The unique zinc-finger domain, zf-4CXXC_R1, of CDCA7 is widely conserved across eukaryotes but is absent from species that lack HELLS and DNA methyltransferases, implying its specialized relation with methylated DNA. Here we demonstrate that zf-4CXXC_R1 acts as a hemimethylated DNA sensor. The zf-4CXXC_R1 domain of CDCA7 selectively binds to DNA with a hemimethylated CpG, but not unmethylated or fully methylated CpG, and ICF disease mutations eliminated this binding. CDCA7 and HELLS interact via their N-terminal alpha helices, through which HELLS is recruited to hemimethylated DNA. While placement of a hemimethylated CpG within the nucleosome core particle can hinder its recognition by CDCA7, cryo-EM structure analysis of the CDCA7-nucleosome complex suggests that zf-4CXXC_R1 recognizes a hemimethylated CpG in the major groove at linker DNA. Our study provides insights into how the CDCA7-HELLS nucleosome remodeling complex uniquely assists maintenance DNA methylation.
Vibrio parahaemolyticus must endure various challenging circumstances while being swallowed by phagocytes of the innate immune system. Moreover, bacteria should recognize and react to environmental signals quickly in host cells. Two-component system (TCS) is an important way for bacteria to perceive external environmental signals and transmit them to the interior to trigger the associated regulatory mechanism. However, the regulatory function of V. parahaemolyticus TCS in innate immune cells is unclear. Here, the expression patterns of TCS in V. parahaemolyticus -infected THP-1 cell-derived macrophages at the early stage were studied for the first time. Based on protein-protein interaction network analysis, we mined and analyzed seven critical TCS genes with excellent research value in the V. parahaemolyticus regulating macrophages, as shown below. VP1503, VP1502, VPA0021, and VPA0182 could regulate the ATP-binding-cassette (ABC) transport system. VP1735, uvrY, and peuR might interact with thermostable hemolysin proteins, DNA cleavage-related proteins, and TonB-dependent siderophore enterobactin receptor, respectively, which may assist V. parahaemolyticus in infected macrophages. Subsequently, the potential immune escape pathways of V. parahaemolyticus regulating macrophages were explored by RNA-seq. The results showed that V. parahaemolyticus might infect macrophages by controlling apoptosis, actin cytoskeleton, and cytokines. In addition, we found that the TCS (peuS/R) could enhance the toxicity of V. parahaemolyticus to macrophages and might contribute to the activation of macrophage apoptosis. IMPORTANCE This study could offer crucial new insights into the pathogenicity of V. parahaemolyticus without tdh and trh genes. In addition, we also provided a novel direction of inquiry into the pathogenic mechanism of V. parahaemolyticus and suggested several TCS key genes that may assist V. parahaemolyticus in innate immune regulation and interaction.
Group II introns are ribozymes that catalyze their self-excision and function as retroelements that invade DNA. As retrotransposons, group II introns form ribonucleoprotein (RNP) complexes that roam the genome, integrating by reversal of forward splicing. Here we show that retrotransposition is achieved by a tertiary complex between a structurally elaborate ribozyme, its protein mobility factor, and a structured DNA substrate. We solved cryo–electron microscopy structures of an intact group IIC intron-maturase retroelement that was poised for integration into a DNA stem-loop motif. By visualizing the RNP before and after DNA targeting, we show that it is primed for attack and fits perfectly with its DNA target. This study reveals design principles of a prototypical retroelement and reinforces the hypothesis that group II introns are ancient elements of genetic diversification.
Proteins are the building blocks for almost all the functions in cells. Understanding the molecular evolution of proteins and the forces that shape protein evolution is essential in understanding the basis of function and evolution. Previous studies have shown that adaptation frequently occurs at the protein surface, such as in genes involved in host-pathogen interactions. However, it remains unclear whether adaptive sites are distributed randomly or at regions associated with particular structural or functional characteristics across the genome, since many proteins lack structural or functional annotations. Here, we seek to tackle this question by combining large-scale bioinformatic prediction, structural analysis, phylogenetic inference, and population genomic analysis of Drosophila protein-coding genes. We found that protein sequence adaptation is more relevant to function-related rather than structure-related properties. Interestingly, intermolecular interactions contribute significantly to protein adaptation. We further showed that intermolecular interactions, such as physical interactions, may play a role in the coadaptation of fast-adaptive proteins. We found that strongly differentiated amino acids across geographic regions in protein-coding genes are mostly adaptive, which may contribute to the long-term adaptive evolution. This strongly indicates that a number of adaptive sites tend to be repeatedly mutated and selected throughout evolution in the past, present, and maybe future. Our results highlight the important roles of intermolecular interactions and coadaptation in the adaptive evolution of proteins both at the species and population levels.
Studying how novel phenotypes originate and evolve is fundamental to the field of evolutionary biology as it allows us to understand how organismal diversity is generated and maintained. However, determining the basis of novel phenotypes is challenging as it involves orchestrated changes at multiple biological levels. Here, we aim to overcome this challenge by using a comparative species framework combining behavioral, gene expression, and genomic analyses to understand the evolutionary novel egg-laying substrate-choice behavior of the invasive pest species Drosophila suzukii. First, we used egg-laying behavioral assays to understand the evolution of ripe fruit oviposition preference in D. suzukii compared with closely related species D. subpulchrella and D. biarmipes as well as D. melanogaster. We show that D. subpulchrella and D. biarmipes lay eggs on both ripe and rotten fruits, suggesting that the transition to ripe fruit preference was gradual. Second, using two-choice oviposition assays, we studied how D. suzukii, D. subpulchrella, D. biarmipes, and D. melanogaster differentially process key sensory cues distinguishing ripe from rotten fruit during egg-laying. We found that D. suzukii's preference for ripe fruit is in part mediated through a species-specific preference for stiff substrates. Last, we sequenced and annotated a high-quality genome for D. subpulchrella. Using comparative genomic approaches, we identified candidate genes involved in D. suzukii's ability to seek out and target ripe fruits. Our results provide detail to the stepwise evolution of pest activity in D. suzukii, indicating important cues used by this species when finding a host, and the molecular mechanisms potentially underlying their adaptation to a new ecological niche.
In a typical biomolecular simulation using Atomic Multipole Optimized Energetics for Biomolecular Applications (AMOEBA) force field, the vast majority molecules in the simulation box consist of water, and these water molecules consume the most CPU power due to the explicit mutual induction effect. To improve the computational efficiency, we here develop two new nonpolarizable water models (with flexible bonds and fixed charges) that are compatible with AMOEBA solute: the 3‐site AW3C and 5‐site AW5C. To derive the force‐field parameters for AW3C and AW5C, we fit to six experimental liquid thermodynamic properties: liquid density, enthalpy of vaporization, dielectric constant, isobaric heat capacity, isothermal compressibility and thermal expansion coefficient, at a broad range of temperatures from 261.15 to 353.15 K under 1.0 atm pressure. We further validate our AW3C and AW5C water models by showing that they can well reproduce the radial distribution function g(r), self‐diffusion constant D, and hydration free energy from the AMOEBA03 water model and the experimental observations. Furthermore, we show that our AW3C and AW5C water models can greatly accelerate (>5 times) the bulk water as well as biomolecular simulations when compared to AMOEBA water. Specifically, we demonstrate that the applications of AW3C and AW5C water models to simulate a DNA duplex lead to a threefold acceleration, and in the meanwhile well maintain the structural properties as the fully polarizable AMOEBA water. We expect that our AW3C and AW5C water models hold great promise to be widely applied to simulate complex bio‐molecules using the AMOEBA force field.
At present, there is still an urgent demand for novel smart materials that can achieve a diverse range of practical applications in syntheticmaterial area. Herein, we developed a simple but versatile aggregation-inducedemission luminogen (1) with a donor-acceptor structure and pronounced intramolecularcharge transfer property. Compound 1 showed aremarkablecolor changein sensitive response to polarity change making it to serve as a promising imagingprobe for detecting environmental polarity in cells and selective visualizationof lipid droplets in live tissues. Additionally, thiscompound exhibited a wide range of thermoresponsive behavior with ratiometric luminescence change and noticeable fluorescencecolor switching.It also can respond to anisotropic shearing force and isotropic hydrostaticpressure with prominent but contrast luminescence conversion due to the distinctdisturbance of the weak intermolecular interactions and charge transfer processes.Meanwhile, compound 1 was sensitive to external electric stimulus and displayedreversibly three-color switched electrochromism and on-to-off electroluminochromism. Such propertyallowed the fabrication of high-performance non-doped OLED with a high externalquantum efficiency of 5.22%. The present results may offer an important guidelinefor multifunctional molecular design and provide an important step forward toexpand the real-life applications of AIE-active smart materials.
Background H2A.B, the most divergent histone variant of H2A, can significantly modulate nucleosome and chromatin structures. However, the related structural details and the underlying mechanism remain elusive to date. In this work, we built atomic models of the H2A.B-containing nucleosome core particle (NCP), chromatosome, and chromatin fiber. Multiscale modeling including all-atom molecular dynamics and coarse-grained simulations were then carried out for these systems. Results It is found that sequence differences at the C-terminal tail, the docking domain, and the L2 loop, between H2A.B and H2A are directly responsible for the DNA unwrapping in the H2A.B NCP, whereas the N-terminus of H2A.B may somewhat compensate for the aforementioned unwrapping effect. The assembly of the H2A.B NCP is more difficult than that of the H2A NCP. H2A.B may also modulate the interactions of H1 with both the NCP and the linker DNA and could further affect the higher-order structure of the chromatin fiber. Conclusions The results agree with the experimental results and may shed new light on the biological function of H2A.B. Multiscale modeling may be a valuable tool for investigating structure and dynamics of the nucleosome and the chromatin induced by various histone variants.
Constructing artificial helical structures through hierarchical self-assembly and exploring the underlying mechanism are important, and they help gain insight from the structures, processes, and functions from the biological helices and facilitate the development of material science and nanotechnology. Herein, the two enantiomers of chiral Au(I) complexes ( S)-1 and ( R)-1 were synthesized, and they exhibited impressive spontaneous hierarchical self-assembly transitions from vesicles to helical fibers. An impressive chirality inversion and amplification was accompanied by the assembly transition, as elucidated by the results of in situ and time-dependent circular dichroism spectroscopy and scanning electron microscope imaging. The two enantiomers could serve as ideal chiral templates to co-assemble with other achiral luminogens to efficiently induce the resulting co-assembly systems to show circularly polarized luminescence (CPL). Our work has provided a simple but efficient way to explore the sophisticated self-assembly process and presented a facile and effective strategy to fabricate architectures with CPL properties.
MDM2 is a well-known oncoprotein overexpressed in a variety of cancers, and the identification of inhibitors that disrupt the MDM2/p53 interaction is of great interest in anticancer drug development. Here we designed a platform for the facile and visualizable identification of inhibitors of MDM2 using co-expressed protein complexes of MDM2/p53. A hexahistidine-tag on MDM2 allows the binding of the protein complex to the Ni-NTA affinity resin, while the fluorescent protein fused to p53 enables the direct visualization of the interaction of p53 with MDM2. Hence, the inhibition of the MDM2/p53 interaction can be observed with the naked eye. The assay can be set up by directly loading cell lysate to the Ni-NTA affinity resin, and no chemical modification of proteins is needed. In addition to the qualitative analyses, the binding affinity of inhibitors to the MDM2 protein can be quantified by fluorescence titration. The applications of this system have been verified using small molecules and peptide inhibitors. As a proof of concept, we screened a small library using this platform. Interestingly, two types of novel inhibitors of MDM2, including cyclohexyl-triphenylamine derivatives and platinum complexes, were identified and their binding affinities were obtained. Quantitative measurements show that these new types of inhibitors demonstrate a high binding affinity (up to Kd = 51.9 nM) to MDM2.