Membrane transport proteins translocate diverse cargos, ranging from small sugars to entire proteins, across cellular membranes1-3. A few structurally distinct protein families have been described that account for most of the known membrane transport processes4-6. However, many membrane proteins with predicted transporter functions remain uncharacterized. Here we determined the structure of Escherichia coli LetAB, a phospholipid transporter involved in outer membrane integrity, and found that LetA adopts a distinct architecture that is structurally and evolutionarily unrelated to known transporter families. LetA localizes to the inner membrane, where it is poised to load lipids into its binding partner, LetB, a mammalian cell entry (MCE) protein that forms an approximately 225 Å long tunnel for lipid transport across the cell envelope. Unexpectedly, the LetA transmembrane domains adopt a fold that is evolutionarily related to the eukaryotic tetraspanin family of membrane proteins, including transmembrane AMPA receptor regulatory proteins (TARPs) and claudins. Through a combination of deep mutational scanning, molecular dynamics simulations, AlphaFold-predicted alternative states and functional studies, we present a model for how the LetA-like family of membrane transporters facilitates the transport of lipids across the bacterial cell envelope.
MOTIVATION:De novo antibody design requires jointly determining the global binding orientation and shaping flexible CDR loops to engage a target epitope. Diffusion-based approaches such as RFantibody are capable of this joint task but frequently produce severe steric clashes requiring extensive post-hoc filtering. Flow-based methods such as IgFlow and FlowDesign offer more stable generation but remain restricted to pre-aligned frames, precluding true de novo design. Achieving structural integrity and epitope specificity simultaneously in this setting remains an open challenge. RESULTS:We propose TiDE-Ab, a conditional SE(3) flow matching framework for de novo epitope-specific antibody design. By conditioning on unpaired antigen and antibody structures without any pre-aligned frame, TiDE-Ab inherits the structural stability of flow matching while enabling global binding pose search from scratch. To further improve epitope targeting, we introduce Time-Dependent Classifier-Free Guidance (TD-CFG), which replaces static conditioning with an adaptive schedule: strong guidance early to establish the global binding pose, followed by gradual relaxation for precise local CDR refinement. On 55 non-redundant benchmark complexes, TiDE-Ab outperforms RFantibody with higher epitope recall (0.935 vs. 0.878) and over 95% fewer steric clashes. In therapeutic case studies on TGF-β and IL-17A, TiDE-Ab reproduced the binding profiles of clinical antibodies across isoform-selective and cross-reactive epitopes, whereas RFantibody consistently failed to produce viable candidates. AVAILABILITY:Source code, data, and trained models are available at https://github.com/SNU-CSSB/TiDE-Ab.
Abstract While AI offers transformative potential for therapeutic antibody design, the lack of ground-truth data fundamentally constrains our ability to model the epistatic topology of fitness landscapes. Here, we establish a high-throughput workflow to characterize tens of thousands of antibody variants per week with gold-standard biophysical precision. By combinatorially assembling functional variants from deep mutational scanning, we charted antibody fitness landscapes comprising over 17,000 data points, which revealed an extremely rugged, non-navigable epistatic topology. Yet, navigating at this unprecedented scale enabled the discovery of rare peak clusters exhibiting simultaneous enhancements in affinity and productivity. Strikingly, ProteinMPNN predicted the CDR-dependent productivity landscape with remarkable accuracy, suggesting that sequence-structure compatibility within CDRs gates cellular productivity. This insight enabled a structure-guided rescue strategy combining AlphaFold3 and ProteinMPNN, which successfully restored the cellular productivity of high-affinity, low-productivity clones via single amino acid substitutions. Two elite variants drawn directly from peak clusters further demonstrated 20- to 100-fold in vivo efficacy gains in a murine psoriasis model. Our findings establish CDR structural fitness as a fundamental determinant of antibody cellular productivity and validate landscape-scale navigation as a powerful framework for therapeutic antibody optimization.
Mannitol is a widely distributed sugar alcohol and a primary carbon source for diverse microbial communities, serving as a substrate for producing high-value metabolites. In Gammaproteobacteria, the conserved mtl operon is essential for mannitol metabolism, encoding the phosphotransferase system enzyme II (EIIMtl; mtlA), mannitol-1-phosphate dehydrogenase (mtlD), and the repressor MtlR (mtlR). While operon transcription is known to be subject to catabolite repression and dependent on cAMP receptor protein (CRP), the molecular mechanism of MtlR-mediated repression has remained unclear, as it lacks a canonical DNA-binding domain. Here, we demonstrate that MtlR represses transcription by utilizing CRP as a DNA-anchoring platform. In the absence of mannitol, MtlR forms a complex with CRP at the mtl operator, with recruitment specificity determined by the precise spacing between adjacent CRP-binding sites. With mannitol, the dephosphorylation of membrane-bound EIIMtl triggers the sequestration of MtlR at the membrane, thereby relieving repression and inducing operon expression. We reveal a dual regulatory mechanism: CRP spacing-dependent MtlR recruitment and EIIMtl-mediated MtlR sequestration. This dual-control strategy is selectively conserved across Gammaproteobacteria, with Enterobacterales retaining both functional modules, while other lineages exhibit distinct evolutionary divergence. This study highlights a non-canonical, broadly conserved strategy for integrating metabolic and environmental signals into bacterial gene control.
FAM19A5 is a secretory protein primarily expressed in neurons. Although its role in synaptic function has been suggested, the precise molecular mechanisms underlying its effects at the synapse remain unclear. Given that synaptic loss is a critical hallmark of Alzheimer’s disease (AD), elucidating the mechanisms involving FAM19A5 could provide valuable insights into reversing synaptic loss in AD. The binding partner of FAM19A5 was identified through co-immunoprecipitation experiments of mouse brain tissue. The effect of FAM19A5 on spine density in hippocampal neurons was evaluated using immunocytochemistry by overexpressing FAM19A5, treating neurons with FAM19A5 protein, and/or an anti-FAM19A5 antibody NS101. Target engagement of NS101 was determined by measuring FAM19A5 levels in mouse, rat, and human plasma at specific time points post NS101 injection using ELISA. Changes in spine density and dynamics in P301S tauopathy mice were assessed via Golgi staining and two-photon microscopy after NS101 administration. The synaptic strengthening of hippocampal neurons in APP/PS1 amyloidopathy mice after NS101 treatment was assessed by measuring miniature excitatory postsynaptic currents (mEPSCs) and field excitatory postsynaptic potentials (fEPSPs). Cognitive performance in AD mice after NS101 treatment was measured using the Y-maze and Morris water maze tests. FAM19A5 binds to LRRC4B, a postsynaptic adhesion molecule, leading to reductions in spine density in mouse hippocampal neurons. Inhibiting FAM19A5 function with NS101 increased spine density. Intravenous administration of NS101 increased spine density in the prefrontal cortex of P301S mice, which initially showed reduced spine density compared to wild-type (WT) mice. NS101 normalized the spine elimination rate in P301S mice, restoring the net spine count to levels comparable to WT mice. NS101 treatment enhanced the frequency of mEPSCs and fEPSPs in the hippocampal synapses of APP/PS1 mice, leading to improved cognitive function. The increases in plasma FAM19A5 levels upon systemic NS101 administration suggest that the antibody effectively engages its target and facilitates the transport of FAM19A5 from the brain. This study demonstrated that inhibiting FAM19A5 function with an anti-FAM19A5 antibody restores synaptic integrity and enhances cognitive function in AD, suggesting a novel therapeutic strategy for AD. https://clinicaltrials.gov/study/NCT05143463 , Identifier: NCT05143463, Release date: 3 December 2021.
Accurate prediction of protein structures is essential for understanding their biological functions. The release of AlphaFold2 in 2021 marked a significant breakthrough, delivering unprecedented accuracy. However, challenges remain, particularly for proteins with limited evolutionary data or complex molecular interactions. This review explores efforts to enhance AlphaFold2's performance through advanced sequence search techniques and alternative approaches, including protein language models and frameworks that integrate diverse biomolecular interactions. We propose that future progress will depend on developing models grounded in fundamental physicochemical principles, offering more accurate and comprehensive predictions across a wider spectrum of biological systems.
Protein-protein interactions (PPIs) are essential for biological function. Coevolutionary analysis and deep-learning (DL)-based protein structure prediction have enabled comprehensive PPI identification in bacteria and yeast, but these approaches have had limited success for the more complex human proteome. We overcame this challenge by enhancing the coevolutionary signals with sevenfold-deeper multiple sequence alignments harvested from 30 petabytes of unassembled genomic data and developing a new DL network trained on augmented datasets of domain-domain interactions from 200 million predicted protein structures. We systematically screened 200 million human protein pairs and predicted 17,849 interactions with an expected precision of 90%, of which 3631 interactions were not identified in previous experimental screens. Three-dimensional models of these predicted interactions provide numerous hypotheses about protein function and mechanisms of human diseases.
Membrane transport proteins translocate diverse cargos, ranging from small sugars to entire proteins, across cellular membranes. A few structurally distinct protein families have been described that account for most of the known membrane transport processes. However, many membrane proteins with predicted transporter functions remain uncharacterized. We determined the structure of E. coli LetAB, a phospholipid transporter involved in outer membrane integrity, and found that LetA adopts a distinct architecture that is structurally and evolutionarily unrelated to known transporter families. LetA functions as a pump at one end of a ~225 Å long tunnel formed by its binding partner, MCE protein LetB, creating a pathway for lipid transport between the inner and outer membranes. Unexpectedly, the LetA transmembrane domains adopt a fold that is evolutionarily related to the eukaryotic tetraspanin family of membrane proteins, including TARPs and claudins. LetA has no detectable homology to known transport proteins, and defines a new class of membrane transporters. Through a combination of deep mutational scanning, molecular dynamics simulations, AlphaFold-predicted alternative states, and functional studies, we present a model for how the LetA-like family of membrane transporters may use energy from the proton-motive force to drive the transport of lipids across the bacterial cell envelope.
The majority of proteins must form higher-order assemblies to perform their biological functions. Despite the importance of protein quaternary structure, there are few machine learning models that can accurately and rapidly predict the symmetry of assemblies involving multiple copies of the same protein chain. Here, we address this gap by training several classes of protein foundation models, including ESM-MSA, ESM2, and RoseTTAFold2, to predict homo-oligomer symmetry. Our best model named Seq2Symm, which utilizes ESM2, outperforms existing template-based and deep learning methods. It achieves an average PR-AUC of 0.48 and 0.44 across homo-oligomer symmetries on two different held-out test sets compared to 0.32 and 0.23 for the template-based method. Because Seq2Symm can rapidly predict homo-oligomer symmetries using a single sequence as input (~ 80,000 proteins/hour), we have applied it to 5 entire proteomes and ~ 3.5 million unlabeled protein sequences to identify patterns in protein assembly complexity across biological kingdoms and species.
Proteome-scale interaction prediction is essential for understanding protein functions and disease mechanisms. Traditional experimental methods are often limited by scale and complexity, driving the need for computational approaches. Deep learning has emerged as a powerful tool, enabling high-throughput, accurate predictions of protein interactions. This review highlights recent advances in deep learning methods for protein-protein and protein-ligand interaction screening, along with datasets used for model training. Despite the progress with deep learning, challenges such as data quality and validation biases remain. We also discuss the increasing importance of integrating structural information to enhance prediction accuracy and how structure-based deep learning approaches can help overcome current limitations, ultimately advancing biological research and drug discovery.
Identification of bacterial protein-protein interactions and predicting the structures of these complexes could aid in the understanding of pathogenicity mechanisms and developing treatments for infectious diseases. Here we developed RoseTTAFold2-Lite, a rapid deep learning model that leverages residue-residue coevolution and protein structure prediction to systematically identify and structurally characterize protein-protein interactions at the proteome-wide scale. Using this pipeline, we searched through 78 million pairs of proteins across 19 human bacterial pathogens and identified 1,923 confidently predicted complexes involving essential genes and 256 involving virulence factors. Many of these complexes were not previously known; we experimentally tested 12 such predictions, and half of them were validated. The predicted interactions span core metabolic and virulence pathways ranging from post-transcriptional modification to acid neutralization to outer-membrane machinery and should contribute to our understanding of the biology of these important pathogens and the design of drugs to combat them.
Methods for predicting bimolecular interactions are seeing tremendous growth, but challenges remain in capturing the full physical complexity of these interactions.
Deep-learning methods have revolutionized protein structure prediction and design but are presently limited to protein-only systems. We describe RoseTTAFold All-Atom (RFAA), which combines a residue-based representation of amino acids and DNA bases with an atomic representation of all other groups to model assemblies that contain proteins, nucleic acids, small molecules, metals, and covalent modifications, given their sequences and chemical structures. By fine-tuning on denoising tasks, we developed RFdiffusion All-Atom (RFdiffusionAA), which builds protein structures around small molecules. Starting from random distributions of amino acid residues surrounding target small molecules, we designed and experimentally validated, through crystallography and binding measurements, proteins that bind the cardiac disease therapeutic digoxigenin, the enzymatic cofactor heme, and the light-harvesting molecule bilin.
Proteomics has been revolutionized by large pre-trained protein language models, which learn unsupervised representations from large corpora of sequences. The parameters of these models are then fine-tuned in a supervised setting to tailor the model to a specific downstream task. However, as model size increases, the computational and memory footprint of fine-tuning becomes a barrier for many research groups. In the field of natural language processing, which has seen a similar explosion in the size of models, these challenges have been addressed by methods for parameter-efficient fine-tuning (PEFT). In this work, we newly bring parameter-efficient fine-tuning methods to proteomics. Using the parameter-efficient method LoRA, we train new models for two important proteomic tasks: predicting protein-protein interactions (PPI) and predicting the symmetry of homooligomers. We show that for homooligomer symmetry prediction, these approaches achieve performance competitive with traditional fine-tuning while requiring reduced memory and using three orders of magnitude fewer parameters. On the PPI prediction task, we surprisingly find that PEFT models actually outperform traditional fine-tuning while using two orders of magnitude fewer parameters. Here, we go even further to show that freezing the parameters of the language model and training only a classification head also outperforms fine-tuning, using five orders of magnitude fewer parameters, and that both of these models outperform state-of-the-art PPI prediction methods with substantially reduced compute. We also demonstrate that PEFT is robust to variations in training hyper-parameters, and elucidate where best practices for PEFT in proteomics differ from in natural language processing. Thus, we provide a blueprint to democratize the power of protein language model tuning to groups which have limited computational resources.
Protein–RNA and protein–DNA complexes play critical roles in biology. Despite considerable recent advances in protein structure prediction, the prediction of the structures of protein–nucleic acid complexes without homology to known complexes is a largely unsolved problem. Here we extend the RoseTTAFold machine learning protein-structure-prediction approach to additionally predict nucleic acid and protein–nucleic acid complexes. We develop a single trained network, RoseTTAFoldNA, that rapidly produces three-dimensional structure models with confidence estimates for protein–DNA and protein–RNA complexes. Here we show that confident predictions have considerably higher accuracy than current state-of-the-art methods. RoseTTAFoldNA should be broadly useful for modeling the structure of naturally occurring protein–nucleic acid complexes, and for designing sequence-specific RNA and DNA-binding proteins.
Antibodies, crucial in adaptive immunity, recognize antigens through specific interactions facilitated by Complementarity Determining Regions (CDRs), diversified via Variable-Diversity-Joining (VDJ) recombination. Traditional antibody development, limited by the scope of animal models and phage display libraries, captures a fraction of the potential antibody-antigen interactions. This underscores a gap in understanding antibody specificity and the relationship between antibody sequence and binding affinity. Here we introduce an approach using the Single-Protein Interaction Detection (SPID) platform, repurposed to systematically map local landscapes of antibody-antigen interactions with unprecedented depth and speed, aiming to rival the precision of methods like Surface Plasmon Resonance (SPR) and Bio-Layer Interferometry (BLI) while significantly boosting throughput. By editing CDR sequences and measuring effects on dissociation constants, we elucidated pathways for optimizing antibody affinity, enhancing predictive models for interactions. Our findings demonstrate the capability of the SPID platform to characterize thousands of variants weekly, offering a deeper insight into antibody-antigen interactions and advancing antibody development with finely-tuned affinities. ### Competing Interest Statement C.C., B.-K.S., J.H.J., J.L., B.Y., and T.-Y.Y. filed patents on these findings [patent number 10-2024-0057002 and 10-2024-0057004]. The other author declares no competing interests.
Protein-protein interactions (PPI) are essential for biological function. Recent advances in coevolutionary analysis and Deep Learning (DL) based protein structure prediction have enabled comprehensive PPI identification in bacterial and yeast proteomes, but these approaches have limited success to date for the more complex human proteome. Here, we overcome this challenge by 1) enhancing the coevolutionary signals with 7-fold deeper multiple sequence alignments harvested from 30 petabytes of unassembled genomic data, and 2) developing a new DL network trained on augmented datasets of domain-domain interactions from 200 million predicted protein structures. These advancements allow us to systematically screen through 200 million human protein pairs and predict 18,316 PPIs with an expected precision of 90%, among which 5,578 are novel predictions. 3D models of these predicted PPIs nearly triple the number of human PPIs with accurate structural information, providing numerous insights into protein function and mechanisms of human diseases. ### Competing Interest Statement The authors have declared no competing interest.
Mapping the ensemble of protein conformations that contribute to function and can be targeted by small molecule drugs remains an outstanding challenge. Here, we explore the use of variational autoencoders for reducing the challenge of dimensionality in the protein structure ensemble generation problem. We convert high-dimensional protein structural data into a continuous, low-dimensional representation, carry out a search in this space guided by a structure quality metric, and then use RoseTTAFold guided by the sampled structural information to generate 3D structures. We use this approach to generate ensembles for the cancer relevant protein K-Ras, train the VAE on a subset of the available K-Ras crystal structures and MD simulation snapshots, and assess the extent of sampling close to crystal structures withheld from training. We find that our latent space sampling procedure rapidly generates ensembles with high structural quality and is able to sample within 1 Å of held-out crystal structures, with a consistency higher than that of MD simulation or AlphaFold2 prediction. The sampled structures sufficiently recapitulate the cryptic pockets in the held-out K-Ras structures to allow for small molecule docking.
Identification of bacterial protein-protein interactions and predicting the structures of the complexes could aid in the understanding of pathogenicity mechanisms and developing treatments for infectious diseases. Here, we developed a deep learning-based pipeline that leverages residue-residue coevolution and protein structure prediction to systematically identify and structurally characterize protein-protein interactions at the proteome-wide scale. Using this pipeline, we searched through 78 million pairs of proteins across 19 human bacterial pathogens and identified 1923 confidently predicted complexes involving essential genes and 256 involving virulence factors. Many of these complexes were not previously known; we experimentally tested 12 such predictions, and half of them were validated. The predicted interactions span core metabolic and virulence pathways ranging from post-transcriptional modification to acid neutralization to outer membrane machinery and should contribute to our understanding of the biology of these important pathogens and the design of drugs to combat them.