The mRNA technology represents a paradigm shift for the biopharmaceutical industry because in vivo delivery of antibody-based therapeutics can expand their clinical applications. Moreover, the mRNA delivery shifts the burden of assuring CMC (chemistry, manufacturing, and control) to the mRNA molecule. However, in vivo developability issues such as intracellular expression, proper protein folding, secretion into extracellular milieu, degradation in physiological conditions, poly-specificity, immunogenicity, safety, efficacy, and pharmacology can still pose significant barriers to the success of these products. Here, we have leveraged insights from the sequence-structural properties of the variable regions (Fvs) from the marketed biotherapeutics to build a medicine-likeness profile for the mRNA-delivered biotherapeutics. Our dataset consists of 122 unique Fvs from 117 marketed biotherapeutics (as of February 2024) and was systematically evaluated using 9 nonredundant sequence and structural parameters to capture a holistic view of developability attributes relevant to in vivo performance of mRNA-delivered antibody therapeutics. To ensure robustness, 25 antibodies that failed to reach approval due to developability issues, as well as those withdrawn from the market post approval, were included as controls. Our findings highlight a complementary relationship between sequence- and structure-based characteristics, leading us to develop a combined scoring system. This integrated "medicine-likeness" profile enables early-stage in silico assessments of mRNA-delivered therapeutic antibody candidates, offering a valuable tool for researchers to predict and optimize their in vivo developability.
Motivation: Molecular visualization is dominated by two families of software. Desktop programs such as PyMOL, VMD and UCSF ChimeraX are powerful but heavy to install and operate. Web viewers such as Jmol, 3Dmol.js, NGL, Mol* and iCn3D are light and installation-free, but typically stop at rendering, lacking deeper functionality needed for even basic protein sequence, structural analyses and design tasks. Tools building upon these often lack comprehensive sequence/structure manipulation, surface, interface and antibody-specific analytics functionality or require an external backend. There is a need for a tool that is as frictionless as a web viewer yet carries the analytical depth normally reserved for the desktop or the command line software. Results: We present Protein Design Viz (PDV), a self-contained molecular visualization studio delivered as a single offline HTML file that runs entirely in the browser. Built as an extensive modification of 3Dmol.js, PDV combines a full visualization workflow: multi-object scenes, representations, coloring palette, linked sequence track, publication-quality outline rendering and portable sessions. Additionally we re-implemented four commonly used macromolecular analyses measures from scratch in client-side JavaScript: a Shrake-Rupley solvent-accessible surface area (SASA) engine, an antibody numbering and germline-assignment engine, a non-covalent interaction detector, and a developability-liability scanner. Each engine is validated against its established reference. PDV's SASA reproduces FreeSASA at Pearson r ≈ 0.997-0.998 across 2,582 structures spanning proteins, nucleic acids and ligands; its numbering reproduces RIOT for over 99.88% of 1.3 million residue positions across 16,996 sequences and four schemes; its interaction detector reproduces PLIP at macro-F1 0.82, matching Arpeggio as closely as PLIP itself does. PDV brings validated, quantitative structural analysis into a no-install, simple to use tool. Availability and implementation: PDV is a single HTML file, free for noncommercial use under the PolyForm Noncommercial License 1.0.0, available from pdv.naturalantibody.com. It requires only a WebGL-capable browser and runs fully offline.
ABSTRACT Protein aggregation is central to amyloid-related disorders and remains a major developability challenge for protein therapeutics. Over the past two decades, significant advances have been made to predict aggregation-prone regions (APRs) and estimate aggregation propensity in proteins and peptides. In contrast, the prediction of aggregation kinetics has received relatively less attention due to the limited availability and heterogeneity of experimental data. Consequently, aggregation propensities from APR prediction algorithms were widely accepted as a means to predict relative changes in the aggregation kinetics of proteins and mutants. Previous studies have demonstrated, using large-scale datasets, that aggregation propensity shows a weak or inconsistent correlation with aggregation kinetics. In the present study, we have integrated complementary state-of-the-art mechanistic and kinetic prediction tools for protein aggregation into a unified, user-friendly web framework entitled “Amylo-Pipe”. Amylo-Pipe also implements practical features that are especially useful for protein engineering, such as gatekeeper-residue mutational scanning to support the design of aggregation-resistant variants. By consolidating multiple prediction tasks in a single interface, Amylo-Pipe enables a more comprehensive assessment of aggregation behavior than APR-only workflows. The web server is freely accessible at: https://web.iitm.ac.in/bioinfo2/amylopipe/ .
Protein structure is closely linked with protein function. For many decades, scientists tried to predict the three-dimensional structure of a protein from its amino acid sequence. Experimental methods such as X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryo-electron microscopy provide reliable structures. However, these methods can be expensive and time-consuming. Computational methods therefore became important alternatives. Early prediction methods were based mainly on sequence similarity, physical energy functions, and structural templates. Their accuracy was limited for many proteins. Deep learning changed this field. AlphaFold2 showed that artificial intelligence can predict the structures of many proteins with near-experimental accuracy (Jumper et al., 2021). RoseTTAFold provided another powerful deep-learning approach (Baek et al., 2021). Protein language models such as ESMFold later showed that structural information can also be learned directly from very large collections of protein sequences (Lin et al., 2023). AlphaFold3 further expanded the field by predicting interactions among proteins, DNA, RNA, ions, small molecules, and modified residues (Abramson et al., 2024). Generative models such as RFdiffusion and ESM3 are now moving the field from structure prediction toward protein design (Watson et al., 2023; Hayes et al., 2025). Recent open models are also increasing access to advanced biomolecular modelling. Despite this progress, major challenges remain. Proteins are dynamic molecules. They may adopt several conformations. Their structures are influenced by ligands, membranes, modifications, and the cellular environment. Future models therefore need to predict not only one structure, but also molecular dynamics, interactions, binding strength, and biological function. This review describes the development of deep learning for protein structure prediction and discusses the major directions that may define the next frontier.
Solute-solvent, solute-solute and solvent-solvent interactions are examined via thermodynamics using apparent molar properties which are temperature dependent and are useful to define the isolated contribution of each component to the non-ideality of the mixture. Apparent molar volumes (V ϕ) and apparent molar adiabatic compressibilities (K ϕ) were investigated for three binary mixtures with different anions: 1-butyl-3-methylimidazolium chloride [Bmim][Cl], 1-butyl-1-methylpyrrolidinium chloride [Bmpym][Cl] and 1-butyl-3-methylimidazolium thiocyanide [Bmim][SCN] with dimethylformamide (DMF) at different temperatures (293.15-343.15) K and at ambient pressure. Density (ρ) and speed of sound (u) of the pure components and their mixtures were recorded. The data was fitted to the Redlich-Mayer polynomial equation to calculate the derived thermodynamic parameters: limiting apparent molar volume (V 0 ϕ), limiting apparent molar expansion (E 0 ϕ), thermal expansion coefficients (α p) and limiting apparent molar adiabatic compressibility (K 0 ϕ) along with their associated parameters (S v, B v, S k, B k). The primary focus of this study was to examine the effect of temperature on the anion and cation interaction of the IL with DMF and how these changes affected the IL structure. The computational investigation further examined the IL-solvent interaction energy and described the type of interaction in all three systems.
Antibody discovery has been successful in designing and progressing molecules to the clinic and market based on largely empirical methods and human experience. The field is now transitioning from classical monospecific antibodies to innovative smart biologics that employ diverse mechanisms of action, such as targeting, antagonism, agonism, and target-independent function. This evolution is being assisted, augmented, and potentially disrupted by artificial intelligence and machine learning (AI/ML) technologies. This perspective is focused on bringing clarity to the strategy and thinking that is required when designing antibody drug candidates and how emerging AI/ML strategies can address the real-world challenges of drug discovery and continue to improve performance.
Light chain amyloidosis is a medical condition characterized by the aggregation of misfolded antibody light chains into insoluble amyloid fibrils in the target organs, causing organ dysfunction, organ failure, and death. Despite extensive research to understand the factors contributing to amyloidogenesis, accurately predicting whether a given protein will form amyloids under specific conditions remains a formidable challenge. In this study, we have conducted a comprehensive analysis to understand the amyloidogenic tendencies within a dataset containing 1828 (348 amyloidogenic and 1480 non-amyloidogenic) antibody light chain variable region (VL) sequences obtained from the AL-Base database. Physicochemical and structural features often associated with protein aggregation, such as net charge, isoelectric point (pI), and solvent-exposed hydrophobic regions did not reveal a consistent association with the aggregation capability of the antibody light chains. However, the solvent-exposed aggregation-prone regions (APRs) occur with higher frequencies among the amyloidogenic light chains when compared with the non-amyloidogenic ones, with the difference ranging from 2% to 15% at various relative solvent-accessible surface area (rASA) cutoffs. We have, for the first time, identified structural gatekeeping residues around the APRs and assessed their impact on the amyloidogenicity of the antibody light chains. The non-amyloidogenic light chains contain these structural gatekeeper residues vicinal to their APRs more often than the amyloidogenic ones. We observed that the rASA cutoff of 35% is optimal for identifying the surface-exposed APRs, and a 4 Å distance cutoff from the APR motif(s) is optimal for identifying the structural gatekeeper residues. Moreover, lambda light chains were found to contain solvent-exposed APRs more often and surrounded by fewer gatekeepers, rendering them more susceptible to aggregation. The insights gained from this report have significant implications for understanding the molecular origins of light-chain amyloidosis in humans and the design of aggregation-resistant therapeutic antibodies.
Antibody-based biotherapeutics make up an important class of biopharmaceuticals. However, their discovery requires resource- and time-consuming laboratory processes. To ameliorate this situation, several computational methods were used to predict the structures of antibody:antigen complexes (Ab:Ag) and identify potential binders, in-silico. However, there is still a general lack of rapid virtual screening methods capable of screening large antibody libraries against a given antigen or group of antigens. In this work, we explore the application of a successful small-molecule drug discovery strategy and adapt pharmacophore-based virtual screening to the world of antibody discovery. Using a nonredundant data set of 874 Ab:Ag complexes, we have developed an automated method to create pharmacophores from the antibody complementarity determining regions. Our method is 98.6% (862 out of 874) successful at reproducing the ground truth, i.e., it can recapitulate the parental antibody:antigen complexes. In a benchmarking comparison with cognate docking, using 33 Ab:Ag complexes of therapeutic interest, the pharmacophore method was not only much faster than cognate docking but also recovered all the native interfacial contacts. In addition, it can also find additional putative antibody binders to a given antigen within clusters of Ab:Ag complexes with similar interfacial structures. Our method has significant implications toward accelerating biotherapeutic drug discovery as well as drug repurposing research. This method was implemented in MOE 2024 and is available to the scientific community.
Understanding the pairing preferences and structural interactions between antibody heavy and light chains can enhance our ability to design more effective and specific therapeutic antibodies. Insights from natural antibody repertoires and conserved contact sites help reduce autoreactivity and improve drug safety and efficacy. Current databases represent only a limited portion of the estimated diversity of unique paired antibody molecules. To address this, we introduce PairedAbNGS, a novel database with paired heavy/light antibody chains. To our knowledge, this is the largest resource for paired natural antibody sequences with 58 bioprojects and over 14 million assembled productive sequences. Using this dataset, we investigated heavy and light chain variable (V) gene pairing preferences and found significant biases beyond gene usage frequencies, possibly due to receptor editing favoring less autoreactive combinations. Analyzing the available antibody structures from the Protein Data Bank, we studied conserved contact residues between heavy and light chains, particularly interactions between the CDR3 region of one chain and the FWR2 region of the opposite chain. Examination of amino acid pairs at key contact sites revealed significant deviations of amino acids distributions compared to random pairings, in the heavy chain's CDR3 region contacting the opposite chain, indicating specific interactions might be crucial for proper chain pairing. This observation is further reinforced by preferential IGHV-IGLJ and IGLV-IGHJ pairing preferences. We hope that both our resources and the findings would contribute to improving the engineering of biological drugs. We make the database accessible at https://naturalantibody.com/paired-ab-ngs as a valuable tool for biological and machine-learning applications.
Although antibody variable regions mediate antigen-specific binding, they can also mediate non-specific interactions with non-cognate antigens, impacting diverse immunological processes and the efficacy, safety, and half-life of antibody therapeutics. To understand the molecular basis of antibody non-specificity, we sorted two dissimilar human naïve antibody libraries against multiple reagents to enrich for variants with different levels of polyreactivity. Sequence analysis of >300,000 paired antibody variable regions revealed that the heavy chain primarily mediates human antibody polyreactivity, and this is due to the high positive charge, high hydrophobicity, and combinations thereof in the corresponding complementarity-determining regions, which can be predicted using a machine learning model developed in this work. Notably, a subset of the most important features governing antibody non-specific interactions, namely those that contain tyrosine, also govern specific antigen recognition. Our findings are broadly relevant for understanding fundamental aspects of antibody molecular recognition and the applied aspects of antibody-drug design.
Antibody-based therapeutics must not undergo chemical modifications that would impair their efficacy or hinder their developability. A commonly used technique to de-risk lead biotherapeutic candidates annotates chemical liability motifs on their sequence. By analyzing sequences from all major sources of data (therapeutics, patents, GenBank, literature, and next-generation sequencing outputs), we find that almost all antibodies contain an average of 3-4 such liability motifs in their paratopes, irrespective of the source dataset. This is in line with the common wisdom that liability motif annotation is over-predictive. Therefore, we have compiled three computational flags to prioritize liability motifs for removal from lead drug candidates: 1. germline, to reflect naturally occurring motifs, 2. therapeutic, reflecting chemical liability motifs found in therapeutic antibodies, and 3. surface, indicative of structural accessibility for chemical modification. We show that these flags annotate approximately 60% of liability motifs as benign, that is, the flagged liabilities have a smaller probability of undergoing degradation as benchmarked on two experimental datasets covering deamidation, isomerization, and oxidation. We combined the liability detection and flags into a tool called Liability Antibody Profiler (LAP), publicly available at lap.naturalantibody.com. We anticipate that LAP will save time and effort in de-risking therapeutic molecules.
EDITORIAL article Front. Mol. Biosci., 07 February 2024Sec. Biological Modeling and Simulation Volume 11 - 2024 | https://doi.org/10.3389/fmolb.2024.1360267
In the past 40 years, therapeutic antibody discovery and development have advanced considerably, with machine learning (ML) offering a promising way to speed up the process by reducing costs and the number of experiments required. Recent progress in ML -guided antibody design and development (D&D) has been hindered by the diversity of data sets and evaluation methods, which makes it dif ficult to conduct comparisons and assess utility. Establishing standards and guidelines will be crucial for the wider adoption of ML and the advancement of the field. This perspective critically reviews current practices, highlights common pitfalls and proposes method development and evaluation guidelines for various ML -based techniques in therapeutic antibody D&D. Addressing challenges across the ML process, best practices are recommended for each stage to enhance reproducibility and progress.
Metal-organic frameworks (MOFs) are a novel class of materials that were produced by combining a metal cluster with an organic linker. Due to the extensive compositional and structural variations that could be introduced thanks to this knowledge, crystalline micro- and mesoporous architectures have a huge structural library to choose from. With their novel properties, MOFs are a class of highly intriguing porous solids that hold great promise for a wide range of potential uses, including biomedicine, health care, diagnosis, therapy, and theragnostics, as well as industrially important processes like gas storage and separation, water harvesting, catalysis, and energy conversion and storage. According to reports, at least one paper on a MOF-related topic is published in every new issue of chemistry journals, and research on the synthesis and applications of MOFs is progressing at an astounding rate. The crystal sizes used to create MOFs are highly tunable and range from a few nanometers to micrometers. The uses of MOFs and their dimensions are somehow linked. As an illustration, bulk MOFs with diameters in micrometers or higher are widely used for gas storage, separation, and catalysis. Since MOFs with dimensions up to 500 nm have also been described as nanoscale MOFs, nanoscale MOFs, like other nanomaterials, cannot be restricted to MOFs with dimensions in the range of 1–100 nm.
Beyond potency, a good developability profile is a key attribute of a biological drug. Selecting and screening for such attributes early in the drug development process can save resources and avoid costly late-stage failures. Here, we review some of the most important developability properties that can be assessed early on for biologics. These include the influence of the source of the biologic, its biophysical and pharmacokinetic properties, and how well it can be expressed recombinantly. We furthermore present in silico, in vitro, and in vivo methods and techniques that can be exploited at different stages of the discovery process to identify molecules with liabilities and thereby facilitate the selection of the most optimal drug leads. Finally, we reflect on the most relevant developability parameters for injectable versus orally delivered biologics and provide an outlook toward what general trends are expected to rise in the development of biologics.