Splicing factors control exon inclusion in messenger RNAs, shaping transcriptome and proteome diversity. Their catalytic activity is regulated by multiple layers, making single-omic measurements on their own fall short in identifying which splicing factors underlie a phenotype. Here, we posit that splicing factor activity, defined as a splicing factor's ability to modulate exon inclusion, can be estimated from changes in exon inclusion signatures. To test this hypothesis, we benchmark methods for constructing splicing factor→exon networks and estimating splicing factor activity. We find that combining RNA-seq perturbation-based networks with VIPER (Virtual Inference of Protein Activity by Enriched Regulon analysis) accurately captures splicing factor activity as modulated by multiple regulatory layers. This approach integrates splicing factor regulation into a single score derived solely from exon inclusion signatures, allowing functional interpretation of heterogeneous conditions. As a proof of concept, we identify recurrent cancer splicing programs, revealing associations with oncogenic- and tumor suppressor-like splicing factors missed by conventional methods. These programs correlate with patient survival and key cancer hallmarks: initiation, proliferation, and immune evasion. Altogether, we show splicing factor activity can be accurately estimated from exon inclusion changes, enabling comprehensive analyses of splicing regulation with minimal data requirements.
Designing antibodies is complex and resource intensive. While deep learning and generative approaches have shown promise in the design of protein binders, achieving high affinity and stability remains challenging. We introduce EvolveX, a structure-based antibody design pipeline leveraging the empirical force field FoldX to design complementarity-determining regions (CDRs) of single-domain antibodies (VHHs). We demonstrate the ability of EvolveX to redesign a VHH targeting mouse Vsig4 (mVsig4) to address two challenges: enhancing stability and affinity for mVsig4 and redesigning it for high affinity to the human ortholog. Notably, EvolveX improved the binding affinity of VHHs to human Vsig4 by over 1,000-fold. Structural analyses by X-ray crystallography confirmed design accuracy. Next-generation sequencing (NGS) analysis further demonstrated the efficiency of FoldX-based design pipeline. Collectively, our study highlights EvolveX's potential to overcome current limitations in antibody design, offering a powerful tool for the development of therapeutics with enhanced specificity, stability, and efficacy.
Splicing factors shape the isoform pool of most transcribed genes, playing a critical role in cellular physiology. Their dysregulation is a hallmark of diseases like cancer, where aberrant splicing contributes to progression. While exon inclusion signatures accurately assess changes in splicing factor activity, systematically mapping disease-driver regulatory interactions at large scale remains challenging. Perturb-seq, which combines CRISPR-based perturbations with single-cell RNA sequencing, enables high-throughput measurement of perturbed gene expression signatures but lacks exon-level resolution, limiting its application for splicing factor activity analysis. Here, we show that shallow artificial neural networks can estimate splicing factor activity from gene expression signatures, bypassing the need for exon-level data. As a case study, we map the genetic interactions regulating splicing factors during carcinogenesis, using the shift in splicing program activity-where oncogenic-like splicing factors become more active than tumor suppressor-like factors-as a molecular reporter of a Perturb-seq screen. Our analysis reveals a cross-regulatory network among splicing factors, involving protein-protein and splicing-mediated interactions, with MYC and additional candidate pathways linking cancer driver mutations to splicing regulation. This regulation recapitulates splicing program dynamics during development. Altogether, we establish a versatile framework for studying splicing factor regulation and demonstrate its utility for uncovering disease mechanisms.
Protein-protein interactions (PPI) are fundamental to cellular signaling, forming robust networks that govern critical biological processes such as immune response, cell growth, and signal transduction. Nanobody-based therapies have emerged as a key strategy for modulating PPIs, offering exceptional potential due to their high specificity, stability, and ability to access challenging epitopes on PPI interfaces inside cells. The rational design of nanobodies relies mainly on understanding and predicting their binding regions, particularly the residues that contribute the most to the binding energy (binding hotspots). Existing computational methods do not fully provide a scalable solution for hotspot identification in nanobody design, leaving a critical gap in the rational design of these therapeutics. Here, we present a scalable and structure-aware algorithm for hotspot prediction in nanobody design. The algorithm queries a curated database of triplets of interacting residues obtained from ~20,000 non-redundant PDB structures. We showed that these triplets contain structural and energetic information, being able to assess the stability effect of residue variations in protein structures, Pearson R = 0.63 (MSE = 1.58 kcal/mol). More important than effects on stability is the ability of the algorithm to predict binding hotspots of protein-protein generic complexes and more specifically in complexes containing nanobodies. HotspotPred reached an accuracy of 0.73 for hotspot residue identification in a protein interaction dataset of 1160 Alanine mutants and correctly identified in 63.4% of the cases we predicted at least 2 residues on the binding surface.
Persistent infection of mosquito cells is essential for the transmission of arboviruses, yet how these viruses produce sufficient progeny without compromising host cell fitness remains unclear. Arbovirus genomes exhibit suboptimal codon usage for both human and mosquito hosts, raising questions about how they achieve efficient translation in such distinct cellular environments. Using chikungunya virus (CHIKV) as a model, we conducted a temporal, genome-wide analysis of transcription, translation, and tRNA modifications in Aedes albopictus C6/36 cells. Unlike in human cells, CHIKV infection does not alter the tRNA modification landscape in mosquitoes to overcome codon bias, nor does it induce widespread degradation of host transcripts. Instead, viral persistence is marked by a progressive, virus-specific repression of CHIKV RNA translation, occurring alongside a recovery of host mRNA translation. This translational balance, maintained independently of the RNAi system, enables sustained viral production without major disruption to host gene expression. Our findings identify translational control as a central mechanism underlying persistent arbovirus infection in mosquito cells. ### Competing Interest Statement The authors have declared no competing interest.
Since AlphaFold2's rise, many deep learning methods for protein design have emerged. Here, we validate widely used and recognized tools, compare them with first-principle methods, and explore their combinations, focusing on their effectiveness in protein redesign and potential for therapeutic repurposing. We address two challenges: evaluating tools and combinations ability to detect the effects of multiple concurrent mutations in protein variants, and leveraging large-scale datasets to compare modeling-free methods, namely force fields, which handle point mutations well with limited backbone rearrangement, and inverse folding tools, which excel at native sequence recovery but may struggle with non-natural proteins. Debuting TriCombine, a tool that identifies residue triangles in input structures, matches them to a structural database, and scores mutants based on substitution frequencies, we shortlisted candidates, modeled them with FoldX, and generated 16 SH3 mutants carrying up to 9 concurrent substitutions. The dataset was expanded to include 36 mutants and 11 crystal structures (7 newly solved), along with a parallel set of multiple non-concurrent mutants from three additional proteins. For broader validation, we analyzed 160,000 four-site GB1 mutants and 163,555 (single and double) variants across 179 natural and de novo domains. We show that combining AI-based modeling tools with force field scoring functions yields the most reliable results. Inverse folding tools perform very well but lose accuracy on less-represented proteins. First-principle force fields like FoldX remain the most accurate for point mutations. All methods perform worse when applied to unsolved de novo models, underscoring the need for hybrid strategies in robust protein design.
ABSTRACTEssentiality studies have traditionally focused on coding regions, often overlooking other small genetic regulatory elements. To address this, we obtained a high-resolution essentiality map at near-single-nucleotide precision of the genome-reduced bacteriumMycoplasma pneumoniae, combining transposon libraries containing promoter or terminator sequences. By integrating temporal transposon-sequencing data, we developed a novel method of essentiality assessment based on k-means unsupervised clustering, which provides dynamic and quantitative information on the fitness contribution of different genomic regions. We compared the insertion tolerance and persistence of the two engineered libraries, assessing the local impact of transcription and termination on cell fitness. Essentiality assessment at the local base-level revealed essential protein domains and small genomic regions that are either essential or inaccessible to transposon insertion. We also identified structural regions within essential genes that tolerate transposon disruptions, resulting in functionally split proteins. Overall, this study presents a nuanced view of gene essentiality, shifting from static and binary models to a more accurate perspective. Additionally, it provides valuable insights for genome engineering and enhances our understanding of the biology of genome-reduced cells.
MOTIVATION:The FoldX force field was originally validated with a database of 1000 mutants at a time when there were few high-resolution structures. Here, we have manually curated a database of 5556 mutants affecting protein stability, resulting in 2484 highly confident mutations denominated FoldX stability dataset (FSD), represented in non-redundant X-ray structures with <2.5 Å resolution, not involving duplicates, metals, or prosthetic groups. Using this database, we have created a new version of the FoldX force field by introducing pi stacking, pH dependency for all charged residues, improving aromatic-aromatic interactions, modifying the Ncap contribution and α-helix dipole, recalibrating the side-chain entropy of methionine, adjusting the H-bond parameters, and modifying the solvation contribution of tryptophan and others. RESULTS:These changes have led to significant improvements for the prediction of specific mutants involving the above residues/interactions and a statistically significant increase of FoldX predictions, as well as for the majority of the 20 aa. Removing all training sets data from FSD [Validation FoldX Stability Dataset (VFSD) dataset] resulted in improved predictions from R = 0.693 (RMSE = 1.277 kcal/mol) to R = 0.706 (RMSE = 1.252 kcal/mol) when compared with the previously released version. FoldX achieves 95% accuracy considering an error of ±0.85 kcal/mol in prediction and an area under the curve = 0.78 for the VFSD, predicting the sign of the energy change upon mutation. AVAILABILITY AND IMPLEMENTATION:FoldX versions 4.1 and 5.1 are freely available for academics at https://foldxsuite.crg.eu/.
Essentiality studies have traditionally focused on coding regions, often overlooking other small genetic regulatory elements. To address this, we combined transposon libraries containing promoter or terminator sequences to obtain a high-resolution essentiality map of a genome-reduced bacterium, at near-single-nucleotide precision when considering non-essential genes. By integrating temporal transposon-sequencing data by k-means unsupervised clustering, we present a novel essentiality assessment approach, providing dynamic and quantitative information on the fitness contribution of different genomic regions. We compared the insertion tolerance and persistence of the two engineered libraries, assessing the local impact of transcription and termination on cell fitness. Essentiality assessment at the local base-level revealed essential protein domains and small genomic regions that are either essential or inaccessible to transposon insertion. We also identified structural regions within essential genes that tolerate transposon disruptions, resulting in functionally split proteins. Overall, this study presents a nuanced view of gene essentiality, shifting from static and binary models to a more accurate perspective. Additionally, it provides valuable insights for genome engineering and enhances our understanding of the biology of genome-reduced cells.
The non-pathogenic Mycoplasma pneumoniae engineered chassis (Mycochassis) has demonstrated the ability to express therapeutic molecules in vitro and to be effective for treatment of lung infectious diseases in in vivo mouse models. However, the expression of heterologous molecules, whether secreted or exposed on the bacterial membrane has not been optimized to ensure sufficient secretion and/or exposure levels to exert a maximum in vivo biological effect. Here, we have improved the currently used secretion signal from MPN142 protein. We found that mutations at P1' position of the signal peptide cleavage site do not abrogate secretion but affect it. Increasing hydrophobicity and mutations at the C-terminal of the signal peptide increases secretion. We tested different lipoprotein signal peptides as possible N-terminal protein anchoring motifs on the Mpn cell surface. Unexpectedly we found that these peptides exhibit variable retention and secretion rates of the protein, with some sequences behaving as full secretion motifs. This raises the question of the biological role of the lipobox motif traditionally thought to anchor membrane proteins without a helical transmembrane domain. These results altogether represent a step forward in chassis optimization, offering different sequences for secretion or membrane retention, which could be used to improve Mycochassis as a delivery vector, and broadening its therapeutic possibilities.
MOTIVATION:Deep learning algorithms applied to structural biology often struggle to converge to meaningful solutions when limited data is available, since they are required to learn complex physical rules from examples. State-of-the-art force-fields, however, cannot interface with deep learning algorithms due to their implementation. RESULTS:We present MadraX, a forcefield implemented as a differentiable PyTorch module, able to interact with deep learning algorithms in an end-to-end fashion. AVAILABILITY AND IMPLEMENTATION:MadraX documentation, together with tutorials and installation guide, is available at madrax.readthedocs.io.
Developing antibodies is complex and resource-intensive, and methods for designing antibodies targeting specific epitopes are lacking. We introduce a de novo antibody design approach leveraging the empirical force field FoldX to design complementarity determining regions (CDRs). Starting from a scaffold VHH, we tackled three challenges of increasing difficulty: 1) design the CDRs to optimize VHH stability and affinity for its original target; 2) design the CDRs for high affinity to the human ortholog; 3) design the CDRs for low nanomolar affinity for a pre-defined epitope on the unrelated human Interleukin-9 receptor alpha, for which no antibodies were previously developed. For each challenge we reached single digit nanomolar affinity in a single design cycle. Our approach allows de novo design of high-affinity VHHs while ensuring specificity and stability. ### Competing Interest Statement RvdK, JS, FR, LSP, JDB, DC. are co-inventors on E.U. provisional patent number EP 24170283.6 which covers the computational antibody design pipeline described here.
The diverse range of organizations contributing to the global research ecosystem is believed to enhance the overall quality and resilience of its output. Mid-sized autonomous research institutes, distinct from universities, play a crucial role in this landscape. They often lead the way in new research fields and experimental methods, including those in social and organizational domains, which are vital for driving innovation. The EU-LIFE alliance was established with the goal of fostering excellence by developing and disseminating best practices among European biomedical research institutes. As directors of the 15 EU-LIFE institutes, we have spent a decade comparing and refining our processes. Now, we are eager to share the insights we've gained. To this end, we have crafted this Charter, outlining 10 principles we deem essential for research institutes to flourish and achieve ground-breaking discoveries. These principles, detailed in the Charter, encompass excellence, independence, training, internationality and inclusivity, mission focus, technological advancement, administrative innovation, cooperation, societal impact, and public engagement. Our aim is to inspire the establishment of new institutes that adhere to these principles and to raise awareness about their significance. We are convinced that they should be viewed a crucial component of any national and international innovation strategies.
A non-pathogenic Mycoplasma pneumoniae-based chassis is leading the development of live biotherapeutics (LBPs) for respiratory diseases. However, reports connecting Guillain-Barré syndrome (GBS) cases to prior M. pneumoniae infections represent a concern for exploiting such a chassis. Galactolipids, especially galactocerebroside (GalCer), are considered the most likely M. pneumoniae antigens triggering autoimmune responses associated with GBS development. In this work, we generated different strains lacking genes involved in galactolipids biosynthesis. Glycolipid profiling of the strains demonstrated that some mutants show a complete lack of galactolipids. Cross-reactivity assays with sera from GBS patients with prior M. pneumoniae infection showed that certain engineered strains exhibit reduced antibody recognition. However, correlation analyses of these results with the glycolipid profile of the engineered strains suggest that other factors different from GalCer contribute to sera recognition, including total ceramide levels, dihexosylceramide (DHCer), and diglycosyldiacylglycerol (DGDAG). Finally, we discuss the best candidate strains as potential GBS-free Mycoplasma chassis.
Identifying open reading frames (ORFs) being translated is not a trivial task. ProTInSeq is a technique designed to characterize proteomes by sequencing transposon insertions engineered to express a selection marker when they occur in-frame within a protein-coding gene. In the bacterium Mycoplasma pneumoniae , ProTInSeq identifies 83% of its annotated proteins, along with 5 proteins and 153 small ORF-encoded proteins (SEPs; ≤100 aa) that were not previously annotated. Moreover, ProTInSeq can be utilized for detecting translational noise, as well as for relative quantification and transmembrane topology estimation of fitness and non-essential proteins. By integrating various identification approaches, the number of initially annotated SEPs in this bacterium increases from 27 to 329, with a quarter of them predicted to possess antimicrobial potential. Herein, we describe a methodology complementary to Ribo-Seq and mass spectroscopy that can identify SEPs while providing other insights in a proteome with a flexible and cost-effective DNA ultra-deep sequencing approach.
While it is recognised that protein functions are determined by their proteoform state, such as mutations and post-translational modifications, methods to determine their differential abundance between conditions are limited. Here, we present a novel workflow for classical immunoprecipitation coupled to mass spectrometry (IP-MS) data that focuses on identifying differential peptidoforms of the bait protein between conditions, providing additional information about protein function.### Competing Interest StatementThe authors have declared no competing interest.
Alternative Splicing (AS) programs serve as instructive signals of cell type specificity, particularly within the brain, which comprises dozens of molecularly and functionally distinct cell types. Among them, retinal photoreceptors stand out due to their unique transcriptome, making them a particularly well-suited system for studying how AS shapes cell type-specific molecular functions. Here, we use the Splicing Regulatory State (SRS) as a novel framework to discuss the splicing factors governing the unique AS pattern of photoreceptors, and how this pattern may aid in the specification of their highly specialized sensory cilia. In addition, we discuss how other sensory cells with ciliated structures, for which data is much scarcer, also rely on specific SRSs to implement a proteome specialized in the detection of sensory stimuli. By reviewing the general rules of cell type- and tissue-specific AS programs, firstly in the brain and subsequently in specialized sensory neurons, we propose a novel paradigm on how SRSs are established and how they can diversify. Finally, we illustrate how SRSs shape the outcome of mutations in splicing factors to produce cell type-specific phenotypes that can lead to various human diseases. This Review discusses the general rules of cell type- and tissue-specific alternative splicing programs in the brain and in specialized sensory neurons and proposes a new paradigm on how ‘splicing regulatory states’ are established, how they can diversify and how they are linked to human diseases.
We investigate the emergence, mutation profile, and dissemination of SARS-CoV-2 lineage B.1.214.2, first identified in Belgium in January 2021. This variant, featuring a 3-amino acid insertion in the spike protein similar to the Omicron variant, was speculated to enhance transmissibility or immune evasion. Initially detected in international travelers, it substantially transmitted in Central Africa, Belgium, Switzerland, and France, peaking in April 2021. Our travel-aware phylogeographic analysis, incorporating travel history, estimated the origin to the Republic of the Congo, with primary European entry through France and Belgium, and multiple smaller introductions during the epidemic. We correlate its spread with human travel patterns and air passenger data. Further, upon reviewing national reports of SARS-CoV-2 outbreaks in Belgian nursing homes, we found this strain caused moderately severe outcomes (8.7% case fatality ratio). A distinct nasopharyngeal immune response was observed in elderly patients, characterized by 80% unique signatures, higher B- and T-cell activation, increased type I IFN signaling, and reduced NK, Th17, and complement system activation, compared to similar outbreaks. This unique immune response may explain the variant's epidemiological behavior and underscores the need for nasal vaccine strategies against emerging variants.