Motivation Identifying antibody binding sites, is crucial for developing vaccines and therapeutic antibodies, processes that are time-consuming and costly. Accurate prediction of the paratope's binding site can speed up the development by improving our understanding of antibody-antigen interactions.Results We present ParaSurf, a deep learning model that significantly enhances paratope prediction by incorporating both surface geometric and non-geometric factors. Trained and tested on three prominent antibody-antigen benchmarks, ParaSurf achieves state-of-the-art results across nearly all metrics. Unlike models restricted to the variable region, ParaSurf demonstrates the ability to accurately predict binding scores across the entire Fab region of the antibody. Additionally, we conducted an extensive analysis using the largest of the three datasets employed, focusing on three key components: (i) a detailed evaluation of paratope prediction for each complementarity-determining region loop, (ii) the performance of models trained exclusively on the heavy chain, and (iii) the results of training models solely on the light chain without incorporating data from the heavy chain.Availability and implementation Source code for ParaSurf, along with the datasets used, preprocessing pipeline, and trained model weights, are freely available at https://github.com/aggelos-michael-papadopoulos/ParaSurf.
Spinocerebellar ataxia type 1 (SCA1) is a neurodegenerative disease caused by the expansion of a polyglutamine (polyQ) tract in the ATXN1 protein. This expansion is thought to be responsible for the gradual aggregation of the mutant protein, which is associated with increased cytotoxicity and neuronal cell death. Apart from the polyQ tract, other domains in ATXN1 are also involved in the initial events of protein aggregation, such as a dimerization domain that promotes protein oligomerization. ATXN1 interacts with various proteins; among them, MED15 significantly enhances the aggregation of the polyQ-expanded protein. Therefore, we set to identify the interaction site between ATXN1 and MED15 and assess whether its chemical targeting would affect polyQ protein aggregation. First, we predicted the structures of ATXN1 and MED15 and simulated their interaction. We experimentally validated that amino acids (aa) 99-163 of ATXN1 and aa548-665 of MED15 are critical for this protein-protein interaction (PPI). We also showed that the aa99-163 domain in ATXN1 is involved in the dimerization of the mutant isoform. Targeting this domain with a chemical compound identified through virtual screening (Chembridge ID: 5755483) inhibited both the interaction of ATXN1 with MED15 and the dimerization of polyQ-expanded ATXN1. These results strengthen our assumption that the aa99-163 domain of ATXN1 may be involved in polyQ protein aggregation and highlight compound 5755483 as a potent first-in-class therapeutic agent for SCA1.
The B cell receptor immunoglobulin (BcR IG) serves as a distinctive molecular marker for each B cell clone, facilitating interactions with antigens that subsequently impact clonal behavior. BcR signaling is crucial for every aspect of B cell physiology, while it plays a significant role in pathological conditions involving B cells, such as B cell malignancies and autoimmune disorders. Hence, understanding the structure of BcR IG complexed with its cognate antigenic epitopes is vital for unraveling the mechanisms underlying BcR-antigen interactions with far-reaching implications extending to the development of targeted immunotherapeutic approaches. Although studying actual protein crystals would be ideal, the crystallographic processes are notoriously lab-intensive and demanding. In order to overcome this limitation, we herein propose an innovative in-silico approach, employing 3D analysis of BcR-antigen interactions.The proposed DeepSurf2-FF extends the work of DeepSurf2.0 (Papadopoulos et al., 2023) by introducing the following innovative features: i) our deep learning model has been modified adopting a new state-of-the-art architecture; ii) a new set of non-geometric properties based on the force fields has been added to our model. The above extensions resulted in significant improvements in the prediction accuracy of our method. Our method was trained and tested on the same well-established paratope prediction benchmark used by PECAN (Pittala et al., 2020), which includes 460 BcR-antigen complexes. The dataset was divided into three sets: the training set (205 complexes), the validation set (103 complexes), and the test set (152 complexes) filtered to ensure that antibodies shared no more than 95% pairwise sequence identity. The dataset included only complexes with paired IG heavy and light chains, having a resolution below 3 Å, and protein antigens. Following previous methods, residues were labeled as binding if any heavy atom was within 4.5 Å of an antigen-heavy atom.Due to the nature of the paratope prediction task, binding site atoms are significantly fewer than non-binding atoms, resulting in a class imbalance. Our method was compared to the previous state-of-the-art Paragraph model (Chinery et al., 2023). The latter had set a high benchmark in paratope prediction, achieving remarkable results on this task. To further enhance its performance, the Paragraph model was also pre-trained on an expanded dataset of 1060 complexes from SabDab. In contrast, DeepSurf2-FF did not require pre-training on such an extensive dataset, since it achieved state-of-the-art results across all metrics, only by training on the official paratope prediction benchmark. DeepSurf2-FF and Paragraph were compared based on four key metrics; AUC-PR which evaluates the performance of imbalance datasets by focusing on the minority class (binding site atoms), providing insights into precision and recall trade-offs, AUC-ROC which measures the model's ability to discriminate between binding and non-binding atoms, F-score which balances the importance of false positives and false negatives and MCC which is a performance metric used to evaluate the quality of binary classifications. Compared to the Paragraph model, DeepSurf2-FF demonstrated significant improvements: a 17.82% increase in PR-AUC (0.820 vs. 0.696), a 4.6% increase in ROC-AUC (0.977 vs. 0.934), a 4.09% increase in F-score (0.713 vs. 0.685), and a 5.66% increase in MCC (0.691 vs. 0.654).Overall, through rigorous training and testing on a well-established paratope prediction benchmark, DeepSurf2-FF outperformed the previous state-of-the-art Paragraph model, demonstrating significant advancements across all evaluation metrics. This significant development highlights the potential of DeepSurf2-FF in enhancing our understanding of BcR-antigen interactions and facilitating future research in the field of (immuno)hematology, offering novel and profound insights into the natural history and management of B cell-related pathologies.
Coronavirus disease 2019 (COVID-19) is caused by a new, highly pathogenic severe-acute-respiratory syndrome coronavirus 2 (SARS-CoV-2) that infects human cells through its transmembrane spike (S) glycoprotein. The receptor-binding domain (RBD) of the S protein interacts with the angiotensin-converting enzyme II (ACE2) receptor of the host cells. Therefore, pharmacological targeting of this interaction might prevent infection or spread of the virus. Here, we performed a virtual screening to identify small molecules that block S-ACE2 interaction. Large compound libraries were filtered for drug-like properties, promiscuity and protein-protein interaction-targeting ability based on their ADME-Tox descriptors and also to exclude pan-assay interfering compounds. A properly designed AI-based virtual screening pipeline was applied to the remaining compounds, comprising approximately 10% of the starting data sets, searching for molecules that could bind to the RBD of the S protein. All molecules were sorted according to their screening score, grouped based on their structure and postfiltered for possible interaction patterns with the ACE2 receptor, yielding 31 hits. These hit molecules were further tested for their inhibitory effect on Spike RBD/ACE2 (19-615) interaction. Six compounds inhibited the S-ACE2 interaction in a dose-dependent manner while two of them also prevented infection of human cells from a pseudotyped virus whose entry is mediated by the S protein of SARS-CoV-2. Of the two compounds, the benzimidazole derivative CKP-22 protected Vero E6 cells from infection with SARS-CoV-2, as well. Subsequent, hit-to-lead optimization of CKP-22 was effected through the synthesis of 29 new derivatives of which compound CKP-25 suppressed the Spike RBD/ACE2 (19-615) interaction, reduced the cytopathic effect of SARS-CoV-2 in Vero E6 cells (IC50 = 3.5 mu M) and reduced the viral load in cell culture supernatants. Early in vitro ADME-Tox studies showed that CKP-25 does not possess biodegradation or liver metabolism issues, while isozyme-specific CYP450 experiments revealed that CKP-25 was a weak inhibitor of the CYP450 system. Moreover, CKP-25 does not elicit mutagenic effect on Escherichia coli WP2 uvrA strain. Thus, CKP-25 is considered a lead compound against COVID-19 infection.
The B cell receptor immunoglobulin (BcR IG) is a unique molecular identity for each B cell clone, underpinning interactions with foreign and (auto)antigens that eventually affect clonal behavior. BcR signaling is crucial for the homeostasis of B cells, affecting all aspects of their physiology including cell activation, proliferation, differentiation and apoptosis. Moreover, it is highly relevant for pathological conditions implicating B cells, e.g. B cell lymphomas and autoimmune disorders. Structural analysis of the BcR IG and its cognate antigenic epitopes is vital in elucidating the mechanisms of BcR-antigen interactions. While analyzing actual protein crystals would be ideal, the crystallographic procedures are notoriously labor-intensive and challenging. Hence, we pivot to an in-silico approach, utilizing 3D analysis of BcR-antigen interactions. Confronted with the inherent variability of BcRs and the arduous nature of experimental analyses, we present a cutting-edge solution: DeepSurf2.0. This innovative computational tool leverages deep learning algorithms to predict Protein-Protein Interactions (PPI) and more specifically BcR-antigen interactions, creating a foundation for fast and accurate protein-protein docking. DeepSurf2.0, specifically tailored for the 3D structures of BcR IG and associated antigens, harnesses the power of deep learning to predict PPI: therefore, a carefully curated dataset is of paramount importance. To achieve the latter, we took advantage of SAbDab, a database containing all the antibody structures available in the Protein Data Bank (PDB), annotated and presented in a consistent fashion. We refined the SAbDab dataset by applying the following filtering steps: (i) we retained only complete BcR IG, i.e. those with available heavy and light chains, (ii) we preserved only one biological assembly from multimeric protein complexes, (iii) we excluded BcRs without associated antigens, and (iv) we constructed each BcR-antigen pair to consist of three chains (one each heavy and light for the BcR and one for the antigen). Through these exacting measures, we created a comprehensive collection of 10,543 BcR-antigen pairs. DeepSurf2.0 was evaluated using two metrics: DCA (Distance between Predicted binding site center and nearest antigen Atom) and OVR (Intersection of real and predicted binding sites divided by their union). A binding site prediction was considered as a hit if DCA < 4 Å. For training purposes, we utilized 9,440 BcR-antigen pairs to optimize DeepSurf2.0. The model was then evaluated on a separate test set of 1,103 BcR-antigen pairs. In this evaluation, DeepSurf2.0 achieved a DCA rate of 33%, which means that a hit was detected in 364 out of 1,103 cases. To measure the quality of these predictions, we assessed the OVR metric that resulted in a rate of 22%. To the best of our knowledge, there are no relevant methods that have been tested in a similar dataset. Existing state-of-the-art PPI prediction approaches achieve similar scores in DCA and OVR; however, the utilized datasets consisted of single chains in receptor and ligand. In contrast, our model incorporates a more complex two-chain receptor paradigm, which is a more challenging task but closer to the reality of BcR-antigen interactions. The aforementioned results not only facilitate understanding molecular interactions but also provide valuable insights into potential BcR docking areas for antigens. This ability to predict and locate the most probable interaction sites has immediate practical implications, significantly expediting the docking process by negating the need for time-consuming blind docking. Since our results are not directly comparable with those of the current state-of-the-art methods, our dataset will be provided publicly as a benchmark to evaluate similar methods in two-chain receptor cases. In conclusion, DeepSurf2.0 serves as a foundation for enabling subsequent docking algorithms to target the predicted interaction binding surface rather than the entire protein structure. This advancement underscores the transformative potential of deep learning within the realm of (immuno)hematology, holding the potential to provide novel insights into the pathogenesis and progression of B cell-related disorders.
This paper presents the methods that have participated in the SHREC 2022 contest on protein–ligand binding site recognition. The prediction of protein- ligand binding regions is an active research domain in computational biophysics and structural biology and plays a relevant role for molecular docking and drug design. The goal of the contest is to assess the effectiveness of computational methods in recognizing ligand binding sites in a protein based on its geometrical structure. Performances of the segmentation algorithms are analyzed according to two evaluation scores describing the capacity of a putative pocket to contact a ligand and to pinpoint the correct binding region. Despite some methods perform remarkably, we show that simple non-machine-learning approaches remain very competitive against data-driven algorithms. In general, the task of pocket detection remains a challenging learning problem which suffers of intrinsic difficulties due to the lack of negative examples (data imbalance problem).
Recent technological advances in the fields of data and computer science have improved significantly the everyday life of people. However, technological advances are also being adopted by criminals to facilitate and expand their illicit actions. The Deep Learning (DL) paradigm has shown a significant potential in analysing complex structured data. However, in the crime detection domain, a limited number of public datasets is available, constrained to specific tasks only, which hinders the research and development of accurate and robust DL-assisted tools. The goal of this work is to extend the well-known UCF-crime dataset to the case of video captioning. To the best of our knowledge, this is the first publicly available crime-related video captioning dataset. A new proposed video captioning approach is compared to a plethora of state-of- the-art-methods in this dataset, while qualitative and quantitative characteristics of the latter are presented.
A video captioning dataset, extracted from the UCF-Crime dataset videos and described at the ICIP 2022 paper "UCF-CAP, video captioning in the wild"
Waste from electrical and electronic equipment is exacerbating the global environmental crisis. There is an urgent need to build a robust infrastructure capable of providing effective e-waste disposal options. In this work, a novel hybrid human-robot and system-agnostic application for relevant waste disassembly and recycling has been developed. Working on cells, collaborative robots, enhanced with state-of-the-art computer vision capabilities, can achieve near-real-time performance and high precision in the disassembly process. Additionally, a new screw dataset suitable for three separate computer vision tasks, namely instance segmentation, object detection, and semantic segmentation, is introduced to facilitate future research, which can be utilized almost for any screwing/unscrewing application beyond the current disassembly topic. Experiments demonstrating the robustness of the visual object detection and robotic 3D deprojection modules, which are the core aspects of the proposed architecture, have been conducted.
MOTIVATION:The knowledge of potentially druggable binding sites on proteins is an important preliminary step toward the discovery of novel drugs. The computational prediction of such areas can be boosted by following the recent major advances in the deep learning field and by exploiting the increasing availability of proper data.RESULTS:In this article, a novel computational method for the prediction of potential binding sites is proposed, called DeepSurf. DeepSurf combines a surface-based representation, where a number of 3D voxelized grids are placed on the protein's surface, with state-of-the-art deep learning architectures. After being trained on the large database of scPDB, DeepSurf demonstrates superior results on three diverse testing datasets, by surpassing all its main deep learning-based competitors, while attaining competitive performance to a set of traditional non-data-driven approaches.AVAILABILITY AND IMPLEMENTATION:The source code of the method along with trained models are freely available at https://github.com/stemylonas/DeepSurf.git.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Automated detection of small objects poses additional challenges, compared to bigger-sized ones, due to the former's limited resolution for extracting discriminative information. In such cases, even a slight misalignment between a candidate region and its ground truth target has a huge impact on their IoU which significantly increases the amount of noisy information. Given the fact that state of the art two-stage detection algorithms generate predefined shaped and sized candidate regions in pixel-level interval, the aforementioned misalignments are very likely to occur. In this work, a scalable object detection approach is introduced -specifically dedicated to small object parts- incorporating both learnable and handcrafted features. In particular, a set of simplified Gabor waveforms (SGWs) is applied to the raw data, ultimately producing an improved set of anchors for the region proposal network. These Gabor filters are further utilized generating a soft attention mask. Additionally, the interaction of a human with the object is also exploited by taking advantage of affordance-based information for further improvement of detection performance. Experiments have been conducted in a newly introduced device disassembly segmentation dataset, demonstrating the robustness of the method in detection of small device components.
The Vision Meets Drone Object Detection in Image Challenge (VisDrone-DET 2020) is the third annual object detector benchmarking activity. Compared with the previous VisDrone-DET 2018 and VisDrone-DET 2019 challenges, many submitted object detectors exceed the recent state-of-the-art detectors. Based on the selected 29 robust detection methods, we discuss the experimental results comprehensively, which shows the effectiveness of ensemble learning and data augmentation in drone captured object detection. The full challenge results are publicly available at the website http://aiskyeye.com/leaderboard/ .
The main goal of this chapter is to develop a system for automatic protein classification. Proteins are classified using CNNs trained on ImageNet, which are tuned using a set of multiview 2D images of 3D protein structures generated by Jmol, which is a 3D molecular graphics program. Jmol generates different types of protein visualizations that emphasize specific properties of a protein’s structure, such as a visualization that displays the backbone structure of the protein as a trace of the Cα atom. Different multiview protein visualizations are generated by uniformly rotating the protein structure around its central X, Y, and Z viewing axes to produce 125 images for each protein. This set of images is then used to fine-tune the pretrained CNNs. The proposed system is tested on two datasets with excellent results. The MATLAB code used in this chapter is available at https://github.com/LorisNanni.
Drug discovery involves extremely costly and time consuming procedures and can be significantly benefited by computational approaches, such as virtual screening (VS). Structure-based VS relies on scoring functions which aim to evaluate the binding of a candidate compound (ligand) on a protein target. Over the last few years, the advancement of the deep learning field has led to the development of novel scoring functions based on convolutional neural networks (CNN), which have achieved state-of-the-art results. In this paper, we present an integrated end-to-end VS pipeline for application on real-world drug discovery scenarios. It combines multiple conformations of the ligand with a new CNN scoring function based on the ResNet architecture, called ResNetVS, which incorporates also the docking output score in its evaluation. After experiments on the DUD-E dataset, it has shown notable performance, especially in early enrichment, where it overcomes current benchmarks. The proposed pipeline is finally applied on the emerging case of COVID-19 pandemic, in a struggle to discover inhibitors for the viral spike protein-ACE2 interaction.
In the current study, a region-based approach for object detection is presented that is suitable for handling very small objects and objects in low-resolution images. To address this challenge, an anchoring mechanism for the region proposal stage of the object detection algorithm is proposed, which boosts the performance in the detection of small objects with an insignificant computational overhead. Our method is applicable to the task of robot-assisted disassembly of Waste Electrical and Electronic devices (WEEE) in an industrial environment. Extensive experiments have been conducted in a newly formed device disassembly segmentation dataset with promising results.
Proteins are natural modular objects usually composed of several domains, each domain bearing a specific function that is mediated through its surface, which is accessible to vicinal molecules. This draws attention to an understudied characteristic of protein structures: surface, that is mostly unexploited by protein structure comparison methods. In the present work, we evaluated the performance of six shape comparison methods, among which three are based on machine learning, to distinguish between 588 multi-domain proteins and to recreate the evolutionary relationships at the protein and species levels of the SCOPe database. The six groups that participated in the challenge submitted a total of 15 sets of results. We observed that the performance of all the methods significantly decreases at the species level, suggesting that shape-only protein comparison is challenging for closely related proteins. Even if the dataset is limited in size (only 588 proteins are considered whereas more than 160,000 protein structures are experimentally solved), we think that this work provides useful insights into the current shape comparison methods performance, and highlights possible limitations to large-scale applications due to the computational cost. (C) 2020 The Author(s). Published by Elsevier Ltd.
The European Union-funded project LASIE, aims to assist forensics investigators through innovative tools for image and video analysis, object detection and tracking, and event detection. These tools exploit the latest advances in machine learning to handle the challenges in processing content from real-world data sources.
This track aimed at retrieving protein evolutionary classification based on their surfaces meshes only. Given that proteins are dynamic, non-rigid objects and that evolution tends to conserve patterns related to their activity and function, this track offers a challenging issue using biologically relevant molecules. We evaluated the performance of 5 different algorithms and analyzed their ability, over a dataset of 5,298 objects, to retrieve various conformations of identical proteins and various conformations of ortholog proteins (proteins from different organisms and showing the same activity). All methods were able to retrieve a member of the same class as the query in at least 94% of the cases when considering the first match, but show more divergent when more matches were considered. Last, similarity metrics trained on databases dedicated to proteins improved the results.
M. Melkemi合作论文数Faculte des Sciences et Techniques3