Artificial intelligence (AI) has recently seen transformative breakthroughs in the life sciences, expanding possibilities for researchers to interpret biological information at an unprecedented capacity, with novel applications and advances being made almost daily. In order to maximise return on the growing investments in AI-based life science research and accelerate this progress, it has become urgent to address the exacerbation of long-standing research challenges arising from the rapid adoption of AI methods. We review the increased erosion of trust in AI research outputs, driven by the issues of poor reusability and reproducibility, and highlight their consequent impact on environmental sustainability. Furthermore, we discuss the fragmented components of the AI ecosystem and lack of guiding pathways to best support Open and Sustainable AI (OSAI) model development. In response, this perspective introduces a practical set of OSAI recommendations directly mapped to over 300 components of the AI ecosystem. Our work connects researchers with relevant AI resources, facilitating the implementation of sustainable, reusable and transparent AI. Built upon life science community consensus and aligned to existing efforts, the outputs of this perspective are designed to aid the future development of policy and structured pathways for guiding AI implementation.
OBJECTIVE:Both susceptibility to, and severity of, rheumatoid arthritis (RA) is associated with the rs26232 C allele. Our primary aim was to identify the biologic mechanism underlying this association. METHODS:Expression of surrounding genes was compared among rs26232 genotypes. Publicly available databases were used to correlate expression with RA inflammation and single-cell synovial distribution. Inhibition of gene expression and activity was achieved using small interfering RNA and a pharmacology agent and effects on RA synovial fibroblasts (RASFs) characteristics in vitro were assayed. The amidated secretome of synovial fibroblasts were characterized by mass spectrometry and enzyme-linked immunosorbent assay. Effects of amidated peptides on macrophage polarity were determined using an RASF-macrophage coculture module. RESULTS:rs26232 C is associated with low expression of peptidylglycine alpha-amidating monooxygenase (PAM) in multiple tissues including RASFs. Synovial PAM is highly expressed in RASFs but not immune cells, and levels are inversely correlated with synovial and systemic levels of inflammation. Inhibition of PAM in RASFs increased tissue-damaging activities such as invasiveness in vitro. The most abundant amidated peptides secreted by RASFs were adrenomedullin (ADM) and pro-ADM N-terminal peptide (PAMP). Incubation of RASFs with either peptide inhibited interleukin-6 (IL-6) and IL-8, increased transforming growth factor β production, and reduced invasiveness in vitro. Inhibition of amidation in an RASF-macrophage coculture model skewed the macrophages to proinflammatory MerTK- phenotypes. CONCLUSION:Genetically determined low PAM reduces the anti-inflammatory and tissue-damaging activities of ADM and PAMP mediated by macrophages and RASFs, explaining the association of rs26232 C with RA severity.
We investigated how seed proteolysis was enhanced by germination, by subsequent homogenisation (disrupting sprout compartments), and by co-incubation of homogenates from different species.Mass spectrometry of released peptides tracked proteolytic signatures from chickpea, lentil, mung and broccoli proteins, in soaked seeds, in sprouted seeds, and after sprout homogenisation followed by incubation alone or in mixture with other sprouts. The proteolytic signatures differed markedly among the four species, and in the different treatment conditions. After homogenisation, legumain-like cleavage (after asparagine) increased in lentils, and proline-rich peptides increased in broccoli. For co-incubated homogenised sprouts, each species' homogenate significantly contributed 6 to 57% of proteolytic patterns in peptides of other species, with chickpea and broccoli homogenates notably releasing metabolic protein peptides from mung and lentils.Thus, germination, homogenisation and homogenate species mixtures can each contribute to proteolysis of seed peptides, potentially increasing digestibility and reducing allergenicity.
ABSTRACT Meiotic recombination is an important means of increasing genetic diversity by generating novel haplotypes in a population. Recombination separates linked loci extremely slowly in some regions, therefore genetic variants in high linkage disequilibrium may become co-adapted. Reciprocal recombination that separates co-adapted variants may generate a deleterious de novo haplotype that contributes to disease. We developed statistical methods to detect genomic regions of recombination excess in two different family-based study designs. We identified recombination in the Simons Simplex Collection in 273 simplex families with one child with autism spectrum disorder (ASD) and at least two unaffected children, in which recombinations can be mapped to the proband and contrasted with the recombination counts in unaffected siblings; and in 1,802 families with two children, where the number of recombinations identified can be contrasted with the expectation from a reference recombination map. Both strategies revealed a tail of low p-values for loci of interest that contrasted with the rest of the distribution. Permutation and bootstrap tests did not identify genome-wide primary findings in either cohort, but the most significant three-child cohort locus of recombination excess (between cadherin genes CDH4 and CDH26 ) replicated in the two-child cohort (p=0.01). While this replication strategy was not defined a priori , five of the most recombination enriched bins identified candidate ASD genes (p=0.02; WWOX, ADAMTS16, INSR, ADARB2, and HS6ST1) . Since the six identified loci were not identified as regions of high de novo copy number variation in the study cohort and no CNVs were detected in any of the recombinant probands in the identified regions, they represent candidates for reciprocal recombinations generating unfavourable haplotypes for these genes. This study highlights a previously unidentified source of clinical genetic variability contributing to the molecular aetiology of ASD. AUTHOR SUMMARY Autism spectrum disorder (ASD) is a constellation of neurodevelopmental disabilities characterised by deficits in social communication and repetitive patterns of behaviour. While ASD is highly heritable, its genetic basis is complex and poorly understood. While some highly penetrant types of genetic variation have been identified, most people with ASD carry a large number of variants that each contribute a small amount to their overall phenotype. In addition to mutations in individual genes, changes in the configuration of genes along a chromosome may contribute to ASD. Here, we describe a method for identifying regions where such new configurations have occurred through recombination and attempt to find regions where such changes are more common in autistic children than in their non-autistic siblings. We explore recombination as a source of genetic variation contributing to autism, which has potential to inform clinicians in providing services to autistic people and their families.
Quantifying model generalization to out-of-distribution data has been a longstanding challenge in machine learning. Addressing this issue is crucial for leveraging machine learning in scientific discovery, where models must generalize to new molecules or materials. Current methods typically split data into train and test sets using various criteria — temporal, sequence identity, scaffold, or random cross-validation — before evaluating model performance. However, with so many splitting criteria available, existing approaches offer limited guidance on selecting the most appropriate one, and they do not provide mechanisms for incorporating prior knowledge about the target deployment distribution(s).To tackle this problem, we have developed a novel metric, AU-GOOD, which quantifies expected model performance under conditions of increasing dissimilarity between train and test sets, while also accounting for prior knowledge about the target deployment distribution(s), when available. This metric is broadly applicable to biochemical entities, including proteins, small molecules, nucleic acids, or cells; as long as a relevant similarity function is defined for them. Recognizing the wide range of similarity functions used in biochemistry, we propose criteria to guide the selection of the most appropriate metric for partitioning. We also introduce a new partitioning algorithm that generates more challenging test sets, and we propose statistical methods for comparing models based on AU-GOOD.Finally, we demonstrate the insights that can be gained from this framework by applying it to two different use cases: developing predictors for pharmaceutical properties of small molecules, and using protein language models as embeddings to build biophysical property predictors.
Bioactive peptides are an important class of natural products with great functional versatility. Chemical modifications can improve their pharmacology, yet their structural diversity presents unique challenges for computational modeling. Furthermore, data for standard peptides (composed of the 20 canonical amino acids) is more abundant than for modified ones. Thus, we set out to identify whether predictive models fitted to standard data are reliable when applied to modified peptides. To do this, we first considered two critical aspects of the modeling problem, namely, choice of similarity function for guiding dataset partitioning and choice of molecular representation. Similarity-based dataset partitioning is an evaluation technique that divides the dataset into train and test subsets, such that the molecules in the test set are different from those used to fit the model.
MOTIVATION:Automated machine learning (AutoML) solutions can bridge the gap between new computational advances and their real-world applications by enabling experimental scientists to build their own custom models. We examine different steps in the development life-cycle of peptide bioactivity binary predictors and identify key steps where automation cannot only result in a more accessible method, but also more robust and interpretable evaluation leading to more trustworthy models. RESULTS:We present a new automated method for drawing negative peptides that achieves better balance between specificity and generalization than current alternatives. We study the effect of homology-based partitioning for generating the training and testing data subsets and demonstrate that model performance is overestimated when no such homology correction is used, which indicates that prior studies may have overestimated their performance when applied to new peptide sequences. We also conduct a systematic analysis of different protein language models as peptide representation methods and find that they can serve as better descriptors than a naive alternative, but that there is no significant difference across models with different sizes or algorithms. Finally, we demonstrate that an ensemble of optimized traditional machine learning algorithms can compete with more complex neural network models, while being more computationally efficient. We integrate these findings into AutoPeptideML, an easy-to-use AutoML tool to allow researchers without a computational background to build new predictive models for peptide bioactivity in a matter of minutes. AVAILABILITY AND IMPLEMENTATION:Source code, documentation, and data are available at https://github.com/IBM/AutoPeptideML and a dedicated web-server at http://peptide.ucd.ie/AutoPeptideML. A static version of the software to ensure the reproduction of the results is available at https://zenodo.org/records/13363975.
Lanthipeptides are a large group of ribosomally encoded peptides cyclized by thioether and methylene bridges, which include the lantibiotics, lanthipeptides with antimicrobial activity. There are over 100 experimentally characterized lanthipeptides, with at least 25 distinct cyclization bridging patterns. We set out to understand the evolutionary dynamics and diversity of lanthipeptides. We identified 977 peptides in 2785 bacterial genomes from short open-reading frames encoding lanthipeptide modifiable amino acids (C, S and T) that lay chromosomally adjacent to genes encoding proteins containing the cyclase domain. These appeared to be synthesized by both known and novel enzymatic combinations. Our predictor of bridging topology suggested 36 novel-predicted topologies, including a single-cysteine topology seen in 179 lanthionine or labionin containing peptides, which were enriched for histidine. Evidence that supported the relevance of the single-cysteine containing lanthipeptide precursors included the presence of the labionin motif among single cysteine peptides that clustered with labionin-associated synthetase domains, and the leader features of experimentally defined lanthipeptides that were shared with single cysteine predictions. Evolutionary rate variation among peptide subfamilies suggests that selection pressures for functional change differ among subfamilies. Lanthipeptides that have recently evolved specific novel features may represent a richer source of potential novel antimicrobials, since their target species may have had less time to evolve resistance.
Despite the importance of grains and legumes in the human diet, little is known regarding peptide release and the temporal changes of protease activities during seed germination. LC/MS-MS peptidomic analysis of two cultivars of germinating chickpea followed by computational analyses indicated cleavage dominated by proteases with a single position preference (mainly before (P1) or after cleavage (P1’): L at P2 (cysEP-like); R or K at P1 (vignain-like), N or Q at P1 (legumain-like); and previously unidentified K, R, A and S at P1’; A at P2’). While P1 N cleavages were relatively constant, P1’ K/R preferences were high in soaked garbanzo (kabuli) seeds, declined by four days, and returned at six days, but were much rarer in the brown (desi) cultivar. Late Embryogenesis Associated (LEA) peptides were markedly released during early germination. Vicilin peptides rich in glutamic acid near their N-termini markedly increased with germination, consistent with strong proteolytic resistance, even to human digestion, as indicated by analyses of separate datasets. Thus, this first peptidomics study of seed germination proteolytic profiles unveils a complex cultivar-specific programme of sequential activation and inactivation of a series of proteases, associated with the differential release of peptides from different protein groups.
Legume seed protein is an important source of nutrition, but generally it is less digestible than animal protein. Poor protein digestibility in legume seeds and seedlings may partly reflect defenses against herbivores. Protein changes during germination typically increase proteolysis and digestibility, by lowering the levels of anti-nutrient protease inhibitors, activating proteases, and breaking down storage proteins (including allergens). Germinating legume sprouts also show striking increases in free amino acids (especially asparagine), but their roles in host defense or other processes are not known. While the net effect of germination is generally to increase the digestibility of legume seed proteins, the extent of improvement in digestibility is species- and strain-dependent. Further research is needed to highlight which changes contribute most to improved digestibility of sprouted seeds. Such knowledge could guide the selection of varieties that are more digestible and also guide the development of food preparations that are more digestible, potentially combining germination with other factors altering digestibility, such as heating and fermentation. Techniques to characterize the shifts in protein make-up, activity and degradation during germination need to draw on traditional analytical approaches, complemented by proteomic and peptidomic analysis of mass spectrometry-identified peptide breakdown products.
During coronavirus infection, three non-structural proteins, nsp3, nsp4, and nsp6, are of great importance as they induce the formation of double-membrane vesicles where the replication and transcription of viral gRNA takes place, and the interaction of nsp3 and nsp4 lumenal regions triggers membrane pairing. However, their structural states are not well-understood. We investigated the interactions between nsp3 and nsp4 by predicting the structures of their lumenal regions individually and in complex using AlphaFold2 as implemented in ColabFold. The ColabFold prediction accuracy of the nsp3–nsp4 complex was increased compared to nsp3 alone and nsp4 alone. All cysteine residues in both lumenal regions were modelled to be involved in intramolecular disulphide bonds. A linker region in the nsp4 lumenal region emerged as crucial for the interaction, transitioning to a structured state when predicted in complex. The key interactions modelled between nsp3 and nsp4 appeared stable when the transmembrane regions of nsp3 and nsp4 were added to the modelling either alone or together. While molecular dynamics simulations (MD) demonstrated that the proposed model of the nsp3 lumenal region on its own is not stable, key interactions between nsp and nsp4 in the proposed complex model appeared stable after MD. Together, these observations suggest that the interaction is robust to different modelling conditions. Understanding the functional importance of the nsp4 linker region may have implications for the targeting of double membrane vesicle formation in controlling coronavirus infection.
Milk-derived peptides are known to confer anti-inflammatory effects. We hypothesised that milk-derived cell-penetrating peptides might modulate inflammation in useful ways. Using computational techniques, we identified and synthesised peptides from the milk protein Alpha-S1-casein that were predicted to be cell-penetrating using a machine learning predictor. We modified the interpretation of the prediction results to consider the effects of histidine. Peptides were then selected for testing to determine their cell penetrability and anti-inflammatory effects using HeLa cells and J774.2 mouse macrophage cell lines. The selected peptides all showed cell penetrating behaviour, as judged using confocal microscopy of fluorescently labelled peptides. None of the peptides had an effect on either the NF-κB transcription factor or TNFα and IL-1β secretion. Thus, the identified milk-derived sequences have the ability to be internalised into the cell without affecting cell homeostatic mechanisms such as NF-κB activation. These peptides are worthy of further investigation for other potential bioactivities or as a naturally derived carrier to promote the cellular internalisation of other active peptides.
The first reported receptor for SARS-CoV-2 on host cells was the angiotensin-converting enzyme 2 (ACE2). However, the viral spike protein also has an RGD motif, suggesting that cell surface integrins may be co-receptors. We examined the sequences of ACE2 and integrins with the Eukaryotic Linear Motif (ELM) resource and identified candidate short linear motifs (SLiMs) in their short, unstructured, cytosolic tails with potential roles in endocytosis, membrane dynamics, autophagy, cytoskeleton, and cell signaling. These SLiM candidates are highly conserved in vertebrates and may interact with the μ2 subunit of the endocytosis-associated AP2 adaptor complex, as well as with various protein domains (namely, I-BAR, LC3, PDZ, PTB, and SH2) found in human signaling and regulatory proteins. Several motifs overlap in the tail sequences, suggesting that they may act as molecular switches, such as in response to tyrosine phosphorylation status. Candidate LC3-interacting region (LIR) motifs are present in the tails of integrin β3 and ACE2, suggesting that these proteins could directly recruit autophagy components. Our findings identify several molecular links and testable hypotheses that could uncover mechanisms of SARS-CoV-2 attachment, entry, and replication against which it may be possible to develop host-directed therapies that dampen viral infection and disease progression. Several of these SLiMs have now been validated to mediate the predicted peptide interactions.
Cationic antimicrobial peptides have raised interest as attractive alternatives to classical antibiotics, and also have utility in preventing food spoilage. We set out to enrich cationic antimicrobial peptides from milk hydrolysates using gels containing various ratios of anionic pectin/alginate. All processes were carried out with foodgrade materials in order to suggest food-safe methods suited for producing food ingredients or supplements. Hydrolysed caseinate peptides retained in the gel fraction, identified by mass spectrometry, were enriched for potential antimicrobial peptides, as judged by a computational predictor of antimicrobial activity. Peptides retained in a 60:40 pectin:alginate gel fraction had a strong antimicrobial effect against 8 tested bacterial strains with a minimal inhibitory concentration of 1.5?5 mg/mL, while the unfractionated hydrolysate only had a detectable effect in one of the eight strains. Among 110 predicted antimicrobial peptides in the gel fraction, four are known antimicrobial peptides, HKEMPFPK, TTMPLW, YYQQKPVA and AVPYPQR. These results highlight the potential of pectin/alginate food-gels based processes as safe, fast, cost-effective methods to separate and enrich for antimicrobial peptides from complex food protein hydrolysates. (funded by Science Foundation Ireland (12/RI/2346 [3]).