This study focuses on understanding the transcriptional heterogeneity of activated platelets and its impact on diseases such as sepsis, COVID-19, and systemic lupus erythematosus (SLE). Recognizing the limited knowledge in this area, our research aims to dissect the complex transcriptional profiles of activated platelets to aid in developing targeted therapies for abnormal and pathogenic platelet subtypes. We analyzed single-cell transcriptional profiles from 47,977 platelets derived from 413 samples of patients with these diseases, utilizing Deep Neural Network (DNN) and eXtreme Gradient Boosting (XGB) to distinguish transcriptomic signatures predictive of fatal or survival outcomes. Our approach included source data annotations and platelet markers, along with SingleR and Seurat for comprehensive profiling. Additionally, we employed Uniform Manifold Approximation and Projection (UMAP) for effective dimensionality reduction and visualization, aiding in the identification of various platelet subtypes and their relation to disease severity and patient outcomes. Our results highlighted distinct platelet subpopulations that correlate with disease severity, revealing that changes in platelet transcription patterns can intensify endotheliopathy, increasing the risk of coagulation in fatal cases. Moreover, these changes may impact lymphocyte function, indicating a more extensive role for platelets in inflammatory and immune responses. This study identifies crucial biomarkers of platelet heterogeneity in serious health conditions, paving the way for innovative therapeutic approaches targeting platelet activation, which could improve patient outcomes in diseases characterized by altered platelet function.
B- and T-lymphocyte attenuator (BTLA; CD272) is an immunoglobulin superfamily member and part of a family of checkpoint inhibitory receptors that negatively regulate immune cell activation. The natural ligand for BTLA is herpes virus entry mediator (HVEM; TNFRSF14), and binding of HVEM to BTLA leads to attenuation of lymphocyte activation. In this study, we evaluated the role of BTLA and HVEM expression in the pathogenesis of systemic lupus erythematosus (SLE), a multisystem autoimmune disease. Peripheral blood mononuclear cells from healthy volunteers (N = 7) were evaluated by mass cytometry by time-of-flight to establish baseline expression of BTLA and HVEM on human lymphocytes compared with patients with SLE during a self-reported flare (N = 5). High levels of BTLA protein were observed on B cells, CD4+, and CD8+ T cells, and plasmacytoid dendritic cells in healthy participants. HVEM protein levels were lower in patients with SLE compared with healthy participants, while BTLA levels were similar between SLE and healthy groups. Correlations of BTLA-HVEM hub genes' expression with patient and disease characteristics were also analyzed using whole blood gene expression data from patients with SLE (N = 1,760) and compared with healthy participants (N = 60). HVEM, being one of the SLE-associated genes, showed an exceptionally strong negative association with disease activity. Several other genes in the BTLA-HVEM signaling network were strongly (negative or positive) correlated, while BTLA had a low association with disease activity. Collectively, these data provide a clinical rationale for targeting BTLA with an agonist in SLE patients with low HVEM expression.
Abstract Legionella pneumophila utilizes the Dot/Icm type IVB secretion system to deliver hundreds of effector proteins inside eukaryotic cells to ensure intracellular replication. Our understanding of the molecular functions of the largest pathogenic arsenal known to the bacterial world remains incomplete. By leveraging advancements in 3D protein structure prediction, we provide a comprehensive structural analysis of 368 L. pneumophila effectors, representing a global atlas of predicted functional domains summarized in a database ( https://pathogens3d.org/legionella-pneumophila ). Our analysis identified 157 types of diverse functional domains in 287 effectors, including 159 effectors with no prior functional annotations. Furthermore, we identified 35 cryptic domains in 30 effector models that have no similarity with experimentally structurally characterized proteins, thus, hinting at novel functionalities. Using this analysis, we demonstrate the activity of thirteen functional domains, including three cryptic domains, predicted in L. pneumophila effectors to cause growth defects in the Saccharomyces cerevisiae model system. This illustrates an emerging strategy of exploring synergies between predictions and targeted experimental approaches in elucidating novel effector activities involved in infection.
The receptor domains of Toll-like receptors (TLRs) are characterized by a solenoid-like structure composed of tandem repeats of α/β units known as Leucine Rich Repeats (LRRs). LRR proteins form large paralogous families, with nearly 400 in the human genome alone, all sharing similar semi-regular solenoid-like structures. Despite this structural similarity, they exhibit remarkable diversity in binding specificity. For TLR receptors, this includes a range of pathogen-associated molecular patterns (PAMPs), while other LRR proteins bind an extensive array of ligands, including proteins, DNA, RNA, and small molecules. The LRR domains contain repeats that have similar, yet not identical, 3D structures and patterns of conserved residues. Through in-depth analysis of sequence and structural conservation in individual repeats of human TLRs, we demonstrate that even subtle variations between these repeats alter the local solenoid structure, leading to significant functional changes. Variations in repeat length and defining patterns result in local changes in curvature and the emergence of structural features such as loops, cavities, or specific interaction interfaces. Understanding how divergence in LRR repeats influences their functional roles can provide deeper insights into their binding mechanisms, including interactions with unknown ligands, both in humans and across the diverse phylogenetic spectrum of animals that rely on their TLR repertoires for immune activation. ### Competing Interest Statement The authors have declared no competing interest.
Since late 2020, SARS-CoV-2 variants have regularly emerged with competitive and phenotypic differences from previously circulating strains, sometimes with the potential to escape from immunity produced by prior ex-posure and infection. The Early Detection group is one of the constituent groups of the US National Institutes of Health National Institute of Allergy and Infectious Dis-eases SARS-CoV-2 Assessment of Viral Evolution pro-gram. The group uses bioinformatic methods to monitor the emergence, spread, and potential phenotypic prop-erties of emerging and circulating strains to identify the most relevant variants for experimental groups within the program to phenotypically characterize. Since April 2021, the group has prioritized variants monthly. Prioritization successes include rapidly identifying most major variants of SARS-CoV-2 and providing experimental groups with-in the National Institutes of Health program easy access to regularly updated information on the recent evolution and epidemiology of SARS-CoV-2 that can be used to guide phenotypic investigations.
Leucine Rich Repeat (LRR) domains, are present in hundreds of thousands of proteins across all kingdoms of life and are typically involved in protein-protein interactions and ligand recognition. LRR domains are classified into eight classes and when examined in three dimensions seven of them form curved solenoid-like super-helices, also described as toruses, with a beta sheet on the concave (inside) and stacked alpha-helices on the convex (outside) of the torus. Here we present an overview of the least characterized 8th class of LRR proteins, the TpLRR-like LRRs, named after the Treponema pallidum protein Tp0225. Proteins from the TpLRR class differ from the proteins in all other known LRR classes by having a flipped curvature, with the beta sheet on the convex side of the torus and irregular secondary structure instead of helices on the opposite, now concave site. TpLRR proteins also present highly divergent sequence pattern of individual repeats and can associate with specific types of additional domains. Several of the characterized proteins from this class, specifically the BspA-like proteins, were found in human bacterial and protozoan pathogens, playing an important role in the interactions between the pathogens and the host immune system. In this paper we surveyed all existing experimental structures and selected AlphaFold models of the best-known proteins containing this class of LRR repeats, analyzing the relation between the pattern of conserved residues, specific structural features and functions of these proteins.
The study focuses on understanding the transcriptional heterogeneity of activated platelets and its impact on diseases like sepsis, COVID-19, and systemic lupus erythematosus (SLE). Recognizing the limited knowledge in this area, our research aims to dissect the complex transcriptional profiles of activated platelets to aid in developing targeted therapies for abnormal and pathogenic platelet subtypes. We analyzed single-cell transcriptional profiles from 47,977 platelets derived from 413 samples of patients with these diseases, utilizing Deep Neural Network (DNN) and eXtreme Gradient Boosting (XGB) to distinguish transcriptomic signatures predictive of fatal or survival outcomes. Our approach included source data annotations and platelet markers, along with SingleR and Seurat for comprehensive profiling. Additionally, we employed Uniform Manifold Approximation and Projection (UMAP) for effective dimensionality reduction and visualization, aiding in the identification of various platelet subtypes and their relation to disease severity and patient outcomes. Our results highlighted distinct platelet subpopulations that correlate with disease severity, revealing that changes in platelet transcription patterns can intensify endotheliopathy, increasing the risk of coagulation in fatal cases. Moreover, these changes also seem to impact lymphocyte function, indicating a more extensive role for platelets in inflammatory and immune responses. This study sheds light on the crucial role of platelet heterogeneity in serious health conditions, paving the way for innovative therapeutic approaches targeting platelet activation, which could potentially improve patient outcomes in diseases characterized by altered platelet function.
Both gender and smoking are correlated with prevalence and outcomes in many types of cancers. Tobacco smoke is a known carcinogen through its genotoxicity but can also affect cancer progression through its effect on the immune system. In this study, we aim to evaluate the hypothesis that the effects of smoking on the tumor immune microenvironment will be influenced differently by gender using large-scale analysis of publicly available cancer datasets. We used The Cancer Genomic Atlas (TCGA) datasets (n = 2724) to analyze effects of smoking on different cancer immune subtypes and the relative abundance of immune cell types between male and female cancer patients. We further validated our results by analyzing additional datasets, including Expression Project for Oncology (expO) bulk RNA-seq dataset (n = 1118) and single-cell RNA-seq dataset (n = 14). Results of our study indicate that in female patients, two immune subtypes, C1 and C2, are respectively over and under abundant in smokers vs. never smokers. In males, the only significant difference is underabundance of the C6 subtype in smokers. We identified gender-specific differences in the population of immune cell types between smokers and never smokers in all TCGA and expO cancer types. Increased plasma cell population was identified as the most consistent feature distinguishing smokers and never smokers, especially in current female smokers based on both TCGA and expO data. Our analysis of existing single-cell RNA-seq data further revealed that smoking differentially affects the gene expression profile of cancer patients based on the immune cell type and gender. In our analysis, female and male smokers show different smoking-induced patterns of immune cells in tumor microenvironment. Besides, our results suggest cancer tissues directly exposed to tobacco smoke undergo the most significant changes, but all other tissue types are affected as well. Findings of current study also indicate that changes in the populations of plasma cells and their correlations to survival outcomes are stronger in female current smokers, with implications for cancer immunotherapy of women smokers. In conclusion, results of this study can be used to develop personalized treatment plans for cancer patients who smoke, particularly women smokers, taking into account the unique immune cell profile of their tumors.
Background Sepsis mortality has remained unchanged for greater than a decade, and early recognition continues to be the most important factor in mortality outcome. Plasma resistin concentration is increased in sepsis, but its mechanism and clinical relevance is unclear. As one function, resistin interacts with toll-like receptor 4 in competition with lipopolysaccharide, a main component of the gram-negative bacterial cell wall. It is not known if the type of infection leading to sepsis influences resistin production. The objective of this study was to investigate whether 1) early plasma resistin concentration can predict mortality, 2) elevated plasma resistin concentration is associated with clinical disease severity scores, such as SOFA, mSOFA and APACHE II, and 3) plasma resistin concentrations differ between gram negative versus other etiologies of sepsis. Methods This was an exploratory study in the framework of a prospective observational design. Peripheral venous blood samples were obtained from subjects admitted to the intensive care unit at clinical recognition of sepsis (0 hour) and at 6 and 24 hours. Vasopressor utilization was not a requirement for inclusion. Plasma was analyzed for resistin concentration by ELISA. Cytokine concentrations including IL-6, IL-8, and IL-10 were determined by cytokine bead array. Cytokine data were evaluated against publicly available sepsis RNA expression datasets to compare protein versus RNA expression levels in predicting clinical disease state. Clinical data were collected from electronic health records for clinical severity index calculations and context for interpretation of resistin and cytokine concentrations. Subjects were followed up to 60 days, or until death, whichever came first. Statistical analysis was completed with R package and SPSS software. Results Resistin levels were elevated in subjects admitted to the intensive care unit with sepsis. Four-hundred subjects were screened with 45 subjects included in the final analysis. Thirteen of 45 patients were non-survivors. Mortality within 60 days correlated with significantly higher resistin concentrations than in survivors. A resistin concentration of >126 ng/mL at clinical recognition of sepsis and >197 ng/mL within the first 24 hours were associated with mortality within 60 days with an area under the curve of 0.82 and 0.88, respectively. Most subjects with resistin concentration greater than these threshold values were deceased prior to 30 days. Resistin concentrations correlated with SOFA, mSOFA, and APACHE II scores in addition to having association with increases in inflammatory and sepsis biomarkers. These associations were validated with analysis of RNA expression datasets. Conclusion Plasma resistin concentrations of >126 ng/mL at clinical recognition of sepsis and >197 ng/mL within the first 24 hours of clinical sepsis recognition are associated with all-cause mortality. Resistin concentration within this timeframe also has comparable mortality association to well-validated clinical severity indices of SOFA, mSOFA, and APACHE II scores.
The global emergence of many severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) variants jeopardizes the protective antiviral immunity induced after infection or vaccination. To address the public health threat caused by the increasing SARS-CoV-2 genomic diversity, the National Institute of Allergy and Infectious Diseases within the National Institutes of Health established the SARS-CoV-2 Assessment of Viral Evolution (SAVE) programme. This effort was designed to provide a real-time risk assessment of SARS-CoV-2 variants that could potentially affect the transmission, virulence, and resistance to infection- and vaccine-induced immunity. The SAVE programme is a critical data-generating component of the US Government SARS-CoV-2 Interagency Group to assess implications of SARS-CoV-2 variants on diagnostics, vaccines and therapeutics, and for communicating public health risk. Here we describe the coordinated approach used to identify and curate data about emerging variants, their impact on immunity and effects on vaccine protection using animal models. We report the development of reagents, methodologies, models and notable findings facilitated by this collaborative approach and identify future challenges. This programme is a template for the response to rapidly evolving pathogens with pandemic potential by monitoring viral evolution in the human population to identify variants that could reduce the effectiveness of countermeasures. The SARS-CoV-2 Assessment of Viral Evolution (SAVE) programme provides a real-time risk assessment of SARS-CoV-2 variants with the potential to affect transmission, virulence and resistance to infection- and vaccine-induced immunity.
The unprecedented growth of publicly available SARS-CoV-2 genome sequence data has increased the demand for effective and accessible SARS-CoV-2 data analysis and visualization tools. The majority of the currently available tools either require computational expertise to deploy them or limit user input to preselected subsets of SARS-CoV-2 genomes. To address these limitations, we developed ViralVar, a publicly available, point-and-click webtool that gives users the freedom to investigate and visualize user-selected subsets of SARS-CoV-2 genomes obtained from the GISAID public database. ViralVar has two primary features that enable: (1) the visualization of the spatiotemporal dynamics of SARS-CoV-2 lineages and (2) a structural/functional analysis of genomic mutations. As proof-of-principle, ViralVar was used to explore the evolution of the SARS-CoV-2 pandemic in the USA in pediatric, adult, and elderly populations (n > 1.7 million genomes). Whereas the spatiotemporal dynamics of the variants did not differ between these age groups, several USA-specific sublineages arose relative to the rest of the world. Our development and utilization of ViralVar to provide insights on the evolution of SARS-CoV-2 in the USA demonstrates the importance of developing accessible tools to facilitate and accelerate the large-scale surveillance of circulating pathogens.
Most attention in the surveillance of evolving SARS-CoV-2 genome has been centered on nucleotide substitutions in the spike glycoprotein. We show that, as the pandemic extends into its second year, the numbers and ratio of genomes with in-frame insertions and deletions (indels) increases significantly, especially among the variants of concern (VOCs). Monitoring of the SARS-CoV-2 genome evolution shows that co-occurrence (i.e., highly correlated presence) of indels, especially deletions on spike N-terminal domain and non-structural protein 6 (NSP6) is a shared feature in several VOCs such as Alpha, Beta, Delta, and Omicron. Indels distribution is correlated with spike mutations associated with immune escape and growth in the number of genomes with indels coincides with the increasing population resistance due to vaccination and previous infections. Indels occur most frequently in the spike, but also in other proteins, especially those involved in interactions with the host immune system. We also showed that indels concentrate in regions of individual SARS-CoV-2 proteins known as hypervariable regions (HVRs) that are mostly located in specific loop regions. Structural analysis suggests that indels remodel viral proteins' surfaces at common epitopes and interaction interfaces, affecting the virus' interactions with host proteins. We hypothesize that the increased frequency of indels, the non-random distribution of them and their independent co-occurrence in several VOCs is another mechanism of response to elevated global population immunity.
The bacterium Vibrio cholerae causes gastrointestinal, wound, and skin infections. The motility-associated killing factor A (MakA) was recently shown to be cytotoxic against colon, prostate, and other cancer cells.
Proteins sample a multitude of different conformations by undergoing small- and large-scale conformational changes that are often intrinsic to their functions. Information about these changes is often captured in the Protein Data Bank by the apparently redundant deposition of independent structural solutions of identical proteins. Here, we mine these data to examine the conservation of large-scale conformational changes between homologous proteins. This is important for both practical reasons, such as predicting alternative conformations of a protein by comparative modeling, and conceptual reasons, such as understanding the extent of conservation of different features in evolution. To study this question, we introduce a novel approach to compare conformational changes between proteins by the comparison of their difference distance maps (DDMs). We found that proteins undergoing similar conformational changes have similar DDMs and that this similarity could be quantified by the correlation between the DDMs. By comparing the DDMs of homologous protein pairs, we found that large-scale conformational changes show a high level of conservation across a broad range of sequence identities. This shows that conformational space is usually conserved between homologs, even relatively distant ones.
The search for drugs against COVID-19 and other diseases caused by coronaviruses focuses on the most conserved and essential proteins, mainly the main (M pro ) and the papain-like (PL pro ) proteases and the RNA-dependent RNA polymerase (RdRp). Nirmatrelvir, an inhibitor for M pro , was recently approved by FDA as a part of a two-drug combination, Paxlovid, and many more drugs are in various stages of development. Multiple candidates for the PL pro inhibitors are being studied, but none have yet progressed to clinical trials. Several repurposed inhibitors of RdRp are already in use. We can expect that once anti-COVID-19 drugs become widely used, resistant variants of SARS-CoV-2 will emerge, and we already see that for the drugs targeting SARS-CoV-2 RdRp. We hypothesize that emergence of such variants can be anticipated by identifying possible escape mutations already present in the existing populations of viruses. Our group previously developed the coronavirus3D server ( https://coronavirus3d.org ), tracking the evolution of SARS-CoV-2 in the context of the three-dimensional structures of its proteins. Here we introduce dedicated pages tracking the emergence of potential drug resistant mutations to M pro and PL pro , showing that such mutations are already circulating in the SARS-CoV-2 viral population. With regular updates, the drug resistance tracker provides an easy way to monitor and potentially predict the emergence of drug resistance-conferring mutations in the SARS-CoV-2 virus.
Systemic infections, especially in patients with chronic diseases, may result in sepsis: an explosive, uncoordinated immune response that can lead to multisystem organ failure with a high mortality rate. Patients with similar clinical phenotypes or sepsis biomarker expression upon diagnosis may have different outcomes, suggesting that the dynamics of sepsis is critical in disease progression. A within-subject study of patients with Gram-negative bacterial sepsis with surviving and fatal outcomes was designed and single-cell transcriptomic analyses of peripheral blood mononuclear cells (PBMC) collected during the critical period between sepsis diagnosis and 6 h were performed. The single-cell observations in the study are consistent with trends from public datasets but also identify dynamic effects in individual cell subsets that change within hours. It is shown that platelet and erythroid precursor responses are drivers of fatal sepsis, with transcriptional signatures that are shared with severe COVID-19 disease. It is also shown that hypoxic stress is a driving factor in immune and metabolic dysfunction of monocytes and erythroid precursors. Last, the data support CD52 as a prognostic biomarker and therapeutic target for sepsis as its expression dynamically increases in lymphocytes and correlates with improved sepsis outcomes. In conclusion, this study describes the first single-cell study that analyzed short-term temporal changes in the immune cell populations and their characteristics in surviving or fatal sepsis. Tracking temporal expression changes in specific cell types could lead to more accurate predictions of sepsis outcomes and identify molecular biomarkers and pathways that could be therapeutically controlled to improve the sepsis trajectory toward better outcomes.
For more than a decade, the Joint Center for Structural Genomics ( JCSG ; www.jcsg.org ) worked toward increased three‐dimensional structure coverage of the protein universe. This coordinated quest was one of the main goals of the four high‐throughput ( HT ) structure determination centers of the Protein Structure Initiative ( PSI ; www.nigms.nih.gov/Research/specificareas/PSI ). To achieve the goals of the PSI , the JCSG made use of the complementarity of structure determination by X‐ray crystallography and nuclear magnetic resonance ( NMR ) spectroscopy to increase and diversify the range of targets entering the HT structure determination pipeline. The overall strategy, for both techniques, was to determine atomic resolution structures for representatives of large protein families, as defined by the Pfam database, which had no structural coverage and could make significant contributions to biological and biomedical research. Furthermore, the experimental structures could be leveraged by homology modeling to further expand the structural coverage of the protein universe and increase biological insights. Here, we describe what could be achieved by this structural genomics approach, using as an illustration the contributions from 20 NMR structure determinations out of a total of 98 JCSG NMR structures, which were selected because they are the first three‐dimensional structure representations of the respective Pfam protein families. The information from this small sample is representative for the overall results from crystal and NMR structure determination in the JCSG . There are five new folds, which were classified as domains of unknown functions ( DUF ), three of the proteins could be functionally annotated based on three‐dimensional structure similarity with previously characterized proteins, and 12 proteins showed only limited similarity with previous deposits in the Protein Data Bank ( PDB ) and were classified as DUF s.
The RNA binding domain abundant in apicomplexans (RAP) is a protein domain identified in a diverse group of proteins, called RAP proteins, many of which have been shown to be involved in RNA binding. To understand the expansion and potential function of the RAP proteins, we conducted a hidden Markov model based screen among the proteomes of 54 eukaryotes, 17 bacteria and 12 archaea. We demonstrated that the domain is present in closely and distantly related organisms with particular expansions in Alveolata and Chlorophyta, and are not unique to Apicomplexa as previously believed. All RAP proteins identified can be decomposed into two parts. In the N-terminal region, the presence of variable helical repeats seems to participate in the specific targeting of diverse RNAs, while the RAP domain is mostly identified in the C-terminal region and is highly conserved across the different phylogenetic groups studied. Several conserved residues defining the signature motif could be crucial to ensure the function(s) of the RAP proteins. Modelling of RAP domains in apicomplexan parasites confirmed an ⍺/β structure of a restriction endonuclease-like fold. The phylogenetic trees generated from multiple alignment of RAP domains and full-length proteins from various distantly related eukaryotes indicated a complex evolutionary history of this family. We further discuss these results to assess the potential function of this protein family in apicomplexan parasites.
Several plastic degrading enzymes have been described in the literature, most notably PETases that are capable of hydrolyzing polyethylene terephthalate (PET) plastic. One of them, the PETase from Ideonella sakaiensis, a bacterium isolated from environmental samples within a PET bottle recycling site, was a subject of extensive studies. To test how widespread PETase functionality is in other bacterial communities, we used a cascade of BLAST searches in the JGI metagenomic datasets and showed that close homologs of I. sakaiensis PETase can also be found in other metagenomic environmental samples from both human-affected and relatively pristine sites. To confirm their classification as putative PETases, we verified that the newly identified proteins have the PETase sequence signatures common to known PETases and that phylogenetic analyses group them with the experimentally characterized PETases. Additionally, docking analysis was performed in order to further confirm the functional assignment of the putative environmental PETases.
There is a wide, and continuously widening, gap between the number of proteins known only by their amino acid sequence versus those structurally characterized by direct experiment. To close this gap, we mostly rely on homology-based inference and modeling to reason about the structures of the uncharacterized proteins by using structures of homologous proteins as templates. With the rapidly growing size of the Protein Data Bank, there are often multiple choices of templates, including multiple sets of coordinates from the same protein. The substantial conformational differences observed between different experimental structures of the same protein often reflect function related structural flexibility. Thus, depending on the questions being asked, using distant homologs, or coordinate sets with lower resolution but solved in the appropriate functional form, as templates may be more informative. The ModFlex server (https://modflex.org/) addresses this seldom mentioned gap in the standard homology modeling approach by providing the user with an interface with multiple options and tools to select the most relevant template and explore the range of structural diversity in the available templates. ModFlex is closely integrated with a range of other programs and servers developed in our group for the analysis and visualization of protein structural flexibility and divergence.