AbstractBackgroundThe paucity of data on the contemporary causes of serious infection among the world’s most vulnerable children means the landscape of emerging paediatric infectious disease remains largely undefined and out of focus on the global vaccine research and development agenda.MethodsWe aimed to partially define the paediatric infectious disease landscape in a typical low-income setting in sub-Saharan Africa in Kilifi, Kenya by simultaneously estimating antibody prevalence for 38 infectious diseases using a longitudinal birth cohort that was sampled between 2002 and 2008 and a paediatric inpatient cohort that was sampled between 2006 and 2017.FindingsAmong the infectious diseases with the highest antibody prevalence in the first year of life were vaccine-preventable diseases such as RSV (57.4%), mumps (31.5%) and influenza H3N2 (37.3%). Antibody prevalence forPlasmodium falciparumshifted substantially over time, from 47% in the mid 2000s to 13% approximately 10 years later corresponding to a documented decline in parasite transmission. A high prevalence of antibodies was also observed in the first year of life for infections for which no licenced vaccines are currently available, including norovirus (34.2%), cytomegalovirus (44.7%), EBV (29.3%) and coxsackie B virus (40.7%). The prevalence to antibodies to vaccine antigens in the local immunisation schedule was generally high but varied by antigen.InterpretationThe data show a high and temporally stable infection burden of RSV, mumps and influenza, providing a compelling evidence base to support progress towards the introduction of these vaccines into the local immunization schedule. The high prevalence of norovirus, EBV, CMV and Coxsackie B provide rationale for increased vaccine research and development investment.FundingThis research was funded by the Wellcome Trust (grant no. WT105882MA).
Strengthening malaria surveillance is a key intervention needed to reduce the global disease burden. Reliable serological markers of recent malaria exposure could improve current surveillance methods by allowing for accurate estimates of infection incidence from limited data. We studied the IgG antibody response to 111 Plasmodium falciparum proteins in 65 adult travellers followed longitudinally after a natural malaria infection in complete absence of re-exposure. We identified a combination of five serological markers that detect exposure within the previous three months with >80% sensitivity and specificity. Using mathematical modelling, we examined the antibody kinetics and determined that responses informative of recent exposure display several distinct characteristics: rapid initial boosting and decay, less inter-individual variation in response kinetics, and minimal persistence over time. Such serological exposure markers could be incorporated into routine malaria surveillance to guide efforts for malaria control and elimination.
Protein microarrays are versatile tools for high throughput study of the human proteome, but systematic and non-systematic sources of bias constrain optimal interpretation and the ultimate utility of the data. Published guidelines to limit technical variability whilst maintaining important biological variation favour DNA-based microarrays that often differ fundamentally in their experimental design. Rigorous tools to guide background correction, the quantification of within-sample variation, normalisation, and batch correction specifically for protein microarrays are limited, require extensive investigation and are not centrally accessible. Here, we develop a generic one-stop-shop pre-processing suite for protein microarrays that is compatible with data from the major protein microarray scanners. Our graphical and tabular interfaces facilitate a detailed inspection of data and are coupled with supporting guidelines that enable users to select the most appropriate algorithms to systematically address bias arising in customized experiments. The localization and distribution of background signal intensities determine the optimal correction strategy. A novel function overcomes the limitations in the interpretation of the coefficient of variation when signal intensities are at the lower end of the detection threshold. We demonstrate essential considerations in the experimental design and their impact on a range of algorithms for normalization and minimization of batch effects. Our user-friendly interactive web-based platform eliminates the need for prowess in programming. The open-source R interface includes illustrative examples, generates an auditable record, enables reproducibility, and can incorporate additional custom scripts through its online repository. This versatility will enhance its broad uptake in the infectious disease and vaccine development community.
To evaluate how clinical chemistry measurements from the REFALS phase 3 trial (NCT03505021) correlate with ALS disease course in order to further elucidate their potential as biomarkers of disease.
Leveraging high-dimensional molecular datasets can help us develop mechanistic insight into associations between genetic variants and complex traits. In this study, we integrated human proteome data derived from brain tissue to evaluate whether targeted proteins putatively mediate the effects of genetic variants on seven neurological phenotypes (Alzheimer disease, amyotrophic lateral sclerosis, depression, insomnia, intelligence, neuroticism, and schizophrenia). Applying the principles of Mendelian randomization (MR) systematically across the genome highlighted 43 effects between genetically predicted proteins derived from the dorsolateral prefrontal cortex and these outcomes. Furthermore, genetic colocalization provided evidence that the same causal variant at 12 of these loci was responsible for variation in both protein and neurological phenotype. This included genes such as DCC, which encodes the netrin-1 receptor and has an important role in the development of the nervous system (p = 4.29 × 10-11 with neuroticism), as well as SARM1, which has been previously implicated in axonal degeneration (p = 1.76 × 10-08 with amyotrophic lateral sclerosis). We additionally conducted a phenome-wide MR study for each of these 12 genes to assess potential pleiotropic effects on 700 complex traits and diseases. Our findings suggest that genes such as SNX32, which was initially associated with increased risk of Alzheimer disease, may potentially influence other complex traits in the opposite direction. In contrast, genes such as CTSH (which was also associated with Alzheimer disease) and SARM1 may make worthwhile therapeutic targets because they did not have genetically predicted effects on any of the other phenotypes after correcting for multiple testing.
Strengthening malaria surveillance is a key intervention needed to reduce the global disease burden. Reliable serological markers of recent malaria exposure could dramatically improve current surveillance methods by allowing for accurate estimates of infection incidence from limited data. We studied the IgG antibody response to 111 Plasmodium falciparum proteins in travellers followed longitudinally after a natural malaria infection in complete absence of re-exposure. We identified a novel combination of five serological markers (GAMA, MSP1, MSPDBL1 C- and N-terminal, and PfSEA1) that detect exposure within the previous 3-months with >80% sensitivity and specificity. Using mathematical modelling, we examined the antibody kinetics and determined that responses informative of recent exposure display several distinct characteristics: rapid initial boosting and decay, less inter-individual variation in response kinetics, and minimal persistence over time. These serological exposure markers can be incorporated into routine malaria surveillance to guide efforts for malaria control and elimination.
Background. Human rhinovirus (HRV) is the most common cause of the common cold but may also lead to more severe respiratory illness in vulnerable populations. The epidemiology and genetic diversity of HRV within a school setting have not been previously described. The objective of this study was to characterize HRV molecular epidemiology in a primary school in a rural location of Kenya. Methods. Between May 2017 and April 2018, over 3 school terms, we collected 1859 nasopharyngeal swabs (NPS) from pupils and teachers with symptoms of acute respiratory infection in a public primary school in Kilifi County, coastal Kenya. The samples were tested for HRV using real-time reverse transcription polymerase chain reaction. HRV-positive samples were sequenced in the VP4/VP2 coding region for species and genotype classification. Results. A total of 307 NPS (16.4%) from 164 individuals were HRV positive, and 253 (82.4%) were successfully sequenced. The proportion of HRV in the lower primary classes was higher (19.8%) than upper primary classes (12.2%; P < .001). HRV-A was the most common species (134/253; 53.0%), followed by HRV-C (73/253; 28.9%) and HRV-B (46/253; 18.2%). Phylogenetic analysis identified 47 HRV genotypes. The most common genotypes were A2 and B70. Numerous (up to 22 in 1 school term) genotypes circulated simultaneously, there was no individual re-infection with the same genotype, and no genotype was detected in all 3 school terms. Conclusions. HRV was frequently detected among school-going children with mild acute respiratory illness symptoms, particularly in the younger age groups (<5-year-olds). Multiple HRV introductions were observed that were characterized by considerable genotype diversity.
HIV-exposed uninfected (HEU) infants are disproportionately at a higher risk of morbidity and mortality, as compared to HIV-unexposed uninfected (HUU) infants. Here, we used transcriptional profiling of peripheral blood mononuclear cells to determine immunological signatures of in utero HIV exposure. We identified 262 differentially expressed genes (DEGs) in HEU compared to HUU infants. Weighted gene co-expression network analysis (WGCNA) identified six modules that had significant associations with clinical traits. Functional enrichment analysis on both DEGs and the six significantly associated modules revealed an enrichment of G-protein coupled receptors and the immune system, specifically affecting neutrophil function and antibacterial responses. Additionally, malaria pathogenicity genes (thrombospondin 1-(THBS 1), interleukin 6 (IL6), and arginine decarboxylase 2 (ADC2)) were down-regulated. Of interest, the down-regulated immunity genes were positively correlated to the expression of epigenetic factors of the histone family and high-mobility group protein B2 (HMGB2), suggesting their role in the dysregulation of the HEU transcriptional landscape. Overall, we show that genes primarily associated with neutrophil mediated immunity were repressed in the HEU infants. Our results suggest that this could be a contributing factor to the increased susceptibility to bacterial infections associated with higher morbidity and mortality commonly reported in HEU infants.
Rationale Pneumonia is a leading cause of mortality in infants and young children. The mechanisms that lead to mortality in these children are poorly understood. Studies of the cellular immunology of the infant airway have traditionally been hindered by the limited sample volumes available from the young, frail children who are admitted to hospital with pneumonia. This is further compounded by the relatively low frequencies of certain immune cell phenotypes that are thought to be critical to the clinical outcome of pneumonia. To address this, we developed a novel in-silico deconvolution method for inferring the frequencies of immune cell phenotypes in the airway of children with different survival outcomes using proteomic data. Methods Using high-resolution mass spectrometry, we identified > 1,000 proteins expressed in the airways of children who were admitted to hospital with clinical pneumonia. 61 of these children were discharged from hospital and survived for more than 365 days after discharge, while 19 died during admission. We used machine learning by random forest to derive protein features that could be used to deconvolve immune cell phenotypes in paediatric airway samples. We applied these phenotype-specific signatures to identify airway-resident immune cell phenotypes that were differentially enriched by survival status and validated the findings using a large retrospective pneumonia cohort. Main Results We identified immune-cell phenotype classification features for 33 immune cell types. Eosinophil-associated features were significantly elevated in airway samples obtained from pneumonia survivors and were downregulated in children who subsequently died. To confirm these results, we analyzed clinical parameters from >10,000 children who had been admitted with pneumonia in the previous 10 years. The results of this retrospective analysis mirrored airway deconvolution data and showed that survivors had significantly elevated eosinophils at admission compared to fatal pneumonia. Conclusions Using a proteomics bioinformatics approach, we identify airway eosinophils as a critical factor for pneumonia survival in infants and young children.
High mortality after discharge from hospital following acute illness has been observed among children with Severe Acute Malnutrition (SAM). However, mechanisms that may be amenable to intervention to reduce risk are unknown. We performed a nested case-control study among HIV-uninfected children aged 2–59 months treated for complicated SAM according to WHO recommendations at four Kenyan hospitals. Blood was drawn from 1778 children when clinically judged stable before discharge from hospital. Cases were children who died within 60 days. Controls were randomly selected children who survived for one year without readmission to hospital. Untargeted proteomics, total protein, cytokines and chemokines, and leptin were assayed in plasma and corresponding biological processes determined. Among 121 cases and 120 controls, increased levels of calprotectin, von Willebrand factor, angiotensinogen, IL8, IL15, IP10, TNFα, and decreased levels of leptin, heparin cofactor 2, and serum paraoxonase were associated with mortality after adjusting for possible confounders. Acute phase responses, cellular responses to lipopolysaccharide, neutrophil responses to bacteria, and endothelial responses were enriched among cases. Among apparently clinically stable children with SAM, a sepsis-like profile is associated with subsequent death. This may be due to ongoing bacterial infection, translocated bacterial products or deranged immune response during nutritional recovery.
Background: High-throughput whole genome sequencing facilitates investigation of minority virus sub-populations from virus positive samples. Minority variants are useful in understanding within and between host diversity, population dynamics and can potentially assist in elucidating person-person transmission pathways. Several minority variant callers have been developed to describe low frequency sub-populations from whole genome sequence data. These callers differ based on bioinformatics and statistical methods used to discriminate sequencing errors from low-frequency variants. Methods: We evaluated the diagnostic performance and concordance between published minority variant callers used in identifying minority variants from whole-genome sequence data from virus samples. We used the ART-Illumina read simulation tool to generate three artificial short-read datasets of varying coverage and error profiles from an RSV reference genome. The datasets were spiked with nucleotide variants at predetermined positions and frequencies. Variants were called using FreeBayes, LoFreq, Vardict, and VarScan2. The variant callers’ agreement in identifying known variants was quantified using two measures; concordance accuracy and the inter-caller concordance. Results: The variant callers reported differences in identifying minority variants from the datasets. Concordance accuracy and inter-caller concordance were positively correlated with sample coverage. FreeBayes identified the majority of variants although it was characterised by variable sensitivity and precision in addition to a high false positive rate relative to the other minority variant callers and which varied with sample coverage. LoFreq was the most conservative caller. Conclusions: We conducted a performance and concordance evaluation of four minority variant calling tools used to identify and quantify low frequency variants. Inconsistency in the quality of sequenced samples impacts on sensitivity and accuracy of minority variant callers. Our study suggests that combining at least three tools when identifying minority variants is useful in filtering errors when calling low frequency variants.
Passive transfer studies in humans clearly demonstrated the protective role of IgG antibodies against malaria. Identifying the precise parasite antigens that mediate immunity is essential for vaccine design, but has proved difficult. Completion of the Plasmodium falciparum genome revealed thousands of potential vaccine candidates, but a significant bottleneck remains in their validation and prioritization for further evaluation in clinical trials. Focusing initially on the Plasmodium falciparum merozoite proteome, we used peer-reviewed publications, multiple proteomic and bioinformatic approaches, to select and prioritize potential immune targets. We expressed 109 P. falciparum recombinant proteins, the majority of which were obtained using a mammalian expression system that has been shown to produce biologically functional extracellular proteins, and used them to create KILchip v1.0: a novel protein microarray to facilitate high-throughput multiplexed antibody detection from individual samples.The microarray assay was highly specific; antibodies against P. falciparum proteins were detected exclusively in sera from malaria-exposed but not malaria-naïve individuals. The intensity of antibody reactivity varied as expected from strong to weak across well-studied antigens such as AMA1 and RH5 (Kruskal–Wallis H test for trend: p < 0.0001). The inter-assay and intra-assay variability was minimal, with reproducible results obtained in re-assays using the same chip over a duration of 3 months. Antibodies quantified using the multiplexed format in KILchip v1.0 were highly correlated with those measured in the gold-standard monoplex ELISA [median (range) Spearman's R of 0.84 (0.65–0.95)]. KILchip v1.0 is a robust, scalable and adaptable protein microarray that has broad applicability to studies of naturally acquired immunity against malaria by providing a standardized tool for the detection of antibody correlates of protection. It will facilitate rapid high-throughput validation and prioritization of potential Plasmodium falciparum merozoite-stage antigens paving the way for urgently needed clinical trials for the next generation of malaria vaccines.
Reconstructing transmission pathways and defining the underlying determinants of virus diversity is critical for developing effective control measures. Whole genome consensus sequences represent the dominant virus subtype which does not provide sufficient information to resolve transmission events for rapidly spreading viruses with overlapping generations. We explored whether the within-host diversity of respiratory syncytial virus quantified from deep sequence data provides additional resolution to inform on who acquires infection from whom based on shared minor variants in samples that comprised epidemiological clusters and that shared similar genetic background. We report that RSV-A infections are characterized by low frequency diversity that occurs across the genome. Shared minor variant patterns alone, were insufficient to elucidate transmission chains within household members. However, they provided inference on potential transmission links where phylogenetic methods were uninformative of transmission when consensus sequences were identical. Interpretation of minor variant patterns was tractable only for small household outbreaks.
A group of investigators gathered to review the landscape of predictive mathematical modelling of RSV intervention programmes, and to identify gaps in knowledge and strategy options being explored. The objective was to set an agenda for future modelling and related research, to explore possible areas for collaborations, and provide an informed status update for various stakeholders.
Conventionally, workflows examining transcription regulation networks from gene expression data involve distinct analytical steps. There is a need for pipelines that unify data mining and inference deduction into a singular framework to enhance interpretation and hypotheses generation. We propose a workflow that merges network construction with gene expression data mining focusing on regulation processes in the context of transcription factor driven gene regulation. The pipeline implements pathway-based modularization of expression profiles into functional units to improve biological interpretation. The integrated workflow was implemented as a web application software (TransReguloNet) with functions that enable pathway visualization and comparison of transcription factor activity between sample conditions defined in the experimental design. The pipeline merges differential expression, network construction, pathway-based abstraction, clustering and visualization. The framework was applied in analysis of actual expression datasets related to lung, breast and prostrate cancer.
Developing database systems connecting diverse species based on omics is the most important theme in big data biology. To attain this purpose, we have developed KNApSAcK Family Databases, which are utilized in a number of researches in metabolomics. In the present study, we have developed a network-based approach to analyze relationships between 3D structure and biological activity of metabolites consisting of four steps as follows: construction of a network of metabolites based on structural similarity (Step 1), classification of metabolites into structure groups (Step 2), assessment of statistically significant relations between structure groups and biological activities (Step 3), and 2-dimensional clustering of the constructed data matrix based on statistically significant relations between structure groups and biological activities (Step 4). Applying this method to a data set consisting of 2072 secondary metabolites and 140 biological activities reported in KNApSAcK Metabolite Activity DB, we obtained 983 statistically significant structure group-biological activity pairs. As a whole, we systematically analyzed the relationship between 3D-chemical structures of metabolites and biological activities.