BackgroundThere has been an unprecedented effort to sequence the SARS-CoV-2 virus and examine its molecular evolution. This has been facilitated by the availability of publicly accessible databases, such as the GISAID (Global Initiative on Sharing All Influenza Data) and GenBank, which collectively hold millions of SARS-CoV-2 sequence records. Genomic epidemiology, however, seeks to go beyond phylogenetic (the study of evolutionary relationships among biological entities) analysis by linking genetic information to patient characteristics and disease outcomes, enabling a comprehensive understanding of transmission dynamics and disease impact. While these repositories include fields reflecting patient-related metadata for a given sequence, the inclusion of these demographic and clinical details is scarce. The current understanding of patient-related metadata in published sequencing studies and its quality remains unexplored. ObjectiveOur review aims to quantitatively assess the extent and quality of patient-reported metadata in papers reporting original whole genome sequencing of the SARS-CoV-2 virus and analyze publication patterns using bibliometric analysis. Finally, we will evaluate the efficacy and reliability of a machine learning classifier in accurately identifying relevant papers for inclusion in the scoping review. MethodsThe National Institutes of Health’s LitCovid collection will be used for the automated classification of papers reporting having deposited SARS-CoV-2 sequences in public repositories, while an independent search will be conducted in MEDLINE and PubMed Central for validation. Data extraction will be conducted using Covidence (Veritas Health Innovation Ltd). The extracted data will be synthesized and summarized to quantify the availability of patient metadata in the published literature of SARS-CoV-2 sequencing studies. For the bibliometric analysis, relevant data points, such as author affiliations, citation metrics, author keywords, and Medical Subject Headings terms will be extracted. ResultsThis study is expected to be completed in early 2025. Our classification model has been developed and we have classified publications in LitCovid published through February 2023. As of September 2024, papers through August 2024 are being prepared for processing. Screening is underway for validated papers from the classifier. Direct literature searches and screening of the results began in October 2024. We will summarize and narratively describe our findings using tables, graphs, and charts where applicable. ConclusionsThis scoping review will report findings on the extent and types of patient-related metadata reported in genomic viral sequencing studies of SARS-CoV-2, identify gaps in the reporting of patient metadata, and make recommendations for improving the quality and consistency of reporting in this area. The bibliometric analysis will uncover trends and patterns in the reporting of patient-related metadata, including differences in reporting based on study types or geographic regions. The insights gained from this study may help improve the quality and consistency of reporting patient metadata, enhancing the utility of sequence metadata and facilitating future research on infectious diseases. Trial RegistrationOSF Registries osf.io/wrh95; https://doi.org/10.17605/OSF.IO/WRH95 International Registered Report Identifier (IRRID)DERR1-10.2196/58567
Objective:Patient metadata exist in published articles, but are often disconnected from genome sequences in databases, limiting their utility for genomic epidemiology. The objective of this study was to develop and evaluate natural language processing methods to facilitate the large-scale detection of patient metadata associated with reports of genome sequencing in published articles, drawing on the case of SARS-CoV-2. Methods:We applied filters to select a sample of 245 PubMed articles (50,918 sentences) in LitCovid for manual annotation of sentences that reported generating SARS-CoV-2 sequences. We trained, deployed, and validated a BERT-based classifier, and selected a sample of 150 predicted articles (22,147 sentences) for manual annotation of sentences that reported patient metadata associated with the sequences. In addition to training BERT-based classifiers, we experimented with a generative AI approach, prompting the Llama-3-70B LLM using zero-shot, role-based, few-shot, chain-of-thought, and reasoning-eliciting prompting. Results:BERT-based models that were pre-trained on corpora in biomedical or, more specifically, COVID-19 domains outperformed those that were pre-trained on corpora in general domains for detecting reports of patient metadata associated with SARS-CoV-2 sequences, achieving the best performance with a classifier based on a BiomedBERT-Large-Abstract model (F1-score = 0.776). While the best performance of our generative AI approach was achieved using role-based, few-shot, and chain-of-thought prompting (F1-score = 0.558), it was nonetheless outperformed by all of our machine learning-based classifiers. Conclusion:Our methods were applied to more than 350,000 published articles and can be used to advance the utility and efficiency of genomic epidemiology for public health responses to virus outbreaks.
By coupling long-range polymerase chain reaction, wastewater-based epidemiology, and pathogen sequencing, we show that adenovirus type 41 hexon-sequence lineages, described in children with hepatitis of unknown origin in the United States in 2021, were already circulating within the country in 2019. We also observed other lineages in the wastewater, whose complete genomes have yet to be documented from clinical samples.
The SARS-CoV-2 pandemic resulted in a scale-up of viral genomic surveillance globally. However, the wet lab constraints (economic, infrastructural, and personnel) of translating novel virus variant sequence information to meaningful immunological and structural insights that are valuable for the development of broadly acting countermeasures (especially for emerging and re-emerging viruses) remain a challenge in many resource-limited settings. Here, we describe a workflow that couples wastewater surveillance, high-throughput sequencing, phylogenetics, immuno-informatics, and virus capsid structure modeling for the genotype-to-serotype characterization of uncultivated picornavirus sequences identified in wastewater. Specifically, we analyzed canine picornaviruses (CanPVs), which are uncultivated and yet-to-be-assigned members of the family Picornaviridae that cause systemic infections in canines. We analyzed 118 archived (stored at −20 °C) wastewater (WW) samples representing a population of ~700,000 persons in southwest USA between October 2019 to March 2020 and October 2020 to March 2021. Samples were pooled into 12 two-liter volumes by month, partitioned (into filter-trapped solids [FTSs] and filtrates) using 450 nm membrane filters, and subsequently concentrated to 2 mL (1000×) using 10,000 Da MW cutoff centrifugal filters. The 24 concentrates were subjected to RNA extraction, CanPV complete capsid single-contig RT-PCR, Illumina sequencing, phylogenetics, immuno-informatics, and structure prediction. We detected CanPVs in 58.3% (14/24) of the samples generated 13,824,046 trimmed Illumina reads and 27 CanPV contigs. Phylogenetic and pairwise identity analyses showed eight CanPV genotypes (intragenotype divergence <14%) belonging to four clusters, with intracluster divergence of <20%. Similarity analysis, immuno-informatics, and virus protomer and capsid structure prediction suggested that the four clusters were likely distinct serological types, with predicted cluster-distinguishing B-cell epitopes clustered in the northern and southern rims of the canyon surrounding the 5-fold axis of symmetry. Our approach allows forgenotype-to-serotype characterization of uncultivated picornavirus sequences by coupling phylogenetics, immuno-informatics, and virus capsid structure prediction. This consequently bypasses a major wet lab-associated bottleneck, thereby allowing resource-limited settings to leapfrog from wastewater-sourced genomic data to valuable immunological insights necessary for the development of prophylaxis and other mitigation measures.
We determine the presence and diversity of rhinoviruses in nasopharyngeal swab samples from 248 individuals who presented with influenza-like illness (ILI) at a university clinic in the Southwest United States between October 1, 2020 and March 31, 2021. We identify at least 13 rhinovirus genotypes (A11, A22, A23, A25, A67, A101, B6, B79, C1, C17, C36, and C56, as well a new genotype [AZ88**]) and 16 variants that contributed to the burden of ILI in the community. We also describe the complete capsid protein gene of a member (AZ88**) of an unassigned rhinovirus A genotype.
There are many studies that require researchers to extract specific information from the published literature, such as details about sequence records or about a randomized control trial. While manual extraction is cost efficient for small studies, larger studies such as systematic reviews are much more costly and time-consuming. To avoid exhaustive manual searches and extraction, and their related cost and effort, natural language processing (NLP) methods can be tailored for the more subtle extraction and decision tasks that typically only humans have performed. The need for such studies that use the published literature as a data source became even more evident as the COVID-19 pandemic raged through the world and millions of sequenced samples were deposited in public repositories such as GISAID and GenBank, promising large genomic epidemiology studies, but more often than not lacked many important details that prevented large-scale studies. Thus, granular geographic location or the most basic patient-relevant data such as demographic information, or clinical outcomes were not noted in the sequence record. However, some of these data was indeed published, but in the text, tables, or supplementary material of a corresponding published article. We present here methods to identify relevant journal articles that report having produced and made available in GenBank or GISAID, new SARS-CoV-2 sequences, as those that initially produced and made available the sequences are the most likely articles to include the high-level details about the patients from whom the sequences were obtained. Human annotators validated the approach, creating a gold standard set for training and validation of a machine learning classifier. Identifying these articles is a crucial step to enable future automated informatics pipelines that will apply Machine Learning and Natural Language Processing to identify patient characteristics such as co-morbidities, outcomes, age, gender, and race, enriching SARS-CoV-2 sequence databases with actionable information for defining large genomic epidemiology studies. Thus, enriched patient metadata can enable secondary data analysis, at scale, to uncover associations between the viral genome (including variants of concern and their sublineages), transmission risk, and health outcomes. However, for such enrichment to happen, the right papers need to be found and very detailed data needs to be extracted from them. Further, finding the very specific articles needed for inclusion is a task that also facilitates scoping and systematic reviews, greatly reducing the time needed for full-text analysis and extraction.
Virus surveillance by wastewater-based epidemiology (WBE) in two Arizona municipalities in Maricopa County, USA (-700,000 people), revealed the presence of six canine picornavirus (CanPV) variants: five in 2019 and one in 2021. Phylogenetic analysis suggests these viruses might be from domestic dog breeds living within or around the area. Phylogenetic and pairwise identity analyses suggest over 15 years of likely enzootic circulation of multiple lineages of CanPV in the USA and possibly globally. Considering <10 CanPV sequences are publicly available in GenBank as of June 2, 2022, the results provided here constitute an increase of current knowledge on CanPV diversity and highlight the need for increased surveillance.
We describe the genome of Microvirus-AZ-2020, which was identified from wastewater in Arizona, USA, in October 2020. Microvirus-AZ-2020 belongs to subfamily Gokushovirinae and contains six (five known and one hypothetical) open reading frames (ORFs), each with >40 codons. HHPred analysis and Colabfold structure prediction suggest that the hypothetical ORF encodes a previously undescribed putative DNA-binding protein.
The use of wastewater-based epidemiology (WBE) for early detection of virus circulation and response during the SARS-CoV-2 pandemic increased interest in and use of virus concentration protocols that are quick, scalable, and efficient. One such protocol involves sample clarification by size fractionation using either low-speed centrifugation to produce a clarified supernatant or membrane filtration to produce an initial filtrate depleted of solids, eukaryotes and bacterial present in wastewater (WW), followed by concentration of virus particles by ultrafiltration of the above. While this approach has been successful in identifying viruses from WW, it assumes that majority of the viruses of interest should be present in the fraction obtained by ultrafiltration of the initial filtrate, with negligible loss of viral particles and viral diversity. We used WW samples collected in a population of ~700,000 in southwest USA between October 2019 and March 2021, targeting three non-enveloped viruses (enteroviruses [EV], canine picornaviruses [CanPV], and human adenovirus 41 [Ad41]), to evaluate whether size fractionation of WW prior to ultrafiltration leads to appreciable differences in the virus presence and diversity determined. We showed that virus presence or absence in WW samples in both portions (filter trapped solids [FTS] and filtrate) are not consistent with each other. We also found that in cases where virus was detected in both fractions, virus diversity (or types) captured either in FTS or filtrate were not consistent with each other. Hence, preferring one fraction of WW over the other can undermine the capacity of WBE to function as an early warning system and negatively impact the accurate representation of virus presence and diversity in a population.
The RNA-binding protein HuD (a.k.a., ELAVL4) is involved in neuronal development and synaptic plasticity mechanisms, including addiction-related processes such as cocaine conditioned-place preference (CPP) and food reward. The most studied function of this protein is mRNA stabilization; however, we have recently shown that HuD also regulates the levels of circular RNAs (circRNAs) in neurons. To examine the role of HuD in the control of coding and non-coding RNA networks associated with substance use, we identified sets of differentially expressed mRNAs, circRNAs and miRNAs in the striatum of HuD knockout (KO) mice. Our findings indicate that significantly downregulated mRNAs are enriched in biological pathways related to cell morphology and behavior. Furthermore, deletion of HuD altered the levels of 15 miRNAs associated with drug seeking. Using these sets of data, we predicted that a large number of upregulated miRNAs form competing endogenous RNA (ceRNA) networks with circRNAs and mRNAs associated with the neuronal development and synaptic plasticity proteins LSAMP and MARK3. Additionally, several downregulated miRNAs form ceRNA networks with mRNAs and circRNAs from MEF2D, PIK3R3, PTRPM and other neuronal proteins. Together, our results indicate that HuD regulates ceRNA networks controlling the levels of mRNAs associated with neuronal differentiation and synaptic physiology.
With the continued adoption of single cell RNA sequencing, evaluation of parameters and approaches are needed to gauge performance of platforms across sample types and disease conditions. In this study, we evaluated single nuclei RNA sequencing (snRNAseq) and single whole cell RNAseq in quadruplicate on the 10x Chromium platform using fresh frozen frontal cortex from a healthy elderly control and Alzheimer's disease (AD) subject. A median of 1,584 cells or nuclei were sequenced across samples with a median of 31,283 mean reads per cell or nuclei. Corroborating other studies, we observed elevated mitochondrial transcripts, lower levels of intronic sequences, a lower number of genes detected, and increased background, in whole cell libraries compared to nuclei libraries. Upon normalizing the number of reads used for analysis, it was revealed that 5’ priming of mRNA transcripts enabled identification of a larger number of unique genes per cell as well as a larger number of total genes compared to 3’ priming. Utilization of replicates for each sample type improved cell population classification and identified populations include neurons (granule, pyramidal, GABAergic), astrocytes, microglia, endothelial cells, and oligodendrocytes. We further identified two cell populations, oligodendrocyte precursor cells and neuronal stem cells, uniquely in the AD subject. Our evaluation reveals that 5’ snRNAseq demonstrates superior performance on the 10x platform when analyzing fresh frozen brain from elderly subjects. While analysis of larger numbers of samples is needed, we also present preliminary data demonstrating discovery of cell populations in AD.