In July 2021, the California Code of Regulations Title 17 required all laboratories performing SARS‑CoV‑2 whole genome sequencing (WGS) to report their sequencing results to the California Department of Public Health (CDPH). These viral genomic data and patient metadata were compiled into the Integrated Genomic Epidemiology Database (IGED). Linking anonymized viral sequences with patient‑level information enabled monitoring of infectiousness, pathogenicity, transmission dynamics, evolution, and vaccine evasion among emerging SARS‑CoV‑2 lineages. Laboratories performing SARS-CoV-2 WGS transmitted sequencing results to CDPH through Electronic Laboratory Reporting (ELR) and non-ELR pathways. CDPH applied uniform reporting requirements but allowed flexibility in specific data formats to accommodate diverse data systems. To preserve data quality and interoperability across heterogeneous sources, CDPH implemented standardization, validation, and deduplication protocols. Snowflake, a cloud‑based data storage and analytics platform, and Posit Connect, a cloud deployment and automation platform, supported the management, processing, and integration of data within the IGED. The IGED established links between SARS‑CoV‑2 WGS data and epidemiologic metadata for 801,418 sequences, representing 81.7% of all sequences reported in California. Lineages reported to the IGED showed strong concordance with lineage proportions in GISAID. Sequences reported to the IGED had average turnaround times longer than one month, and the majority of sequencing was performed in Southern California and Los Angeles. The IGED enhanced genomic surveillance through predictive modeling and monitoring concerning evolutionary trends such as recombination and saltations in persistent infections. Development of the IGED highlighted the need for standardized data requirements, sustained funding for sequencing, incentives for data submission, and interdisciplinary collaboration to build an effective genomic surveillance system. This framework for linking genomic and epidemiologic data has not only generated critical insights for SARS‑CoV‑2 but also provided the foundation for CDPH and other public health organizations to develop similar IGED‑like systems for other priority pathogens as genomic surveillance expands.
Supplementary Figure from Human Papillomavirus Integration Strictly Correlates with Global Genome Instability in Head and Neck Cancer
Introduction:The clinical incidence of antimicrobial-resistant fungal infections has dramatically increased in recent years. Certain fungal pathogens colonize various body cavities, leading to life-threatening bloodstream infections. However, the identification and characterization of fungal isolates in laboratories remain a significant diagnostic challenge in medicine and public health. Whole-genome sequencing provides an unbiased and uniform identification pipeline for fungal pathogens but most bioinformatic analysis pipelines focus on prokaryotic species. To this end, TheiaEuk_Illumina_PE_PHB (TheiaEuk) was designed to focus on genomic analysis specialized to fungal pathogens.Methods:TheiaEuk was designed using containerized components and written in the workflow description language (WDL) to facilitate deployment on the cloud-based open bioinformatics platform Terra. This species-agnostic workflow enables the analysis of fungal genomes without requiring coding, thereby reducing the entry barrier for laboratory scientists. To demonstrate the usefulness of this pipeline, an ongoing outbreak of C. auris in southern Nevada was investigated. We performed whole-genome sequence analysis of 752 new C. auris isolates from this outbreak. Furthermore, TheiaEuk was utilized to observe the accumulation of mutations in the FKS1 gene over the course of the outbreak, highlighting the utility of TheiaEuk as a monitor of emerging public health threats when combined with whole-genome sequencing surveillance of fungal pathogens.Results:A primary result of this work is a curated fungal database containing 5,667 unique genomes representing 245 species. TheiaEuk also incorporates taxon-specific submodules for specific species, including clade-typing for Candida auris (C. auris). In addition, for several fungal species, it performs dynamic reference genome selection and variant calling, reporting mutations found in genes currently associated with antifungal resistance (FKS1, ERG11, FUR1). Using genome assemblies from the ATCC Mycology collection, the taxonomic identification module used by TheiaEuk correctly assigned genomes to the species level in 126/135 (93.3%) instances and to the genus level in 131/135 (97%) of instances, and provided zero false calls. Application of TheiaEuk to actual specimens obtained in the course of work at a local public health laboratory resulted in 13/15 (86.7%) correct calls at the species level, with 2/15 called at the genus level. It made zero incorrect calls. TheiaEuk accurately assessed clade type of Candida auris in 297/302 (98.3%) of instances.Discussion:TheiaEuk demonstrated effectiveness in identifying fungal species from whole genome sequence. It further showed accuracy in both clade-typing of C. auris and in the identification of mutations known to associate with drug resistance in that organism.
The capacity for pathogen genomics in public health expanded rapidly during the coronavirus disease 2019 (COVID- 19) pan-demic, but many public health laboratories did not have the infrastructure in place to handle the vast amount of severe acute respiratory syndrome coronavirus 2 (SARS-CoV- 2) sequence data generated. The California Department of Public Health, in partnership with Theiagen Genomics, was an early adopter of cloud -based resources for bioinformatics and genomic epi-demiology, resulting in the creation of a SARS- CoV- 2 genomic surveillance system that combined the efforts of more than 40 sequencing laboratories across government, academia and industry to form California COVIDNet, California's SARS- CoV- 2 Whole-Genome Sequencing Initiative. Open-source bioinformatics workflows, ongoing training sessions for the public health workforce, and automated data transfer to visualization tools all contributed to the success of California COVIDNet. While chal-lenges remain for public health genomic surveillance worldwide, California COVIDNet serves as a framework for a scaled and successful bioinformatics infrastructure that has expanded beyond SARS- CoV- 2 to other pathogens of public health importance,
We have adopted an open bioinformatics ecosystem to address the challenges of bioinformatics implementation in public health laboratories (PHLs). Bioinformatics implementation for public health requires practitioners to undertake standardized bioinformatic analyses and generate reproducible, validated and auditable results. It is essential that data storage and analysis are scalable, portable and secure, and that implementation of bioinformatics fits within the operational constraints of the laboratory. We address these requirements using Terra, a web-based data analysis platform with a graphical user interface connecting users to bioinformatics analyses without the use of code. We have developed bioinformatics workflows for use with Terra that specifically meet the needs of public health practitioners. These Theiagen workflows perform genome assembly, quality control, and characterization, as well as construction of phylogeny for insights into genomic epidemiology. Additonally, these workflows use open-source containerized software and the WDL workflow language to ensure standardization and interoperability with other bioinformatics solutions, whilst being adaptable by the user. They are all open source and publicly available in Dockstore with the version-controlled code available in public GitHub repositories. They have been written to generate outputs in standardized file formats to allow for further downstream analysis and visualization with separate genomic epidemiology software. Testament to this solution meeting the requirements for bioinformatic implementation in public health, Theiagen workflows have collectively been used for over 5 million sample analyses in the last 2 years by over 90 public health laboratories in at least 40 different countries. Continued adoption of technological innovations and development of further workflows will ensure that this ecosystem continues to benefit PHLs.
Human papillomavirus (HPV)-positive head and neck cancers, predominantly oropharyngeal squamous cell carcinoma (OPSCC), exhibit epidemiologic, clinical, and molecular characteristics distinct from those OPSCCs lacking HPV. We applied a combination of whole-genome sequencing and optical genome mapping to interrogate the genome structure of HPV-positive OPSCCs. We found that the virus had integrated in the host genome in two thirds of the tumors examined but resided solely extrachromosomally in the other third. Integration of the virus occurred at essentially random sites within the genome. Focal amplification of the virus and the genomic sequences surrounding it often occurred subsequent to integration, with the number of tandem repeats in the chromosome accounting for the increased copy number of the genome sequences flanking the site of integration. In all cases, viral integration correlated with pervasive genome-wide somatic alterations at sites distinct from that of viral integration and comprised multiple insertions, deletions, translocations, inversions, and point mutations. Few or no somatic mutations were present in tumors with only episomal HPV. Our data could be interpreted by positing that episomal HPV is captured in the host genome following an episode of global genome instability during tumor development. Viral integration correlated with higher grade tumors, which may be explained by the associated extensive mutation of the genome and suggests that HPV integration status may inform prognosis.IMPLICATIONS:Our results indicate that HPV integration in head and neck cancer correlates with extensive pangenomic structural variation, which may have prognostic implications.
Accurately predicting chromatin loops from genome-wide interaction matrices such as Hi-C data is critical to deepening our understanding of proper gene regulation. Current approaches are mainly focused on searching for statistically enriched dots on a genome-wide map. However, given the availability of orthogonal data types such as ChIA-PET, HiChIP, Capture Hi-C, and high-throughput imaging, a supervised learning approach could facilitate the discovery of a comprehensive set of chromatin interactions. Here, we present Peakachu, a Random Forest classification framework that predicts chromatin loops from genome-wide contact maps. We compare Peakachu with current enrichment-based approaches, and find that Peakachu identifies a unique set of short-range interactions. We show that our models perform well in different platforms, across different sequencing depths, and across different species. We apply this framework to predict chromatin loops in 56 Hi-C datasets, and release the results at the 3D Genome Browser.
3-Hydroxy-3-methylglutaryl coenzyme A reductase is associated with monitoring cholesterol levels. The presence of the single-nucleotide polymorphism rs3846662 introduces alternative splicing at exon 13; the exclusion of this exon leads to a reduction in total cholesterol levels. Lower cholesterol levels are linked to a reduction in Alzheimer's disease (AD) risk. The major allele of rs3846662, which encourages the splicing of exon 13, has recently been shown to act as a preventative allele for AD, especially in women. The purpose of our research was to replicate and confirm this finding. Using logistic regressions and survival curves, we found a significant association between AD and rs3846662, with a stronger association in individuals who carry the APOE e4 allele, supporting previously published work. The effect of rs3846662 on women is insignificant in our cohort. We confirmed that rs3846662 is associated with reduced risk for AD without gender differences; however, we failed to detect association between rs3846662 and delayed mild cognitive impairment conversion to AD for either of the APOE e4 allelic groups.
3-hydroxyl-3-methylglutaryl coenzyme a reductase (HMGCR) is associated with monitoring cholesterol levels. The presence of the single-nucleotide polymorphism rs3846662 introduces alternative splicing at exon 13; the exclusion of this exon leads to a reduction in total cholesterol levels. Lower cholesterol levels have been linked to a reduction in Alzheimer's disease (AD) risk. The major allele of rs3846662, which encourages the splicing of exon 13, has recently been shown to act as a preventative allele for AD, especially in women. The purpose of our research was to replicate and confirm this finding in the Cache County Study on Memory in Aging to test for association with AD status and rate of cognitive decline We will use generalized linear model and survival curves to determine an association between AD and rs3846662. ANCOVA and multiple mix model regression will be used to determine the association of rs3846662 with covariant like sex, APOE4 allele, rate of cognitive decline Using generalized linear model regressions and survival curves, we found a significant association between AD and rs3846662 (p-value = 0.049), with a stronger association in individuals that also express the APOE4 allele (p-value = 0.016). The effect of this allele on women was insignificant (p-value = 0.129) in our findings. However, we failed to detect association between rs3846662 and delayed mild cognitive impairment (MCI) conversion to AD for either of the APOE4 allelic groups (APOE4 positive p-value = 0.663; APOE4 negative p-value = 0.671) We were able to confirm that rs3846662 is associated with reduced risk for AD without sex differences using the Cache County Study on Memory in Aging.
It is well-documented that codon usage biases affect gene translational efficiency; however, it is less known if viruses share their host's codon usage motifs.We determined that human-infecting viruses share similar codon usage biases as proteins that are expressed in tissues the viruses infect.By performing 7,052,621 pairwise comparisons of genes from humans versus genes from 113 viruses that infect humans, we determined which codon usage motifs were most highly correlated.We found that 16 viruses averaged a significant correlation in codon usage with over 500 human genes per viral gene, 58 viruses were highly correlated with an average of at least 100 human genes per viral gene, and 37 viruses were significantly correlated with an average of at least one human gene per viral gene at an alpha level of 7.09 x (0.05 alpha / 7,052,621 comparisons).Only two viruses were not highly correlated with an average of one human gene per viral gene.While relatively few of the interactions were previously documented, the high statistical correlations suggest that researchers may be able to determine which tissues a virus is most likely to infect by analyzing codon usage biases.
ABSTRACT Isolates of the lactic acid bacterium Leuconostoc citreum are a major part of fermentation processes, especially in Korean kimchi. Here, we present the genome of L. citreum DmW_111, isolated from wild Drosophila melanogaster ; analysis of this genome will expand the diversity of genome sequences for non- Lactobacillus spp. isolated from D. melanogaster .