Introduction ClinVar is an NIH database of human genetic variants and classifications of those variants for inherited disease, cancer, and drug responses. Since its introduction in 2013, the vast majority of ClinVar data has been germline variants. With the increase in genetic testing for cancer and publication of standards for classification of somatic variants, ClinVar made significant updates to provide distinct classifications of variants in the germline and somatic contexts. These updates were made available in January 2024; this abstract reviews uptake of submissions of somatic variants in ClinVar. Methods ClinVar submission processing was updated to accept somatic variants classified for clinical impact and for oncogenicity. This feature was advertised in the clinical genetics community through ClinVar’s GitHub repository (https://github.com/ncbi/clinvar); NCBI’s social media accounts; presentations to ClinGen, the Variant Interpretation in Cancer Consortium (VICC); and the Cancer Genome Consortium (CGC). Statistics for submissions of somatic variants were collected by querying the ClinVar website, using Advanced search. Results As of March 1, 2025, 11 organizations have submitted somatic variants to ClinVar, or fewer than 1% of total submitting organizations (3004). The majority of somatic variants (83%) were submitted based on results from clinical testing. 15% of somatic variants were submitted based on results from research, and 0.5% (4 variants) were submitted based on curation results from the submitter. The majority of variants (96%) were submitted without a comment on the classification that describes the evidence. A small percentage of variants (8.5%) were submitted with citations from PubMed as part of the evidence for the classification. Discussion and Conclusion Somatic-specific classifications have been available in ClinVar for just over a year. However, uptake of submissions has not been robust; to date, only 11 organizations have submitted somatic variants. In contrast, ClinVar’s first release in 2013 included data from 29 organizations, at a time when ClinVar was barely known. Most of the somatic variant classifications in ClinVar are based on clinical testing. Additionally, most of the somatic variants were not submitted with citations from PubMed as part of the evidence. These two observations highlight one of the strengths of ClinVar which is to broadly share results from clinical testing which are not typically published in scientific journals. However, the lack of citations in PubMed coupled with the lack of comments to explain the classification suggests that the submitted variants may be of limited utility to ClinVar users, or that users will need to contact submitters to understand each variant classification. In conclusion, more effort is needed to onboard cancer testing laboratories for ClinVar submission, and to encourage submitters to include PubMed citations and a comment on the classification so that the data is of optimal use to the clinical genetics testing community.
ClinVar (www.ncbi.nlm.nih.gov/clinvar/) is a free, public database of human genetic variants and their relationships to disease, with >3 million variants submitted by >2800 organizations across the world. The database was recently updated to have three types of classifications: germline, oncogenicity and clinical impact for somatic variants. As for germline variants, classifications for somatic variants can be submitted in batches in a file submission or through the submission API; variants can also be submitted and updated one at a time in online submission forms. The ClinVar XML files were redesigned to allow multiple classification types. Both old and new formats of the XML are supported through the end of 2024. Data for somatic classifications were also added to the ClinVar VCF files and to several tab-delimited files. The ClinVar VCV pages were updated to display the three types of classifications, both as it was submitted and as it was aggregated by ClinVar. Clinical testing laboratories and others in the cancer community are invited to share their classifications of somatic variant classifications through ClinVar to provide transparency in genomic testing and improve patient care.
Colorectal cancer (CRC) is one of the top five most common and life-threatening malignancies worldwide. Most CRC develops from advanced colorectal adenoma (ACA), a precancerous stage, through the adenoma-carcinoma sequence. However, its underlying mechanisms, including how the tumor microenvironment changes, remain elusive. Therefore, we conducted an integrative analysis comparing RNA-seq data collected from 40 ACA patients who visited Dongguk University Ilsan Hospital with normal adjacent colons and tumor samples from 18 CRC patients collected from a public database. Differential expression analysis identified 21 and 79 sequentially up- or down-regulated genes across the continuum, respectively. The functional centrality of the continuum genes was assessed through network analysis, identifying 11 up- and 13 down-regulated hub-genes. Subsequently, we validated the prognostic effects of hub-genes using the Kaplan-Meier survival analysis. To estimate the immunological transition of the adenoma-carcinoma sequence, single-cell deconvolution and immune repertoire analyses were conducted. Significant composition changes for innate immunity cells and decreased plasma B-cells with immunoglobulin diversity were observed, along with distinctive immunoglobulin recombination patterns. Taken together, we believe our findings suggest underlying transcriptional and immunological changes during the adenoma-carcinoma sequence, contributing to the further development of pre-diagnostic markers for CRC.
Background Male-pattern baldness (MPB) is the most common cause of hair loss in men. It can be categorized into three types: type 2 (T2), type 3 (T3), and type 4 (T4), with type 1 (T1) being considered normal. Although various MPB-associated genetic variants have been suggested, a comprehensive study for linking these variants to gene expression regulation has not been performed to the best of our knowledge. Results In this study, we prioritized MPB-related tissue panels using tissue-specific enrichment analysis and utilized single-tissue panels from genotype-tissue expression version 8, as well as cross-tissue panels from context-specific genetics. Through a transcriptome-wide association study and colocalization analysis, we identified 52, 75, and 144 MPB associations for T2, T3, and T4, respectively. To assess the causality of MPB genes, we performed a conditional and joint analysis, which revealed 10, 11, and 54 putative causality genes for T2, T3, and T4, respectively. Finally, we conducted drug repositioning and identified potential drug candidates that are connected to MPB-associated genes. Conclusions Overall, through an integrative analysis of gene expression and genotype data, we have identified robust MPB susceptibility genes that may help uncover the underlying molecular mechanisms and the novel drug candidates that may alleviate MPB.
Background Juvenile idiopathic arthritis (JIA) is one of the most prevalent rheumatic disorders in children and is classified as an autoimmune disease (AID). While a robust genetic contribution to JIA etiology has been established, the exact pathogenesis remains unclear. Methods To prioritize biologically interpretable susceptibility genes and proteins for JIA, we conducted transcriptome-wide and proteome-wide association studies (TWAS/PWAS). Then, to understand the genetic architecture of JIA, we systematically analyzed single-nucleotide polymorphism (SNP)-based heritability, a signature of natural selection, and polygenicity. Next, we conducted HLA typing using multi-ethnicity RNA sequencing data. Additionally, we examined the T cell receptor (TCR) repertoire at a single-cell level to explore the potential links between immunity and JIA risk. Results We have identified 19 TWAS genes and two PWAS proteins associated with JIA risks. Furthermore, we observe that the heritability and cell type enrichment analysis of JIA are enriched in T lymphocytes and HLA regions and that JIA shows higher polygenicity compared to other AIDs. In multi-ancestry HLA typing, B*45:01 is more prevalent in African JIA patients than in European JIA patients, whereas DQA1*01:01, DQA1*03:01, and DRB1*04:01 exhibit a higher frequency in European JIA patients. Using single-cell immune repertoire analysis, we identify clonally expanded T cell subpopulations in JIA patients, including CXCL13 + BHLHE40 + T H cells which are significantly associated with JIA risks. Conclusion Our findings shed new light on the pathogenesis of JIA and provide a strong foundation for future mechanistic studies aimed at uncovering the molecular drivers of JIA.
Psoriasis is a chronic inflammatory skin disease characterized by cutaneous eruptions and pruritus. Because the genetic backgrounds of psoriasis are only partially revealed, an integrative and rigorous study is necessary. We conducted a transcriptome-wide association study (TWAS) with the new Genotype-Tissue Expression version 8 reference panels, including some tissue and multi-tissue panels that were not used previously. We performed tissue-specific heritability analyses on genome-wide association study data to prioritize the tissue panels for TWAS analysis. TWAS and colocalization (COLOC) analyses were performed with eight tissues from the single-tissue panels and the multi-tissue panels of context-specific genetics (CONTENT) to increase tissue specificity and statistical power. From TWAS, we identified the significant associations of 101 genes in the single-tissue panels and 64 genes in the multi-tissue panels, of which 26 genes were replicated in the COLOC. Functional annotation and network analyses identified that the genes were associated with psoriasis and/or immune responses. We also suggested drug candidates that interact with jointly significant genes through a conditional and joint analysis. Together, our findings may contribute to revealing the underlying genetic mechanisms and provide new insights into treatments for psoriasis.
Chronic obstructive pulmonary disease (COPD) is a respiratory disease characterized by airflow limitation and chronic inflammation of the lungs that is a leading cause of death worldwide. Since the complete pathological mechanisms at the single-cell level are not fully understood yet, an integrative approach to characterizing the single-cell-resolution landscape of COPD is required. To identify the cell types and mechanisms associated with the development of COPD, we conducted a meta-analysis using three single-cell RNA-sequencing datasets of COPD. Among the 154,011 cells from 16 COPD patients and 18 healthy subjects, 17 distinct cell types were observed. Of the 17 cell types, monocytes, mast cells, and alveolar type 2 cells (AT2 cells) were found to be etiologically implicated in COPD based on genetic and transcriptomic features. The most transcriptomically diversified states of the three etiological cell types showed significant enrichment in immune/inflammatory responses (monocytes and mast cells) and/or mitochondrial dysfunction (monocytes and AT2 cells). We then identified three chemical candidates that may potentially induce COPD by modulating gene expression patterns in the three etiological cell types. Overall, our study suggests the single-cell level mechanisms underlying the pathogenesis of COPD and may provide information on toxic compounds that could be potential risk factors for COPD.
Submissions to ClinVar increase each year. Today, the database holds more than 1.8 million submitted records representing 1.1 million variants. Most data are submitted to ClinVar in batches of many records. Submission processing includes several types of data validation, including verifying that minimal fields for submission are included; validating the variant description; and confirming that allowed values are provided for other fields like the interpretation of the variant. Prior to 2021, all submissions required a curator to shepherd the submission through the validation and processing steps and return errors to the submitter.
Antimicrobial peptides (AMPs) show promises as valuable compounds for developing therapeutic agents to control the worldwide health threat posed by the increasing prevalence of antibiotic-resistant bacteria. Animal venom can be a useful source for screening AMPs due to its various bioactive components. Here, the deep learning model was developed to predict species-specific antimicrobial activity. To overcome the data deficiency, a multi-task learning method was implemented, achieving F1 scores of 0.818, 0.696, 0.814, 0.787, and 0.719 for Bacillus subtilis, Escherichia coli, Pseudomonas aeruginosa, Staphylococcus aureus, and Staphylococcus epidermidis, respectively. Peptides PA-Full and PA-Win were identified from the model using different inputs of full and partial sequences, broadening the application of transcriptome data of the spider Pardosa astrigera. Two peptides exhibited strong antimicrobial activity against all five strains along with cytocompatibility. Our approach enables excavating AMPs with high potency, which can be expanded into the fields of biology to address data insufficiency.
Three chitinolytic, Gram-negative, light pink, capsule-forming, rod-shaped bacterial strains with gliding motion (MYSH2T, MJ1aT and dk17T) were isolated from seashells, soil and foxtail, respectively. Phylogenetic analysis of the 16S rRNA gene sequences and concatenated alignment of 92 core genes indicated that strains MYSH2T, MJ1aT and dk17T were novel species of the genus Mucilaginibacter and exhibited a high 16S rRNA sequence similarity (i.e. more than 97.2 %) among each other. These novel strains contained summed feature 3 (C16:1 ω7c and/or C16:1 ω6), iso-C15:0 and MK-7 as the predominant fatty acids and menaquinone. According to the CAZys coding gene of KAAS, MYSH2T and MJ1aT were interpreted as strains containing both GH18 and 19 family coding genes, except for dk17T, which shows only GH19 family genes. These strains likely degrade chitin to chitobiose or directly to N-acetyl-d-glucosamine, which may enhance their chitinolytic capacity, thus making these stains potentially useful for industrial chitin degradation. Based on distinct morphological, physiological, chemotaxonomic and phylogenetic differences from their closest phylogenetic neighbours, we propose that strains MYSH2T, MJ1aT and dk17T represent three novel species in the genus Mucilaginibacter, for which the names Mucilaginibacter conchicola sp. nov. (=KACC 19716T=JCM 32787T), Mucilaginibacter achroorhodeus sp. nov. (=KACC 19906T=NBRC 113667T) and Mucilaginibacter pallidiroseus sp. nov. (=KACC 19907T=NBRC 113666T) are proposed. An emended description of the genus Mucilaginibacter is proposed.
Uterine fibroid is one of the most prevalent benign tumors in women, with high socioeconomic costs. Although genome-wide association studies (GWAS) have identified several loci associated with uterine fibroid risks, they could not successfully interpret the biological effects of genomic variants at the gene expression levels. To prioritize uterine fibroid susceptibility genes that are biologically interpretable, we conducted a transcriptome-wide association study (TWAS) by integrating GWAS data of uterine fibroid and expression quantitative loci data. We identified nine significant TWAS genes including two novel genes, RP11-282O18.3 and KBTBD7, which may be causal genes for uterine fibroid. We conducted functional enrichment network analyses using the TWAS results to investigate the biological pathways in which the overall TWAS genes were involved. The results demonstrated the immune system process to be a key pathway in uterine fibroid pathogenesis. Finally, we carried out chemical-gene interaction analyses using the TWAS results and the comparative toxicogenomics database to determine the potential risk chemicals for uterine fibroid. We identified five toxic chemicals that were significantly associated with uterine fibroid TWAS genes, suggesting that they may be implicated in the pathogenesis of uterine fibroid. In this study, we performed an integrative analysis covering the broad application of bioinformatics approaches. Our study may provide a deeper understanding of uterine fibroid etiologies and informative notifications about potential risk chemicals for uterine fibroid.
Cerebral adrenoleukodystrophy (cALD) is a rare neurodegenerative disease characterized by inflammatory demyelination in the central nervous system. Another neurodegenerative disease with a high prevalence, Alzheimer's disease (AD), shares many common features with cALD such as cognitive impairment and the alleviation of symptoms by erucic acid. We investigated cALD and AD in parallel to study the shared pathological pathways between a rare disease and a more common disease. The approach may expand the biological understandings and reveal novel therapeutic targets. Gene set enrichment analysis (GSEA) and weighted gene correlation network analysis (WGCNA) were conducted to identify both the resemblance in gene expression patterns and genes that are pathologically relevant in the two diseases. Within differentially expressed genes (DEGs), GSEA identified 266 common genes with similar up- or down-regulation patterns in cALD and AD. Among the interconnected genes in AD data, two gene sets containing 1,486 genes preserved in cALD data were selected by WGCNA that may significantly affect the development and progression of cALD. WGCNA results filtered by functional correlation via protein-protein interaction analysis overlapping with GSEA revealed four genes (annexin A5, beta-2-microglobulin, CD44 molecule, and fibroblast growth factor 2) that showed robust associations with the pathogeneses of cALD and AD, where they were highly involved in inflammation, apoptosis, and the mitogen-activated protein kinase pathway. This study provided an integrated strategy to provide new insights into a rare disease with scant publicly available data (cALD) using a more prevalent disorder with some pathological association (AD), which suggests novel druggable targets and drug candidates.
A Gram-negative, moderately halophilic bacterium, designated as strain Y3S6T, was isolated from a surface seawater sample collected from Dongangyoeng cave, Udo-myeon, Jeju-si, Jeju-do, Repulic of Korea. Cells of strain Y3S6T were aerobic, rod-shaped, non-sporulated, yellow, catalase- negative, oxidase-negative and motile with one polar flagellum. Growth of strain Y3S6T occurred at 15-40 °C (optimum: 25-30 °C), at pH 6.0-9.0 (optimum: pH 7.0) and in the presence of 0-13% NaCl (optimum: 1-6 %, w/v). The novel strain was able to produce carotenoids. Its chemotaxonomic and morphological characteristics were consistent with those of members of the genus Halomonas. Phylogenetic analysis of the 16S rRNA gene sequence revealed that strain Y3S6T formed a clade with Halomonas pellis L5T (98.97 %) and Halomonas saliphila LCB169T(98.90%). The average nucleotide identity and digital DNA-DNA hybridization values of strain Y3S6T with the most closely related strains for which whole genomes are publicly available were 82.3-85.2% and 62.8-66.1 %, respectively. The major fatty acids in strain Y3S6T were C16 : 0, C19 : 0 cyclo ω8c and summed feature 8 (composed of C18 : 1 ω7c and/or C18 : 1 ω6c), and the predominant quinone was Q-9. Its polar lipid profile consisted of diphosphatidylglycerol, phosphatidylglycerol, phosphatidylethanolamine, two unidentified phosphoglycolipid, one unidentified phosphoaminoglycolipid and one unidentified phospholipid. The genomic DNA G+C content based on the draft genome sequence was 64.2 mol%. The results of physiological and biochemical tests and 16S rRNA sequence analysis clearly revealed that strain Y3S6T represents a novel species in the genus Halomonas, for which the name Halomonas antri sp. nov. has been proposed. The type strain is Y3S6T (=KACC 21536T=NBRC 114315=TBRC 15164T).
Silver nanomaterials have potent antibacterial properties that are the foundation for their wide commercial use as well as for concerns about their unintended environmental impact. The nanoparticles themselves are relatively biologically inert but they can undergo oxidative dissolution yielding toxic silver ions. A quantitative relationship between silver material structure and dissolution, and thus antimicrobial activity, has yet to be established. Here, this dissolution process and associated biological activity is characterized using uniform nanoparticles with variable dimension, shape, and surface chemistry. From this, a phenomenological model emerges that quantitatively relates material structure to both silver dissolution and microbial toxicity. Shape has the most profound influence on antibacterial activity, and surprisingly, surface coatings the least. These results illustrate how material structure may be optimized for antimicrobial properties and suggest strategies for minimizing silver nanoparticle effects on microbes.
Ion channels, which can be modulated by peptides, are promising drug targets for neurological, metabolic, and cardiovascular disorders. Because it is expensive and labor-intensive to experimentally screen ion channel-modulating peptides (IMPs), in-silico approaches can serve as excellent alternatives. In this study, we present PrIMP, prediction models for screening IMPs that can target sodium, potassium, and calcium ion channels, as well as nicotine acetylcholine receptors (nAChRs). To overcome the data insufficiency of the IMPs, we utilized two types of knowledge transfer approaches: multi-task learning (MTL) and transfer learning (TL). MTL enabled model training for four target tasks simultaneously with hard parameter sharing, thereby increasing model generalization. TL transferred knowledge of pre-trained model weights from antimicrobial peptide data, which was a much larger, naturally-occurring functional peptide dataset that could potentially improve the model performance. MTL and TL successfully improved the prediction performance of prediction models. In addition, a hybrid approach by implementing deep learning along with traditional machine learning was utilized, with additional performance improvements. PrIMP achieved F1 scores of 0.924 (sodium ion channel), 0.937 (potassium ion channel), 0.898 (calcium ion channel), and 0.931 (nAChRs). The pre-processed dataset and proposed model are available at https://github.com/bzlee-bio/PrIMP .
Atopic dermatitis (AD) is one of the most common inflammatory skin diseases, which significantly impact the quality of life. Transcriptome-wide association study (TWAS) was conducted to estimate both transcriptomic and genomic features of AD and detected significant associations between 31 expression quantitative loci and 25 genes. Our results replicated well-known genetic markers for AD, as well as 4 novel associated genes. Next, transcriptome meta-analysis was conducted with 5 studies retrieved from public databases and identified 5 additional novel susceptibility genes for AD. Applying the connectivity map to the results from TWAS and meta-analysis, robustly enriched perturbations were identified and their chemical or functional properties were analyzed. Here, we report the first research on integrative approaches for an AD, combining TWAS and transcriptome meta-analysis. Together, our findings could provide a comprehensive understanding of the pathophysiologic mechanisms of AD and suggest potential drug candidates as alternative treatment options.
Systemic juvenile idiopathic arthritis (sJIA) is a rare subtype of juvenile idiopathic arthritis, whose clinical features are systemic fever and rash accompanied by painful joints and inflammation. Even though sJIA has been reported to be an autoinflammatory disorder, its exact pathogenesis remains unclear. In this study, we integrated a meta-analysis with a weighted gene co-expression network analysis (WGCNA) using 5 microarray datasets and an RNA sequencing dataset to understand the interconnection of susceptibility genes for sJIA. Using the integrative analysis, we identified a robust sJIA signature that consisted of 2 co-expressed gene sets comprising 103 up-regulated genes and 25 down-regulated genes in sJIA patients compared with healthy controls. Among the 128 sJIA signature genes, we identified an up-regulated cluster of 11 genes and a down-regulated cluster of 4 genes, which may play key roles in the pathogenesis of sJIA. We then detected 10 bioactive molecules targeting the significant gene clusters as potential novel drug candidates for sJIA using an in silico drug repositioning analysis. These findings suggest that the gene clusters may be potential genetic markers of sJIA and 10 drug candidates can contribute to the development of new therapeutic options for sJIA.