Single-cell RNA sequencing (scRNA-seq) offers precise quantification of the transcriptome at an individual cell level, surpassing traditional bulk RNA-seq. Despite its advancements, the high-dimensional nature of scRNA-seq data complicates the extraction and visualization of underlying biological information. Various dimensionality reduction algorithms have been developed to aid biologists in uncovering cellular relationships, especially through clustering. However, the stability of single-cell dimensionality reduction has been largely overlooked, particularly the variation of neighbor relations or cluster quality in the reduced space. In this study, we generate alternative datasets by randomly shuffling rows or columns of the single-cell expression matrix. These alternative datasets are processed similarly to the original data. Stability is measured by comparing results from the original and alternative data using multiple metrics, including knn-preservation for neighbor relations, the Calinski-Harabasz, Davies-Bouldin, and Xie-Beni indexes for cluster internal evaluation, and the Jaccard Index for clustering consistency. Additionally, the RF-hierarchical metric evaluated the preservation of global metainformation. We employed Monocle 3, an R toolkit, to assess the stability of scRNA-seq visualizations using two popular methods: t-SNE and UMAP. Our evaluation involved six datasets, including two from the C. elegans Monocle 3 tutorial, two from the Single Cell Portal (PBMC ID 345 and islet ID 1526), and two comprising Mouse Retinal and Brain samples. Internal cluster scores varied after data shuffling, with the Jaccard Index of C. elegans-con dropping below 0.6, indicating instability, and the RF distance of SCP1526 reaching the maximum value. Our findings suggest that dimensionality reduction techniques vary in stability with shuffled inputs, indicating that claims of their robustness may be premature without considering input variations.
Background:The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). Results:CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.
BACKGROUND:Classical CD14+CD16- monocytes are elevated in biliary atresia (BA); however, their specific role in bile duct injury and the underlying regulatory mechanisms remain unclear. This study aimed to define their contribution to BA pathogenesis, focusing on the NF-κB signaling pathway. METHODS:Liver tissues and blood samples from patients with BA and controls were analyzed by single-cell RNA sequencing, flow cytometry, and immunofluorescence. A rhesus rotavirus-induced BA mouse model was used for anti-Ly6C monocyte depletion and NF-κB inhibition (dehydroxymethylepoxyquinomicin). Transcriptomic profiling and cytokine analysis revealed key molecular mechanisms. RESULTS:Classical monocytes were significantly enriched near the damaged bile ducts in patients with BA and positively correlated with liver injury severity. These monocytes exhibited NF-κB hyperactivation, marked by the upregulation of TNF, IL-1β, Cxcl2, and NLRP3 inflammasome components. RNA-seq revealed BA-specific monocyte clusters with enriched NF-κB signatures. The depletion of classical monocytes (anti-Ly6C) in rhesus rotavirus-induced BA mice reduced biliary inflammation, restored bile duct patency, and improved survival. Pharmacological NF-κB inhibition (dehydroxymethylepoxyquinomicin) similarly attenuated inflammation and liver dysfunction and improved survival in rhesus rotavirus-induced BA mice. CONCLUSIONS:Classical CD14+CD16- monocytes are spatially enriched and exhibit NF-κB hyperactivation in BA. Targeting these cells or their NF-κB axis represents a promising therapeutic strategy to mitigate disease progression.
Abstract Principal component analysis (PCA) of the Hi-C Pearson correlation matrix is the standard approach for identifying A/B chromatin compartments. Despite its widespread use, the relationship between the first principal component (PC1) and the underlying compartment structure remains insufficiently characterized, and computing PC1 can become computationally expensive for high-resolution Hi-C data. Here we investigate the role of the PC1 explained variance ratio in compartment analysis and show that chromosomes with strong compartment organization typically exhibit a dominant PC1 signal. Based on this observation, we propose HiCPEP, a heuristic algorithm that estimates the sign pattern and relative magnitude of PC1 directly from the Hi-C Pearson covariance matrix without performing explicit eigenvector decomposition. The method can operate from either a dense Pearson matrix for fast approximation or a sparse observed/expected (O/E) matrix to reduce memory usage. Furthermore, because many covariance columns exhibit PC1-like patterns when the compartment signal is strong, HiCPEP can be accelerated using random sampling without substantially reducing accuracy. Across multiple Hi-C datasets, HiCPEP consistently recovered compartment patterns with high similarity to reference PC1 vectors produced by standard PCA-based methods. Benchmark experiments show that HiCPEP achieves comparable accuracy while reducing computational cost in terms of runtime or memory usage. These results suggest that HiCPEP provides a practical alternative for efficient chromatin compartment analysis from large-scale Hi-C datasets. The HiCPEP implementation is freely available at https://github.com/ZhiRongDev/HiCPEP .
Cell states' complexity and heterogeneity pose significant challenges in uncovering biological patterns in high-dimensional single-cell data. To address this, we developed scGHSOM, an enhanced framework based on the Growing Hierarchical Self-Organizing Map (GHSOM), for hierarchical clustering and visualization of high-dimensional datasets such as Mass Cytometry by Time-Of-Flight (CyTOF) and single-cell RNA sequencing. scGHSOM organizes data hierarchically, expanding clusters to satisfy within- and between-cluster variation thresholds. We propose a novel Significant Attributes Identification algorithm within the scGHSOM framework to identify features that minimize intra-cluster variation while maximizing inter-cluster variation, enabling targeted data analysis. To enhance interpretability, scGHSOM introduces two visualization tools: the Cluster Feature Map, which highlights feature distributions across hierarchical clusters, and the Cluster Distribution Map, which visualizes leaf clusters as circles sized by data volume and filled with different colors to represent features such as cell types or other attributes. Performance evaluation on three CyTOF datasets demonstrates that scGHSOM is compatible with state-of-the-art methods. Specifically, it achieves the best CH index in two of the three datasets. Furthermore, the proposed visualization tools significantly improve clarity and efficiency in interpreting scGHSOM results, effectively revealing clustering patterns and features.
Neuroblastoma, the most common extracranial solid tumor in children, demonstrates heterogeneous genetic susceptibility. Long noncoding RNAs (lncRNAs) are emerging cancer regulators. LincRNA TDRG1 is implicated in carcinogenesis. However, the potential impact of its functional variant rs8506 C > T on neuroblastoma susceptibility has not been assessed previously. This case–control study enrolled 402 pathologically confirmed neuroblastoma patients and 473 age- and sex-matched controls from Jiangsu Province, China. TaqMan assays with rigorous quality control were used for genotyping. Genetic associations were evaluated using five genetic models (homozygous, heterozygous, dominant, recessive, and additive) via unconditional logistic regression. Stratified analyses by age (≤ 18 months vs. > 18 months), sex (male vs. female), tumor site (adrenal gland, retroperitoneal, or mediastinum), and clinical stage (stages I + II + 4 s vs. stages III + IV) were performed to evaluate the subgroup-specific effects of rs8506 C > T. In silico expression quantitative trait locus (eQTL) analysis was performed using genotype-tissue expression (GTEx) Project data, which encompasses a wide array of human tissues and cell types, such as the adrenal gland and neural-related tissues. Controls were in Hardy–Weinberg equilibrium (P = 0.712). The rs8506 CT genotype demonstrated a significant protective effect against neuroblastoma in the heterozygous model [adjusted odds ratio (OR) = 0.71, 95
The CCAT2 gene is associated with carcinogenesis, but its effect on neuroblastoma, the most common extracranial tumor in children, remains unclear. We conducted a case–control study involving 402 children with neuroblastoma and 473 children without neuroblastoma. TaqMan genotyping of two CCAT2 polymorphisms (rs3843549 A > G and rs6983267 T > G) was conducted for all participants. Correlations were analyzed by calculating the odds ratio (OR) and 95
Neuroblastoma, developed from the sympathetic nervous system, is a deadly childhood malignancy. There is an urgent need to elucidate its intricated etiology. MYCN amplification leads to aggressive neuroblastoma and represents a powerful marker of poor prognosis. However, the correlation between MYCN gene polymorphisms and neuroblastoma susceptibility remains largely unknown in Chinese Han children. We conducted a case-control study to evaluate the associations between MYCN gene polymorphisms and neuroblastoma susceptibility, involving 402 cases and 473 controls from Jiangsu Province, China. The association strength between the studied polymorphisms and neuroblastoma susceptibility was quantified using odds ratios and 95
Objective:Neuroblastoma is the most common extracranial solid tumor in children and has complex genetic underpinnings. Previous genome-wide association studies (GWASs) have identified many loci associated with neuroblastoma susceptibility; however, their application in risk prediction for Chinese children has not been systematically explored. This study seeks to enhance neuroblastoma risk prediction by validating these loci and evaluating their performance in polygenic risk models. Methods:We validated 35 GWAS-identified neuroblastoma susceptibility loci in a cohort of Chinese children, consisting of 402 neuroblastoma patients and 473 healthy controls. Genotyping these polymorphisms was conducted via the TaqMan method. Univariable and multivariable logistic regression analyses revealed the genetic loci significantly associated with neuroblastoma risk. We constructed polygenic risk models by combining these loci and assessed their predictive performance via area under the curve (AUC) analysis. We also established a polygenic risk scoring (PRS) model for risk prediction by adopting the PLINK method. Results:Fourteen loci, including ten protective polymorphisms from CASC15, BARD1, LMO1, HSD17B12, and HACE1, and four risk variants from BARD1, RSRC1, CPZ and MMP20 were significantly associated with neuroblastoma risk. Compared with single-gene model, the 8-gene model (AUC=0.72) and 13-gene model (AUC=0.73) demonstrated superior predictive performance. Additionally, a PRS incorporating six significant loci achieved an AUC of 0.66, effectively stratifying individuals into distinct risk categories regarding neuroblastoma susceptibility. A higher PRS was significantly associated with advanced International Neuroblastoma Staging System (INSS) stages, suggesting its potential for clinical risk stratification. Conclusions:Our findings validate multiple loci as neuroblastoma risk factors in Chinese children and demonstrate the utility of polygenic risk models, particularly the PRS, in improving risk prediction. These results suggest that integrating multiple genetic variants into a PRS can enhance neuroblastoma risk stratification and potentially improve early diagnosis by guiding targeted screening programs for high-risk children.
The N1-adenosine methylation (m1A) modification plays a significant role in various cancers. However, the functions of m1A modification genes and their variants in neuroblastoma remain to be elucidated. We conducted a case-control study involving 402 neuroblastoma patients and 473 cancer-free controls from China via the TaqMan genotyping method to evaluate m1A modification gene polymorphisms. Multivariate logistic regression analysis was conducted to estimate odds ratios (ORs) and 95
OBJECTIVE:Modification of 5-methylcytosine (m5C) exerts regulatory effects on RNA functionality, governing critical processes that include cell migration, survival, and differentiation. NSUN4, a demethylase responsible for generating the m5C modification, plays a pivotal role in carcinogenesis and cellular differentiation. To date, there have been no documented reports on the role of NSUN4 gene polymorphisms in neuroblastoma. METHODS:The authors investigated 402 neuroblastoma patients and 473 control subjects and identified 4 potential functional polymorphisms (rs10736428 A>C, rs3737744 G>A, rs10252 G>A, and rs41294484 C>T) with the TaqMan assay. Logistic regression analysis assessed the correlation in terms of the OR and 95% CI. Furthermore, rs10736428 and rs41294484 were stratified to assess their potential associations with increased risk of neuroblastoma. RESULTS:Individuals carrying the rs10736428 CC genotype exhibited a markedly increased risk of neuroblastoma development (adjusted OR 2.06, 95% CI 1.02-4.14, p = 0.044). Further stratified analyses revealed that individuals with the rs10736428 CC genotype exhibited heightened predisposition to neuroblastoma, particularly within the subgroups of male patients, patients with mediastinal tumors, and patients with tumors classified under the International Neuroblastoma Staging System as stages 3 and 4. Moreover, children with 1-4 risk genotypes also showed positive associations with mediastinal tumors. CONCLUSIONS:A strong association between the NSUN4 rs10736428 polymorphism and increased susceptibility to neuroblastoma has been identified.
Background:Neuroblastoma is the predominant extracranial solid tumor occurring in children, and genetic factors like genetic polymorphism play a crucial role in its etiology. In this study, we investigated the associations between three NEFL polymorphisms (rs11994014 G>A, rs2979704 T>C, and rs1059111 A>T) and neuroblastoma susceptibility in a cohort of 402 neuroblastoma patients and 473 controls from Jiangsu Province. Methods:Genotyping was determined using the TaqMan method. Genotype distributions between cases and controls were assessed via both univariate and multivariate logistic regression models to assess the associations between NEFL polymorphisms and neuroblastoma risk. Stratified analyses were performed based on age, sex, clinical stage, and site of origin to explore potential effect modifications and subgroup-specific associations. Results:In the overall analysis, no significant associations were found between any of the three NEFL polymorphisms and neuroblastoma risk. When subjects were grouped on the basis of the number of risk genotypes, no significant alteration in susceptibility was observed in children carrying three risk genotypes compared with controls carrying fewer risk genotypes. Stratified analyses based on age, sex, clinical stage, and site of origin also revealed no significant results. Conclusions:Our findings suggest that NEFL polymorphisms do not significantly modify neuroblastoma susceptibility in this population, suggesting that the previously reported neuroblastoma susceptibility loci in NEFL in Caucasians may not be consistent across different populations. Further research, including larger, more diverse cohorts, is necessary to clarify the potential role of NEFL and other genetic factors in neuroblastoma etiology.
Neuroblastoma is the most common extracranial solid tumour in children, and genetic susceptibility plays a crucial role in its development. The impact of tRNA Dimethyltransferase 1 (TRDMT1), a primary methyltransferase catalysing 5-methylcytosine (m5C) RNA modification, on neuroblastoma susceptibility remains unexplored. We conducted a case-control study involving 402 neuroblastoma patients and 473 controls from Jiangsu, China. TRDMT1 polymorphisms (rs7074891 T>C, rs10904887 T>C and rs2273734 C>T) were genotyped via the TaqMan assay. Logistic regression was used to assess odds ratios (ORs) and 95% confidence intervals (CIs), while stratification analysis and expression quantitative trait locus (eQTL) analysis were used to examine subgroup-specific effects and regulatory impacts. Additionally, clinical correlation analysis and survival analysis were performed on neuroblastoma datasets used to evaluate. The rs7074891 TC/CC genotype reduced neuroblastoma risk (adjusted OR = 0.75, 95% CI = 0.57-0.98, p = 0.036), especially in children aged ≤ 18 months and those with mediastinal-origin tumours. Conversely, the rs10904887 CC (adjusted OR = 1.73, 95% CI = 1.27-2.38, p = 0.0006) and rs2273734 TT genotypes (adjusted OR = 1.80, 95% CI = 1.09-2.97, p = 0.023) were associated with increased risk, with distinct subgroup-specific effects. Combined 1-3 risk genotypes further confirmed increased susceptibility (adjusted OR = 1.81, 95% CI = 1.38-2.37, p < 0.0001), particularly in males and older children. eQTL analysis revealed that the rs7074891 C and rs10904887 C alleles increased TRDMT1 expression, whereas the rs2273734 T allele decreased it. Elevated TRDMT1 expression was correlated with poor prognosis and high-risk clinical features. TRDMT1 polymorphisms are significantly associated with neuroblastoma susceptibility, providing insights into their genetic and epigenetic mechanisms and potential as biomarkers and therapeutic targets.
To evaluate the potential association between single nucleotide polymorphisms in the RAS gene and neuroblastoma risk, we examined four candidate SNPs within this gene. Our hospital-based case–control study included 402 cases and 473 controls. Four SNPs (rs12587 G > T, rs7973450 A > G, and rs7312175 G > A in KRAS and rs2273267 A > T in NRAS) were genotyped using the TaqMan assay. The association between RAS gene polymorphisms and neuroblastoma susceptibility was assessed through odds ratios and 95
Neuroblastoma, “ a malignancy originating from neural crest cells, is most commonly diagnosed in children and adolescents. Polymorphisms within the long noncoding RNA (lncRNA) HOXA distal transcript antisense RNA (HOTTIP) are believed to have the capacity to alter an individual’s susceptibility to various cancers. This study aimed to investigate the link between HOTTIP gene polymorphisms and neuroblastoma susceptibility. We identified the genotypes of two prevalent polymorphisms (rs3807598 and rs1859168) within the HOTTIP via the TaqMan assay in a cohort comprising 402 individuals diagnosed with neuroblastoma and 473 healthy controls. Logistic regression was used to evaluate the associations between the HOTTIP polymorphisms and the likelihood of neuroblastoma susceptibility. Additionally, the genotype-tissue expression (GTEx) database was used to investigate how these HOTTIP gene variations influence gene expression across different tissues. Our findings demonstrated a significant association between the rs1859168 C > A polymorphism and reduced neuroblastoma susceptibility (CA vs. CC: adjusted odds ratio (OR) = 0.55, 95
Abstract Background Neuroblastoma is a common malignant tumor stemming from the sympathetic nervous system in children, which is often life‐threatening. The genetics of neuroblastoma remains unclear. Studies have shown that miRNAs participate in the regulation of a broad spectrum of biological pathways. The abnormity in the miRNA is associated with the risk of various cancers, including neuroblastoma. However, research on the relationship of miRNA polymorphisms with neuroblastoma susceptibility is still in the initial stage. Methods In this research, a retrospective case–control study was conducted to explore whether miR‐100 rs1834306 A > G polymorphism is associated with neuroblastoma susceptibility. We enrolled 402 cases and 473 controls for the study. The logistic regression analysis was adopted to calculate odds ratios (ORs) and 95% confidence intervals (CIs) for the association between miR‐100 rs1834306 A > G and neuroblastoma risk. Results Our results elucidated that the miR‐100 rs1834306 A > G polymorphism was associated with the decreased risk of neuroblastoma (AG versus AA: adjusted OR = 0.72, 95% CI = 0.53–0.98, and P = 0.038). The subsequent stratified analysis further found that rs1834306 AG/GG genotype reduced the risk of neuroblastoma in the subgroup with tumors of the mediastinum origin (adjusted OR = 0.63, 95% CI = 0.41–0.95, and P = 0.029). Conclusions In summary, miR‐100 rs1834306 A > G polymorphism was shown to associate with decreased neuroblastoma risk in Chinese children, especially for neuroblastoma of mediastinum origin. This conclusion needs to be verified in additional large‐size case–control studies.
Background and objectivesand aims MicroRNAs (miRNAs) are endogenous small noncoding RNAs that regulate gene expression by either degrading or inhibiting the translation of mRNAs and have significant roles in the development of various tumors. A polymorphism (rs2682818) in miR-618 has been confirmed to be correlated with susceptibility to various cancers. Nonetheless, its role has not been investigated in neuroblastoma to date. Therefore, we assessed whether the miR-618 rs2682818 C>A polymorphism was correlated with neuroblastoma risk in the Chinese population.
Common genetic mutations are absent in neuroblastoma, one of the most common childhood tumours. As a demethylase of 5-methylcytosine (m5C) modification, TET1 plays an important role in tumourigenesis and differentiation. However, the association between TET1 gene polymorphisms and susceptibility to neuroblastoma has not been reported. Three TET1 gene polymorphisms (rs16925541 A > G, rs3998860 G > A and rs12781492 A > C) in 402 Chinese patients with neuroblastoma and 473 cancer-free controls were assessed using TaqMan. Multivariate logistic regression analysis was used to evaluate the association between TET1 gene polymorphisms and susceptibility to neuroblastoma. The GTEx database was used to analyse the impact of these polymorphisms on peripheral gene expression. The relationship between gene expression and prognosis was analysed using Kaplan–Meier analysis with the R2 platform. We found that both rs3998860 G > A and rs12781492 A > C were significantly associated with increased neuroblastoma risk. Stratified analysis further showed that rs3998860 G > A and rs12781492 A > C significantly increased neuroblastoma risk in certain subgroups. In the combined risk genotype model, 1–3 risk genotypes significantly increased risk of neuroblastoma compared with the 0 risk genotype. rs3998860 G > A and rs12781492 A > C were significantly associated with increased STOX1 mRNA expression in adrenal and whole blood, and high expression of STOX1 mRNA in adrenal and whole blood was significantly associated with worse prognosis. In summary, TET1 gene polymorphisms are significantly associated with increased neuroblastoma risk; further research is required for the potential mechanism and therapeutic prospects in neuroblastoma.
Drosophila insulators were the first DNA elements found to regulate gene expression by delimiting chromatin contacts. We still do not know how many of them exist and what impact they have on the Drosophila genome folding. Contrary to vertebrates, there is no evidence that fly insulators block cohesin-mediated chromatin loop extrusion. Therefore, their mechanism of action remains uncertain. To bridge these gaps, we mapped chromatin contacts in Drosophila cells lacking the key insulator proteins CTCF and Cp190. With this approach, we found hundreds of insulator elements. Their study indicates that Drosophila insulators play a minor role in the overall genome folding but affect chromatin contacts locally at many loci. Our observations argue that Cp190 promotes cobinding of other insulator proteins and that the model, where Drosophila insulators block chromatin contacts by forming loops, needs revision. Our insulator catalog provides an important resource to study mechanisms of genome folding.
Background Biliary atresia (BA) is a type of severe cholestatic childhood disease that may have a genetic component. miR-100 plays a key role in regulating cell apoptosis, proliferation, and inflammatory reactions. A single-nucleotide polymorphism in miR-100 has been proven to modulate susceptibility to various diseases. Methods We conducted a case-control retrospective study to explore the correlation between miR-100 gene polymorphism (rs1834306 A>G) and biliary atresia susceptibility in 484 Chinese patients and 1445 matched control subjects. Results Our results showed that rs1834306 A>G was correlated with a significantly increased risk for BA (GG vs. AA: adjusted odds ratio (OR) = 1.44, 95%confidence interval (CI) = 1.02–2.03, p = 0.041; and GG vs. AA/AG: adjusted OR = 1.39, 95%CI = 1.02–1.89, p = 0.036). Conclusions Our results showed that the rs1834306 A>G polymorphism is associated with an increased risk for BA and contributes to BA susceptibility.
Ting-Yi Sung合作论文数Institute of Information Science
Academia Sinica, Taiwan4