The coronavirus disease 2019 (COVID-19) epidemic continues to spread rapidly around the world and nearly 20 millions people are infected. This paper utilises both single-locus analysis and joint-SNPs analysis for detection of significant single nucleotide polymorphisms (SNPs) in the phenotypes of symptomatic versus asymptomatic, the early collection time versus the late collection time, the old versus the young, and the male versus the female. Also, this paper analyses the relationship between any two SNPs via linkage disequilibrium analysis, and visualises the patterns of cumulative mutations of SNPs over collection time. The results are in three folds. First, the SNP which locates at the nucleotide position 4321 is found to be an independent significant locus associated with all the first three phenotypes. Moreover, 12 significant SNPs are found in the first two studies. Second, gene orf1ab containing SNP-4321 is detected to be significantly associated with the first three phenotypes, and the three genes S, ORF3a, and N, are detected to be significant in the first two phenotypes. Third, some of the detected genes or SNPs are related to the SARS-COV-2 as supported by literature survey, which indicates that the results here may be helpful for further investigation.
Traditional pipeline for the task of detecting phenotypic biomarkers is a two-stage implementation, i.e., differentially expressed candidates are identified by NT tests, and then a subset of the candidates are further detected by phenotype-targeted tests (PT test) for significant phenotypic features, where N is short for Normal data and T is for Treatment/Trouble data. Such a two-stage procedure has low detection power as they do not make full use of the information contained in the (T, N). In this paper, we apply the two-variate PT test which jointly considers tumor-adjacent data and tumor data for improving the detection power. We investigate its performance by experiments on real-world datasets of breast cancer, considering phenotypes including BMI, overall survival time, pathologic stage, and tumor size. The results show that the method has high detection power and is more reliable, and the tumor-adjacent normal data plays an important role in the detection of phenotypic biomarkers. Finally, we obtain a new finding that the gene TCTEX1D2 is significantly related to tumor size in breast cancer.
Many multivariate statistical methods have been applied to detect the difference between case and control population. However, it is difficult to control the false positive rate, especially under small sample size. Traditional family-wise error rate or false discovery rate adjusts the p values based on the distribution or ranks of p value in the same multiple testing. In this paper, we investigated the performance of integrating the Data-space boundary-based test (BBT) and Statistics-space BBT to control the false positive rate, under a previous proposed framework called Integrative Hypothesis Tests (IHT). The classification accuracy rate by Data-space BBT provides valuable information complementary to the p value from Statistics-space BBT. The simulation results demonstrated that the integration effectively controls the false positive rate even for small-sample-size cases. Experiments on the real-world dataset of bipolar disorder also validated the effectiveness of the integration.
Different from the common approaches that use either hypothesis test or classifier for biomarker discovery, we applied the integrative hypothesis test (IHT) that combined both to identifying miRNAs for differentiation between lung cancer and Chronic Obstructive Pulmonary Disease (shortly L-C differentiation) on GEO data set GSE24709, and further extended IHT implementation by bootstrapping aided ranking and mean-variance based reliability check, which outputs a list of the top-15 differentially expressed miRNAs that confirmed the previously reported 14 miRNAs for L-C differentiation from a very different perspective plus an additional one. Moreover, we conducted a literature survey for a further explanation via dividing the 15 miRNAs into subclasses based on known relevances to the two diseases. Also, every pair of 15 miRNAs is exhaustively examined on their joint effect via p-value, misclassification, and correlation, which identifies core pairs and linked cliques as joint miRNAs biomarkers.
Instead of the single nucleotide variants (SNVs) analysis, many joint-SNVs analysis methods were proposed to tackle the 'missing heritability problem' in the genome-wide association studies (GWASs). In this paper, we performed a comparative study on five typical methods for joint-SNVs analysis and a recently proposed method called Statistics-space Boundary-based test (S-space BBT). For a fair and comprehensive comparison, we conducted simulation experiments by considering dominant single variant, effect direction, minor allele frequency (MAF), odds ratio (OR) and the linkage disequilibrium (LD). The results indicated that the S-space BBT not only does not swamp the significant SNV but also maintains stronger detection power under different configurations. As a result, we applied the S-space BBT to the dataset of bipolar disorder and obtained a list of biomarkers. Besides, the literature researches were conducted to validate the reliability of the results.
Many joint-SNVs (single-nucleotide variants) analysis methods were proposed to tackle the 'missing heritability' problem, which emphasizes that the joint genetic variants can explain more heritability of traits and diseases. However, there is still lack of a systematic comparison and investigation on the relative strengths and weaknesses of these methods. In this paper, we evaluated their performance on extensive simulated data generated by varying sample size, linkage disequilibrium (LD), odds ratios (OR), and minor allele frequency (MAF), which aims to cover almost all scenarios encountered in practical applications. Results indicated that a method called Statistics-space Boundary Based Test (S-space BBT) showed stronger detection power than other methods. Results on a real dataset of gastric cancer for Korean population also validate the effectiveness of the S-space BBT method.
Single nucleotide variants (SNVs) have been discovered that they play crucial roles in disease pathogenesis as genetic factors. Featured by analyzing multiple SNVs in a biological module (e.g. exon, gene, etc.) collectively, the joint-SNVs studies are increasingly attractive in genome-wide association studies (GWASs), for which extensive efforts have been devoted to pursue effective multivariate methods. In this paper, we first reviewed several main streams of existing methods and their limitations in joint-SNVs studies. Then, we introduced a recently proposed novel method, namely statistic-space boundary based test (S-space BBT) to tackle these limitations. Via computational experiments on simulation datasets, not only we figured out the applicable scenarios for the six methods in considering the effect direction and whether the single significant is involved in, but also demonstrated the strong detecting sensitivity of S-space BBT under the different conditions of odds ratio, minor allele frequency, and the linkage disequilibrium. We anticipate that our study may provide clues for multivariate method selection, and that S-space BBT may play a promising role in the joint-SNVs analysis.