Unsupervised feature selection is an important problem, especially for high-dimensional data. However, until now, it has been scarcely studied and the existing algorithms cannot provide satisfying performance. Thus, in this paper, we propose a new unsupervised feature selection algorithm using similarity-based feature clustering, Feature Selection-based Feature Clustering (FSFC). FSFC removes redundant features according to the results of feature clustering based on feature similarity. First, it clusters the features according to their similarity. A new feature clustering algorithm is proposed, which overcomes the shortcomings of K-means. Second, it selects a representative feature from each cluster, which contains most interesting information of features in the cluster. The efficiency and effectiveness of FSFC are tested upon real-world data sets and compared with two representative unsupervised feature selection algorithms, Feature Selection Using Similarity (FSUS) and Multi-Cluster-based Feature Selection (MCFS) in terms of runtime, feature compression ratio, and the clustering results of K-means. The results show that FSFC can not only reduce the feature space in less time, but also significantly improve the clustering performance of K-means.
The Influenza A virus is prone to mutation and the ongoing research on its evolution is of great significance to its prevention and control. The WHO did not update the recommended vaccine after A/California/07/2009 was recommended as the vaccine strain, but the virus has been mutating. This paper proposes an Integrated-Clustering-Analysis (ICA) model to study the distribution and evolution of the influenza A (H1N1) virus. We discover the following interesting facts. Every year there is one major type of virus sequences, the number of which is the overwhelming majority of all sequences. Viral sequences after 2009 undergo cumulative changes as they deviate from the viral vaccine strain over time. According to the drift rate, the evolution process can be divided into three stages. The first stage is a high-speed mutation period from 2009 to 2011. In the second stage, from 2012 to 2014, the mutation speed drops continuously and keeps at a low level. The third stage, from 2015 to 2017, the mutation speed starts with a year jump and follows by two years trough. It seems that the evolution of the influenza A virus has a three years cycle, so we cautiously guess that the drift rate in 2018 would jump up again. The ICA model proposed in this paper can intuitively observe the process of virus type change.
SESSION TITLE: Lung Cancer II SESSION TYPE: Original Investigation Poster PRESENTED ON: Saturday, April 16, 2016 at 11:45 AM - 12:45 PM PURPOSE: Predict and identify the HLA-A*0201 restricted CTL epitopes with the strongest immunogengenicity of SOX2 antigen. METHODS: Predict and screen the natural epitopes with bioinformatics method. Synthesize peptides in vitro. T2 cells were induced by peptides, and then were used in T2 binding assay. The effective cells were acquired by inducing PBMCs with peptides, and the target cells were acquired by inducing T2 cells with peptides. Then they were used in ELISPOT hIFN-γ and LDH releasing tests. RESULTS: Natural epitopes were obtained: SMYLPGAEV (P1), TLMKKDKYT (P2) and ALGSMGSVV (P3). The binding force between P1 and MHC I increased with the increase of the concentration of peptide, and when the concentration was 50μmol/L and 100μmol/L P1 group was the highest. Effect of killing the target cells can be observed in P1 and P2 group, and the former can secrete more IFN-γ. CONCLUSIONS: SMYLPGAEV is the strongest HLA-A*0201 restricted CTL natural epitope of SOX2 antigen, and may induce a strong immune response in patients with tumor. CLINICAL IMPLICATIONS: This peptide may play an important role in tumor vaccine. DISCLOSURE: The following authors have nothing to disclose: Yu Wang, Jingyan Yuan, Jingyuan Xiao, Wei Li, Na Fan, Wenjing Deng, Boxuan Liu, Ping Fang, Yujie Zhong, Shuanying Yang No Product/Research Disclosure Information
The purpose of this research work was to investigate whether such a workflow including the phosphoproteins enriched by affinity purification, isolated and quantified by two-dimensional difference gel electrophoresis, the differential protein spots identified by mass spectrometry, was suitable to dissect the intact phosphoprotein profile that mediated cell signaling pathways. Endothelial cells with or without vascular endothelial growth factor (VEGF) induction were used in the experiment. The accidental error introduced by enrichment column was assayed by control-to-control 2D-DIGE analysis and the efficiency of the workflow was tested by control-to-sample 2D-DIGE analysis. The data indicated that a high level of accidental error was introduced by different phosphoprotein enrichment columns and could be reduced by slowing down the flow rate of the mobile phase at the expense of sensitivity and specificity of enrichment column. Dephosphorylation did not occur for most high abundance phosphoproteins based on the sensitivity of 2D-DIGE detection. The advantage of this workflow was that multiple phosphorylated proteins could be visualized on the 2D-gel directly, but DIGE minimal label only quantified the limited high or medium abundance phosphoproteins. The sensitivity and accuracy of 2D-DIGE measurement were still not good enough to dissect the intact phosphoprotein profile that involved in cell signal pathway. Consequently, phosphoproteomic laboratory workflow based on phosphoproteins enrichment strategy combined with 2D-DIGE quantification does not have enough sensitivity and accuracy to dissect the intact phosphoprotein profile that mediated cell signaling pathways. Low abundance and ionization of the phosphopeptides are the main reasons that lead to the failure of protein identification even if the phosphoproteins were enriched.
Qinbao Song (宋擒豹)合作论文数Faculty of Electronic and Information Engineering, Xi'an Jiaotong University1