Selection of reliable cancer biomarkers is crucial for gene expression profile-based precise diagnosis of cancer type and successful treatment. However, current studies are confronted with overfitting and dimensionality curse in tumor classification and false positives in the identification of cancer biomarkers. Here, we developed a novel gene-ranking method based on neighborhood rough set reduction for molecular cancer classification based on gene expression profile. Comparison with other methods such as PAM, ClaNC, Kruskal-Wallis rank sum test, and Relief-F, our method shows that only few top-ranked genes could achieve higher tumor classification accuracy. Moreover, although the selected genes are not typical of known oncogenes, they are found to play a crucial role in the occurrence of tumor through searching the scientific literature and analyzing protein interaction partners, which may be used as candidate cancer biomarkers.
Peptide fragments that serve as the cytotoxic T lymphocyte (CTL) epitopes are processed from antigens by the proteasome and then are transported to the endoplasmic reticulum through transporter associated with antigen processing (TAP) before being loaded onto the MHC class I molecule. Here, we studied TAP specificity by a neighborhood rough set (NRS) model based feature selection and prioritization method. By means of binary, amino acid properties, and binary plus properties of amino acids encoding, respectively, we adopted NRS based feature selection method to select multiple optimal feature sets for TAP binding peptides binary classification. The features in these optimal sets were ranked according to their occurrence frequency. Results show that the NRS is effective for prediction improvement and analysis of the specificity of TAP transporter. The proposed method can be used as a tool for predicting TAP binding peptides and be useful for subunit vaccine rational design and related bioinformatics cases.
We develop a knowledge-based statistical energy function on residual level for quantitatively predicting the affinity of protein-protein complexes by using 20 residue types and a distance-free reference state. The correlation coefficients between experimentally measured protein-protein binding affinities (PPIA) and the predicted affinities by our approach are 0.74 for 82 proteinprotein (peptide) complexes. Compared to the published results of two other volume corrected knowledge-based scoring functions on atomic level, the proposed approach not only is the simplest but also yields the comparable correlation between theoretical and experimental binding affinities of the test sets with the reported best methods.