Molecular signatures have been excessively reported for diagnosis of many cancers during the last 20 years. However, false-positive signatures are always found using statistical methods or machine learning approaches, and that makes subsequent biological experiments fail. Therefore, signature discovery has gradually become a non-mainstream work in bioinformatics. Actually, there are three critical weaknesses that make the identified signature unreliable. First of all, a signature is wrongly thought to be a gene set, each component of which keeps differential expressions between or among sample groups. Second, there may be many false-positive genes expressed differentially found, even if samples derived from cancer or normal group can be separated in one-dimensional space. Third, cross-platform validation results of a discovered signature are always poor. In order to solve these problems, we propose a new feature selection framework based on ensemble classification to discover signatures for cancer diagnosis. Meanwhile, a procedure for data transform among different expression profiles across different platforms is also designed. Signatures are found on simulation and real data representing different carcinomas across different platforms. Besides, false positives are suppressed. The experimental results demonstrate the effectiveness of our method.
In order to identify plant pentatricopeptide repeat (PPR) proteins, a framework of variable selection has been proposed. In fact, it is an effective feature selection strategy that focuses on the performance of classification. Random forest has been used as the classifier with certain variables automatically selected for discrimination between PPR functional and non-functional proteins. However, it is found that samples regarded as PPR functional proteins are wrongly classified in a high rate. In this paper, we plan to improve the framework in order to achieve better classification results. Modifications are made on the framework for better identifying PPR functional proteins. Instead of random forest, a hybrid ensemble classifier is built with its base classifiers derived from six different classification methods. Besides, an incremental strategy and a clustering by search in descending order are alternatively used for feature selection, which can effectively select the most representative variables for identification on PPR proteins. In addition, it can be found that different base classifiers alternately play an important role in the ensemble classifier with feature dimension increasing. The experimental results demonstrate the effectiveness of our improvements.
Machine learning plays an important role in computational intelligence and has been widely used in many engineering fields. Surface voids or bugholes frequently appearing on concrete surface after the casting process make the corresponding manual inspection time consuming, costly, labor intensive, and inconsistent. In order to make a better inspection of the concrete surface, automatic classification of concrete bugholes is needed. In this paper, a variable selection strategy is proposed for pursuing feature interpretability, together with an automatic ensemble classification designed for getting a better accuracy of the bughole classification. A texture feature deriving from the Gabor filter and gray-level run lengths is extracted in concrete surface images. Interpretable variables, which are also the components of the feature, are selected according to a presented cumulative voting strategy. An ensemble classifier with its base classifier automatically assigned is provided to detect whether a surface void exists in an image or not. Experimental results on 1000 image samples indicate the effectiveness of our method with a comparable prediction accuracy and model explicable.
Background: Diagnosis of hip joint plays an important role in early screening of hip diseases such as coxarthritis, heterotopic ossification, osteonecrosis of the femoral head, etc. Early detection of hip dysplasia on X-ray films may probably conduce to early treatment of patients, which can help to cure patients or relieve their pain as much as possible. There has been no method or tool for automatic diagnosis of hip dysplasia till now. Results: A semi-automatic method for diagnosis of hip dysplasia is proposed. Considering the complexity of medical imaging, the contour of acetabulum, femoral head, and the upper side of thigh-bone are manually marked. Feature points are extracted according to marked contours. Traditional knowledge-driven diagnostic criteria is abandoned. Instead, a data-driven diagnostic model for hip dysplasia is presented. Angles including CE, sharp, and Tonnis angle which are commonly measured in clinical diagnosis, are automatically obtained. Samples, each of which consists of these three angle values, are used for clustering according to their densities in a descending order. A three-dimensional normal distribution derived from the cluster is built and regarded as the parametric model for diagnosis of hip dysplasia. Experiments on 143 X-ray films including 286 samples (i.e., 143 left and 143 right hip joints) demonstrate the effectiveness of our method. According to the method, a computer-aided diagnosis tool is developed for the convenience of clinicians, which can be downloaded at http://www.bio-nefu.com/HIPindex/. The data used to support the findings of this study are available from the corresponding authors upon request. Conclusions: This data-driven method provides a more objective measurement of the angles. Besides, it provides a new criterion for diagnosis of hip dysplasia other than doctors' experience deriving from knowledge-driven clinical manual, which actually corresponds to very different way for clinical diagnosis of hip dysplasia.