Publisher: School of Statistics, Renmin University of China, Journal: Journal of Data Science, Title: Building an Honest Tree for Mass Spectra Classification Based on Prior Logarithm Normal Distribution, Authors: Cheng-Jian Xu, Ping He, Yi-Zeng Liang
Publisher: School of Statistics, Renmin University of China, Journal: Journal of Data Science, Title: The Matrix Expression, Topological Index and Atomic Attribute of Molecular Topological Structure, Authors: Qian-Nan Hu, Yi-Zeng Liang, Kai-Tai Yi-Zeng
Publisher: School of Statistics, Renmin University of China, Journal: Journal of Data Science, Title: Data Mining in Chemometrics - Sub-structures Learning via Peak Combinations Searching in Mass Spectra, Authors: Yu Tang, Yi-Zeng Liang
Publisher: School of Statistics, Renmin University of China, Journal: Journal of Data Science, Title: Application of Orthogonal Block Variables and Canonical Correlation Analysis in Modeling Pharmacological Activity of Alkaloids from Plant Medicines, Authors: Qian-Nan Hu, Yi-Zeng Liang
Variable selection is an important tool in QSAR. In this article, we employ three known techniques: sliced inverse regres- sion (SIR), principal components regression (PCR) and partial least squares regression (PLSR) for models to predict the boiling points of 530 saturated hydrocarbons. With 122 topological indices as in- put variables our results show that these three methods have good performance and perform better than some existing methods in the
Radix Sophorae Subprostratae (traditional Chinese medicine) was evaluated in a fixed system consisting of NaBrO3, H2SO4, Me2CO and MnSO4 do det. its fingerprint. Three oscillation areas were established. The results agreed with chromatog. detd. data.
Seven various Chinese green teas were characterized by using the nonlinear chem. fingerprint method with electro-chem. detection. The aq. suspensions of tea were treated with H2SO4, MnSO4 and Me2CO. The 1st grade Maojian tea (green bamboo mountain tea) showed the highest quality.
Partial least squares (PLS) have gained wide applications especially in chemometrics, metabolomics/metabonomics as well as bioinformatics. To our knowledge, an integrated PLS library that include not only basic PLS modeling algorithms but also advanced and/or recently developed methods on model assessment, outlier detection and variable selection is in lack. Here we present libPLS which provides an integrated platform for developing PLS regression and/or discriminant analysis (PLS-DA) models. This library is written in MATLAB and freely available at www.libpls.net.
Manifold learning classification, as an advanced semisupervised learning algorithm in recent years, has gained great popularity in a variety of fields. Moreover, kernel methods are a group of algorithms for pattern analysis, the task of which is to find and study general types of relations in datasets. Thus, under the framework of kernel methods, manifold learning classifier has been introduced and explored to directly detect the intrinsic similarity by local and global information hidden in datasets. Two validation approaches were used to evaluate the performance of our models. Experiments indicate that the proposed model can be considered as an effective and alternative modeling algorithm, and it could be further applied to the areas of biochemical science, environmental analysis, clinical, etc.
The rapid increase in the use of metabolite profiling/fingerprinting techniques to resolve complicated issues in metabolomics has stimulated demand for data processing techniques, such as alignment, to extract detailed information. In this study, a new and automated method was developed to correct the retention time shift of high-dimensional and high-throughput data sets. Information from the target chromatographic profiles was used to determine the standard profile as a reference for alignment. A novel, piecewise data partition strategy was applied for the determination of the target components in the standard profile as markers for alignment. An automated target search (ATS) method was proposed to find the exact retention times of the selected targets in other profiles for alignment. The linear interpolation technique (LIT) was employed to align the profiles prior to pattern recognition, comprehensive comparison analysis, and other data processing steps. In total, 94 metabolite profiles of ginseng were studied, including the most volatile secondary metabolites. The method used in this article could be an essential step in the extraction of information from high-throughput data acquired in the study of systems biology, metabolomics, and biomarker discovery.
In this study,a gold nanocrystal colloid was used as the enhancement factor for surfaceenhanced Raman scattering(SERS).Raman spectra were transformed by continuous wavelet transform (CWT),and Mexican hat wavelet were chosen as the wavelet basis.This procedure could be used to alleviate the influence of baseline variations and random noise,and find peak positions and the best scale wavelet coefficients of signal.Reverse search method was proposed to compare the spectrum of an unknown sample with a spectrum of standard using the information in wavelet space.Reverse match quality(RMQ) could be obtained automatically to determine whether a substance is present.It was used to identify colorants in a variety of food successfully.The colorants could be identified with 99 percent accuracy.It shows a better performance compared with traditional hit quality index (HQI).The study confirmed that the wavelet-based reverse search is feasible and accurate in qualitative analysis.
Identifying a small subset of genes that can classify disease samples from healthy controls plays an import role for evaluating disease risk and facilitating diagnosis. Existing methods often provide a single metric to assess predictive performances of genes. Also, model-based gene importance is conditioned on the subset of genes used to build multivariate models, and is thus model/context-specific. Existing methods often do not take into account such context-specific effects. Here we present a novel gene selection approach that evaluates predictive performance of genes using two criteria by taking into account gene interactions and project them onto four different regions in a 2-dimensional plot, like a phase diagram (PHADIA) in chemistry. Using two publicly available microarray datasets, we showed that PHADIA achieves comparable or better classification accuracies compared to reported results in the literature. The source codes are freely available at: www.libpls.net.
Near infrared spectroscopy (NIRS) is a kind of indirect analysis technology, whose application depends on the setting up of relevant calibration model. In order to improve interpretability, accuracy and modeling efficiency of the prediction model, wavelength selection becomes very important and it can minimize redundant information of near infrared spectrum. Intelligent optimization algorithm is a sort of commonly wavelength selection method which establishes algorithm model by mathematical abstraction from the background of biological behavior or movement form of material, then iterative calculation to solve combinatorial optimization problems. Its core strategy is screening effective wavelength points in multivariate calibration modeling by using some objective functions as a standard with successive approximation method. In this work, five intelligent optimization algorithms, including ant colony optimization (ACO), genetic algorithm (GA), particle swarm optimization (PSO), random frog (RF) and simulated annealing (SA) algorithm, were used to select characteristic wavelength from NIR data of tobacco leaf for determination of total nitrogen and nicotine content and together with partial least squares (PLS) to construct multiple correction models. The comparative analysis results of these models showed that, the total nitrogen optimums models of dataset A and B were PSO-PLS and GA-PLS models. GA-PLS and SA-PLS models were the optimums for nicotine, respectively. Although not all predicting performance of these optimization models was superior to that of full spectrum PLS models, they were simplified greatly and their forecasting accuracy, precision, interpretability and stability were improved. Therefore, this research will have great significance and plays an important role for the practical application. Meanwhile, it could be concluded that the informative wavelength combination for total nitrogen were 4 587~4 878 and 6 700~7 200 cm(-1), and that for tobacco nicotine were 4 500~4 700 and 5 800~6 000 cm(-1). These selected wavelengths have actually physical significance.
In contrast to the conventional tools,the recently developed software named “ChemDataSolution” has obvious advantages for multivariate modeling on the basis of intelligent strategy and methodology for data analysis.These features include,but not limited to,automatic loading of different types of file formats within a file folder,algorithm flow including a batch of methods for rapid and integrated data processing,simultaneous modeling,validation and prediction to multi-source data,and many other excellent characteristics for data handling.This makes ChemDataSolution become a promising tool to treat multivariate scientific instrumental data.
Kernel partial least squares (KPLS) has become popular techniques for chemical and biological modeling, which is a nonlinear extension of linear PLS. Training samples are transformed into a feature space via a nonlinear mapping, and then PLS algorithm can be carried out in the feature space. However, one of the main limitations of KPLS is that each feature is given the same importance in the kernel matrix, thus explaining the poor performance of KPLS for data with many irrelevant features. In this study, we provide a new strategy incorporated variable importance into KPLS, which is termed as the WKPLS approach. The WKPLS approach by modifying the kernel matrix provides a feasible way to differentiate between the true and noise variables. On the basis of the fact that the regression coefficients of the PLS model reflect the importance of variables, we firstly obtain the normalized regression coefficients by establishing the PLS model with all the variables. Then, Variable importance is incorporated into primary kernel. The performance of WKPLS is investigated with one simulated dataset and two structure–activity relationship (SAR) datasets. Compared with standard linear kernel PLS and Gaussian kernel PLS, The results show that WKPLS yields superior prediction performances to standard KPLS. WKPLS could be considered as a good mechanism by introducing extra information to improve the performance of KPLS for modeling SAR.
In order to minimize the influence of artificial experience on flue-cured tobacco leaf grading in purchasing process,a rapid grading method using near-infrared (NIR) spectroscopy combined with extreme learning machine (ELM) algorithm was proposed.A grouping method based on principle of similar quality and close price of flue-cured tobacco leaves was put forward.Cross validation was used to optimize the number of hidden nodes of ELM.The method was compared with commonly used multi-class classification algorithms,including K nearest neighbor (KNN),support vector machine (SVM),and random forest (RF) algorithm.Results showed that ELM classification model was superior to other methods with automatic optimization parameters,short training time,and high stability and predictability.The classification prediction accuracy of tobacco dataset A and B into high,medium,and low groups was 95.77% and 94.23%,respectively.Furthermore,classification accuracy of subdividing high,medium,and low groups of tobacco prediction samples A was 85.71%,86.67%,and 100%,respectively,and subdivision accuracy of tobacco prediction samples B was 100%,92.86% and 92.86%,respectively.Therefore,application of NIR technology combined with ELM could accurately determine flue-cured tobacco leaf grade,providing a promising tool for quality evaluation in flue-cured tobacco leaf purchasing process.
In order to achieve rapid quantitative analysis of fiber component content in cotton polyester spandex three components fabric,near infrared spectroscopy signature of 138 samples were collected.Model population analysis method was adopted to eliminate the abnormal samples.46 key wavelengths of the samples were screened out.The forecasting model of the rapid quantitative analysis fiber component content in cotton polyester spandex three components fabric was established by using partial least squares method.Correction factor of the established model is 0.96,cross-validation mean square residuum is 4.50 and prediction error root mean square is 4.59.The model can meet the precision requirement of quantitative analysis.
A simple, reliable and rapid isocratic liquid chromatography (LC)-mass spectrometric detection (MS) coupled with electrospray ionization (ESI) method for simultaneous separation and determination of calycosin-7-O-β-D-glucoside, ononin, calycosin and formonometin in Astragali Radix was developed. After the samples were extracted with ethanol, the optimum separation conditions for these analytes were achieved using water and acetonitrile (70:30, v/v) containing 0.2% (v/v) acetic acid as a mobile phase and a 2.0 mm×150 mm Hypersil-Keystone C18 column. Selective ion monitoring (SIM) mode and [M+H]+ ions at m/z 447, 431, 285 and 269 were used for quantitative analysis of four main active components above mentioned. The calibration curves were linear in the range of 0.4−175.0 μg/mL for calycosin-7-O-β-D-glucoside, 0.2−146.0 μg/mL for ononin, 0.4−210.0 μg/mL for calycosin and 0.5−217.0 μg/mL for formonetion, respectively. The limits of quantification (LOQ) and detection (LOD) were 0.4 μg/mL and 0.08 μg/mL for calycosin-7-O-β-D-glucoside, 0.2 μg/mL and 0.06 μg/mL for ononin, 0.4 μg/mL and 0.1 μg/mL for calycosin, 0.5 μg/mL and 0.1 μg/mL formonetion, respectively. The standard recoveries were in the range of 96.5%−104.7%. The developed method has successfully been used for the determination of four main flavonoids in Astragali Radix from various sources and can be used for identification, differentiation and quality evaluation of Astragali Radix.