目的 探索基于压缩感知理论变量筛选方法在小样本量蛋白质组学研究中应用的效果和特点,为小样本量的蛋白质组学的变量筛选提供更灵敏、可靠的方法.方法 模拟实验比较基于CS理论的变量筛选方法与偏最小二乘(PLS)及随机森林(RF)筛选变量的能力,通过灵敏度、特异度及平衡准确度评价其变量筛选效果;利用CS变量筛选方法筛选非小细胞肺癌两亚型组(腺癌和鳞状细胞癌)的差异蛋白.结果 模拟实验表明,CS理论的变量筛选方法在样本量较小时具有较好的变量筛选效果,灵敏度、特异度及平衡准确度均较高;利用基于CS理论的变量筛选方法筛选,获得肺腺癌和鳞状细胞癌间差异表达蛋白22种,被证明是肺腺癌和鳞状细胞癌间有差异的蛋白为:Cytokeratin 6A、Cytokeratin 6B、Cytokeratin 6C、PKP1、P63、MCT1.结论 基于CS理论的变量筛选方法在样本量特别少时,筛选变量的效果优于PLS和RF,更适用于小样本蛋白质组学数据变量筛选研究.
目的 建立权重概率主成分分析模型,通过模拟实验进行模型评价,选择最优模型进行代谢组学数据分析,为代谢组学数据分析提供降噪优化的分析方法.方法 使用折刀抽样法计算变量载荷的置信区间和变异系数,利用变量载荷的变异信息设计倒数式、开根式、对数式三种加权方式进行原始数据中的变量加权,结合概率主成分分析模型建立权重概率主成分分析模型;通过模拟实验从第一主成分载荷的估计和预测效能进行模型评价,选择最优权重概率主成分分析模型;绘制代谢组学数据主成分得分图,利用中心距离比较权重概率主成分分析模型与概率主成分分析模型在可视化分组效果.结果 倒数式加权概率模型在第一主成分载荷的估计和模型预测方面优于另外两种权重概率模型.在可视化方面,权重概率主成分分析不仅缩小了模型估计的不确定性,而且增大组间的中心距离.结论 构建了权重概率主成分分析模型,不仅结果解释和可视化优于概率主成分分析模型,而且为差异变量的筛选提供了一个较小的参考范围.
Objective This paper aims to identify differentially expressed metabolites between patients of purpura ne-phritis and those with anaphylactoid purpura.Methods Four ranked lists from Student’s t test,Wilcoxon rank sum test,partial least square,and the random forest,respectively,are aggregated to get a combined single ranking list for identifying differentially expressed variables between groups.The ranking aggregation model was cross-validated and compared with the Least absolute shrinkage and selection operator(LASSO)model using simulation,and then applied to an real dataset,in which the ability of i-dentifying differentially expressed metabolites between patients of purpura nephritis and those with anaphylactoid purpura were compared between these two methods.Results The simulation experiment shows that:(1)when the number of observations and differential variables are small,the average area under the curve(AUC)value from the rank aggregation model is greater than that from LASSO;(2)when the number of observations and differential variables are larger,two AUCs are similar to each other;(3) no matter how the parametersare set,the number of differential variables selected by the rank aggregation model is basically less than that selected by LASSO.The real dataset analysis finds that 12 differential variables in patients with purpura nephritis are i-dentified to be differentially expressed as compared with patients with anaphylactoid purpurausing the rank aggregation model, with the AUC reaching its highest value of 0. 96.Conclusion Comparing with the LASSO model,the rank aggregation model is more reliable and accurate when being used to select differentially expressed variables,which may provide us with a new idea in metabolomics data analysis.
代谢组学的概念自20世纪90年代被正式提出[1],已被广泛应用于医学研究领域,其一般研究流程包括样本采集、样本检测、数据预处理、数据分析和生物学解释等.常用的样本检测技术有核磁共振(nuclear magnetic resonance,NMR)和高分辨率色谱-质谱联用技术[2],本文所述方法针对后者.经色谱-质谱联用平台检测的数据具有以下特点:高维度,小样本,变量间高度相关,高噪声,高缺失,以及高度变异性.基于以上数据特征,在代谢组学数据分析之前需要对数据进行预处理[3],以消除或减小数据中高噪声、高缺失和高度变异性对统计分析结果的干扰.数据预处理是数据分析的前提,有利于统计分析过程的多变量模型构建;若数据未经过充分预处理而直接用于统计分析,可能会掩盖数据中某些真实特征(如组间差异),使研究结果模糊化.数据预处理包括缺失数据处理,数据标准化,以及数据的中心化、标度化和转换等内容.本文将全面介绍基于色谱-质谱联用平台的代谢组学数据预处理方法,并进行各方法比较,为研究者选择合适的数据预处理方法提供思路.
Objective To compare the effect of one cross-validation and multiple cross-validations on PLSDA optimal model and discuss the effect of multiple cross-validations on stability of the optimal model when a few individuals are wrong grouped and when all individuals are right grouped,respectively.Methods The order of individuals in one dataset was disor-ganized to perform multiple cross-validations.Simulative data and real data were analyzed using one cross-validation and multiple cross-validations.The variation and stability of the models were tested using parameters like principal component number and MSEP.Results For simulative data,the principal component number of one cross-validation is 3 and MSEP is 0.3792;for re-sult of 5000 cross-validations when the data is not disordered,the range of principal component number is 2~6 and the range of MSEP is 0.2569~0.5794;for result of 5000 cross-validations when the data is 5%disordered,the range of principal component number is 1 ~8 and the range of MSEP is 0.2061 ~0.6463;for result of 10000 times cross-validation of real data,the range of principal component number is 4~10 and the range of MSEP is 0.0802 ~0.3761.Conclusion PLSDA models built by one cross-validation are not stable whereas multiple cross-validations can help build PLSDA models more stably when a few individ-uals are wrong grouped.So multiple cross-validation is recommended to ensure the stability of PLSDA model.