BACKGROUND:The identification of new diagnostic or prognostic biomarkers is one of the main aims of clinical cancer research. Technologies like mass spectrometry are commonly being used in proteomic research. Mass spectrometry signals show the proteomic profiles of the individuals under study at a given time. These profiles correspond to the recording of a large number of proteins, much larger than the number of individuals. These variables come in addition to or to complete classical clinical variables. The objective of this study is to evaluate and compare the predictive ability of new and existing models combining mass spectrometry data and classical clinical variables. This study was conducted in the context of binary prediction.RESULTS:To achieve this goal, simulated data as well as a real dataset dedicated to the selection of proteomic markers of steatosis were used to evaluate the methods. The proposed methods meet the challenge of high-dimensional data and the selection of predictive markers by using penalization methods (Ridge, Lasso) and dimension reduction techniques (PLS), as well as a combination of both strategies through sparse PLS in the context of a binary class prediction. The methods were compared in terms of mean classification rate and their ability to select the true predictive values. These comparisons were done on clinical-only models, mass-spectrometry-only models and combined models.CONCLUSIONS:It was shown that models which combine both types of data can be more efficient than models that use only clinical or mass spectrometry data when the sample size of the dataset is large enough.
The identification of new diagnostic or prognostic biomarkers is one of the main aims of clinical cancer research. In recent years, there has been a growing interest in using mass spectrometry for the detection of such biomarkers. The MS signal resulting from MALDI-TOF measurements is contaminated by different sources of technical variations that can be removed by a prior pre-processing step. In particular, denoising makes it possible to remove the random noise contained in the signal. Wavelet methodology associated with thresholding is usually used for this purpose. In this study, we adapted two multivariate denoising methods that combine wavelets and PCA to MS data. The objective was to obtain better denoising of the data so as to extract the meaningful proteomic biological information from the raw spectra and reach meaningful clinical conclusions. The proposed methods were evaluated and compared with the classical soft thresholding denoising method using both real and simulated data sets. It was shown that taking into account common structures of the signals by adding a dimension reduction step on approximation coefficients through PCA provided more effective denoising when combined with soft thresholding on detail coefficients.
L'identification de nouveaux biomarqueurs diagnostiques ou pronostiques est un des objectifs majeurs en recherche clinique. L'utilisation des technologies a haut debit comme la spectrometrie de masse est prometteuse pour l'identification de tels marqueurs. A partir d'un prelevement de sang ou de tumeur par exemple, cette technologie permet de traduire sous forme de spectres le profil proteique des individus. Le signal biologique observe dans les spectres est masque par differentes sources de variabilites techniques, qu'une phase prealable de pretraitement doit permettre de retirer. La methode classique permettant de retirer le bruit aleatoire de mesure de ce signal combine la methodologie des ondelettes et un seuillage, ceci spectre a spectre. L'utilisation des ondelettes permet la separation du bruit aleatoire (coefficients de detail) et du signal biologique (coefficients d'approximation). Le seuillage des coefficients de details permet ensuite d'annuler un certain nombre d'entre eux. Nous proposons dans ce travail d'ameliorer le debruitage classique des donnees en tenant compte de la structure commune des signaux. Nous avons pour cela adapte aux donnees spectrometriques deux methodes de debruitage qui combinent les ondelettes, le seuillage et les analyses en composantes principales classique ou creuse. Les methodes proposees ont ete evaluees et comparees a la methode de seuillage univariee classique, ceci pour des donnees reelles et simulees. Il a ete montre que l'ajout d'une etape de reduction de la dimension sur les approximations par une analyse en composantes principales en plus d'un seuillage classique sur les details ameliore le debruitage.