Sparse non-Gaussian component analysis is an unsupervised linear method of extracting any structure from high-dimensional distributed data based on estimating a low-dimensional non-Gaussian data component. In this paper we discuss a new approach with known apriori reduced dimension to direct estimation of the projector on the target space using semidefinite programming. The new approach avoids the estimation of the data covariance matrix and overcomes the traditional separation of element estimation of the target space and target space reconstruction. This allows to reduced the sampling size while improving the sensitivity to a broad variety of deviations from normality. Moreover the complexity of the new approach is limited to O(dlogd). We also discuss the procedures which allows to recover the structure when its effective dimension is unknown.
Non-Gaussian component analysis (NGCA) introduced in offered a method for high-dimensional data analysis allowing for identifying a low-dimensional non-Gaussian component of the whole distribution in an iterative and structure adaptive way. An important step of the NGCA procedure is identification of the non-Gaussian subspace using principle component analysis (PCA) method. This article proposes a new approach to NGCA called sparse NGCA which replaces the PCA-based procedure with a new the algorithm we refer to as convex projection.
Non-gaussian component analysis (NGCA) introduced in [24] offered a method for high dimensional data analysis allowing for identifying a low-dimensional non-Gaussian component of the whole distribution in an iterative and structure adaptive way. An important step of the NGCA procedure is identification of the non-Gaussian subspace using Principle Component Analysis (PCA) method. This article proposes a new approach to NGCA called sparse NGCA which replaces the PCAbased procedure with a new the algorithm we refer to as convex projection. keywords: reduction of dimensionality, model reduction, sparsity, variable selection, principle component analysis, structural adaptation, convex projection Mathematical Subject Classification: 62G05, 60G10, 60G35, 62M10, 93E10 Supported by DFG research center Matheon ”Mathematics for key technologies” (FZT 86) in Berlin.
In technical chemistry, systems biology and biotechnology, the construction of predictive models has become an essential step in process design and product optimization. Accurate modelling of the reactions requires detailed knowledge about the processes involved. However, when concerned with the development of new products and production techniques for example, this knowledge often is not available due to the lack of experimental data. Thus, when one has to work with a selection of proposed models, the main tasks of early development is to discriminate these models. In this article, a new statistical approach to model discrimination is described that ranks models wrt. the probability with which they reproduce the given data. The article introduces the new approach, discusses its statistical background, presents numerical techniques for its implementation and illustrates the application to examples from biokinetics.
Background: The computer-assisted detection of small molecules by mass spectrometry in biological samples provides a snapshot of thousands of peptides, protein fragments and proteins in biological samples. This new analytical technology has the potential to identify disease associated proteomic patterns in blood serum. However, the presently available bioinformatic tools are not sensitive enough to identify clinically important low abundant proteins as hormons or tumor markers with only low blood concentrations. Aim: Find, analyze and compare serum proteom patterns in groups of human subjects having different properties such as disease status with a new workflow to enhance sensitivity and specificity. Problems: Mass data acquired from high-throughput platforms frequently are blurred and noisy. This complicates the reliable identification of peaks in general and very small peaks even below noise level in particular. However, this statement is only valid for single or few spectra. If the algorithm has access to a large number of spectra (e.g. N > 1000), new possibilities arise, one of such being a statistical approach. Approach: Apply signal preprocessing steps followed by statistical analyses of the blurred data and the region below the typical noise threshold to identify signals usually hidden below this "barrier". Results: A new analysis workflow has been developed that is able to accurately identify, analyze and determine peaks and their parameters even below noise level which other tools can not detect. A Comparison to commercial software has clearly proven this gain in sensitivity. These additional peaks can be used in subsequent steps to build better peak patterns for proteomic pattern analysis. We belive that this new approach will foster identification of new biomarkers having not been detectable by most algorithms currently available.
Background: The computer-assisted detection of small molecules by mass spectrometry in biological samples provides a snapshot of thou- sands of peptides, protein fragments and proteins in biological samples. This new analytical technology has the potential to identify disease asso- ciated proteomic patterns in blood serum. However, the presently avail- able bioinformatic tools are not sensitive enough to identify clinically important low abundant proteins as hormons or tumor markers with only low blood concentrations. Aim: Find, analyze and compare serum proteom patterns in groups of human subjects having different properties such as disease status with a new workflow to enhance sensitivity and specificity. Problems: Mass data acquired from high-throughput platforms fre- quently are blurred and noisy. This complicates the reliable identification of peaks in general and very small peaks even below noise level in par- ticular. 1 However, this statement is only valid for single or few spectra. If the algorithm has access to a large number of spectra (e.g. N > 1000), new possibilities arise, one of such being a statistical approach. Approach: Apply signal preprocessing steps followed by statistical ana- lyses of the blurred data and the region below the typical noise threshold to identify signals usually hidden below this "barrier". Results: A new analysis workflow has been developed that is able to accurately identify, analyze and determine peaks and their parameters even below noise level which other tools can not detect. A Comparison to commercial software 2 has clearly proven this gain in sensitivity. These additional peaks can be used in subsequent steps to build better peak patterns for proteomic pattern analysis. We belive that this new approach will foster identification of new biomarkers having not been detectable by most algorithms currently available.