Classifying nodes in networks is a task with a wide range of applications. It can be particularly useful in anomaly and fraud detection. Many resources are invested in the task of fraud detection due to the high cost of fraud, and being able to automatically detect potential fraud quickly and precisely allows human investigators to work more efficiently. Many data analytic schemes have been put into use; however, schemes that bolster link analysis prove promising. This work builds upon the belief propagation algorithm for use in detecting collusion and other fraud schemes. We propose an algorithm called SNARE (Social Network Analysis for Risk Evaluation). By allowing one to use domain knowledge as well as link knowledge, the method was very successful for pinpointing misstated accounts in our sample of general ledger data, with a significant improvement over the default heuristic in true positive rates, and a lift factor of up to 6.5 (more than twice that of the default heuristic). We also apply SNARE to the task of graph labeling in general on publicly-available datasets. We show that with only some information about the nodes themselves in a network, we get surprisingly high accuracy of labels. Not only is SNARE applicable in a wide variety of domains, but it is also robust to the choice of parameters and highly scalable-linearly with the number of edges in a graph.
In recent years, there have been several large accounting frauds where a company's financial results have been intentionally misrepresented by billions of dollars. In response, regulatory bodies have mandated that auditors perform analytics on detailed financial data with the intent of discovering such misstatements. For a large auditing firm, this may mean analyzing millions of records from thousands of clients. This paper proposes techniques for automatic analysis of company general ledgers on such a large scale, identifying irregularities - which may indicate fraud or just honest errors - for additional review by auditors. These techniques have been implemented in a prototype system, called Sherlock, which combines aspects of both outlier detection and classification. In developing Sherlock, we faced three major challenges: developing an efficient process for obtaining data from many heterogeneous sources, training classifiers with only positive and unlabeled examples, and presenting information to auditors in an easily interpretable manner. In this paper, we describe how we addressed these challenges over the past two years and report on experiments evaluating Sherlock.
Summary: Using replicated human serum samples, we applied an error model for proteomic differential expression profiling for a high-resolution liquid chromatography-mass spectrometry (LC-MS) platform. The detailed noise analysis presented here uses an experimental design that separates variance caused by sample preparation from variance due to analytical equipment. An analytic approach based on a two-component error model was applied, and in combination with an existing data driven technique that utilizes local sample averaging, we characterized and quantified the noise variance as a function of mean peak intensity. The results indicate that for processed LC-MS data a constant coefficient of variation is dominant for high intensities, whereas a model for low intensities explains Poisson-like variations. This result leads to a quadratic variance model which is used for the estimation of sample preparation noise present in LC-MS data.
There is a well-recognized but unmet need for biological markers to characterize disease type, status, progression, and response to therapy in autoimmune diseases. We are developing and applying an integrated bioanalytical platform and clinical research program to facilitate comprehensive differential phenotyping of patient samples and enable the discovery of biomarkers. Our measurement platform includes microvolume laser scanning cytometry for the quantification of hundreds of cellular parameters in whole blood and other samples, liquid chromatography-mass spectrometry and gas chromatography-mass spectrometry for the quantification of proteins and low molecular weight biomolecules in serum and other fluids or tissues, and specific immunoassays for the quantification of trace proteins in serum. We describe the technologies and discuss initial applications to the analysis of subjects with rheumatoid arthritis (RA) and healthy controls.
A liquid chromatography-mass spectrometry (LC-MS) proteomics and metabolomics platform is presented for quantitative differential expression analysis. Proteome profiles obtained from 1.5μL of human serum show ∼5000 de-isotoped and quantifiable molecular ions. Approximately 1500 metabolites are observed from 100μL of serum. Quantification is based on reproducible sample preparation and linear signal intensity as a function of concentration. The platform is validated using human serum, but is generally applicable to all biological fluids and tissues. The median coefficient of variation (CV) for ∼5000 proteomic and ∼1500 metabolomic molecular ions is approximately 25%. For the case of C-reactive protein, results agree with quantification by immunoassay. The independent contributions of two sources of variance, namely sample preparation and LC-MS analysis, are respectively quantified as 20.4 and 15.1% for the proteome, and 19.5 and 13.5% for the metabolome, for median CV values. Furthermore, biological diversity for ∼20 healthy individuals is estimated by measuring the variance of ∼6500 proteomic and metabolomic molecular ions in sera for each sample; the median CV is 22.3% for the proteome and 16.7% for the metabolome. Finally, quantitative differential expression profiling is applied to a clinical study comparing healthy individuals and rheumatoid arthritis (RA) patients.
A new method is presented for quantifying proteomic and metabolomic profile data by liquid chromatography-mass spectrometry (LC-MS) with electrospray ionization. This biotechnology provides differential expression measurements and enables the discovery of biological markers (biomarkers). Work presented here uses human serum but is applicable to any fluid or tissue. The approach relies on linearity of signal versus molecular concentration and reproducibility of sample processing. There is no use of isotopic labeling or chemically similar standard materials. Linear standard curves are reported for a variety of compounds introduced into human serum. As a measure of analytical reproducibility for proteome and metabolome sampling, median coefficients of variation of 25.7 and 23.8%, respectively, were determined for approximately 3400 molecular ions (not counting their numerous isotopes) from 25 independently processed human serum samples, corresponding to a total of 85000 individual molecular ion measurements.
We present a graphical method for evaluating the quality of a feature extraction mapping. Based on the Bilipschitz criterion, this Bilipschitz Criterion Plot (BCP) can be used to evaluate dimension reducing mappings for relative quality and to estimate the injectivity of the reduction map (as well as the associated reconstruction map). It can also be used to survey regions where the map is locally an expansion or contraction map. The plot is easy and fast to construct, and gives much more insight than any single value can, such as the distance preservation error. We demonstrate the value of such a mapping when examining the quality of the Sammon map, Neuroscale, the autoassociative map, and a recent technique that is designed to optimize the BCP in a linear fashion, the adaptive secant basis algorithm.
We propose a tool for filtering multivariate time series that was initially developed for analysing multi-spectral satellite imagery. The basic technique, known as the maximum noise fraction (MNF) method (Green et al. 1988), may be used to provide a subspace decomposition of a multivariate time series in terms of basis vectors which contain maximum noise (or maximum signal). We demonstrate the utility of the method for filtering nonsmooth multivariate data that includes high variance bands such as climate data. The methodology is applied to the reduction of data on noisy manifolds. A comparison of the approach to independent component analysis (ICA) is also provided.
A model validation test based on simple linear autocorrelation is proposed as an objective method to determine the optimal number of units in the hidden layer of a radial basis function network. The data to be fitted is assumed to consist of a signal with additive iid noise. A novel stopping criteria is introduced based on the statistics of the residuals rather than on ad hoc parameters. Consequently, this network is shown to neither overfit nor underfit the data. In addition, each new unit is adjusted to respond locally to the target data.
In this paper, we outline the relationship between the Maximum Noise Fraction (MNF) method–an algorithm first proposed for cleaning noise from multispectral image data– and Blind Signal Separation (BSS). In particular we demonstrate under what conditions these methods are equivalent and indicate that MNF may be viewed as an extension to BSS for the case of subspace mixing. We present several examples and compare the results of the MNF method to algorithms for performing independent component analysis (ICA).
We propose a clustering algorithm that dynamically inserts and relocates cluster units based only on their interaction with neighboring clusters and data points. This leads to update and allocation procedures for centers locations based on local data distributions. These local data distributions can be uncovered by examining neighboring clusters or local interconnections between center locations. The consequence of only adapting nearest centers to a newly inserted cluster unit is a significant reduction in the necessary computational power for finding the center distribution that reduces the global distortion error. The proposed algorithm inserts new cluster units based on local distortion errors and utility measures, and uses a local LBG routine to integrate the new unit. Experiments have shown that it is not necessary to let the LBG routine converge in order to achieve integration of the new unit; the number of necessary iterations is instead determined by the center distribution in the neighborhood of new units. The algorithm thus offers a considerable speedup compared to conventional clustering algorithms that take the entire data set into account when inserting or relocating cluster units.
Mary Mcglohon合作论文数Google Pittsburgh1