Computational phenotyping allows for unsupervised discovery of subgroups of patients as well as corresponding co-occurring medical conditions from electronic health records (EHR). Typically, EHR data contains demographic information, diagnoses and laboratory results. Discovering (novel) phenotypes has the potential to be of prognostic and therapeutic value. Providing medical practitioners with transparent and interpretable results is an important requirement and an essential part for advancing precision medicine. Low-rank data approximation methods such as matrix (e.g., non-negative matrix factorization) and tensor decompositions (e.g., CANDECOMP/PARAFAC) have demonstrated that they can provide such transparent and interpretable insights. Recent developments have adapted low-rank data approximation methods by incorporating different constraints and regularizations that facilitate interpretability further. In addition, they offer solutions for common challenges within EHR data such as high dimensionality, data sparsity and incompleteness. Especially extracting temporal phenotypes from longitudinal EHR has received much attention in recent years. In this paper, we provide a comprehensive review of low-rank approximation-based approaches for computational phenotyping. The existing literature is categorized into temporal vs. static phenotyping approaches based on matrix vs. tensor decompositions. Furthermore, we outline different approaches for the validation of phenotypes, i.e., the assessment of clinical significance.
The purpose of this study is to uncover cervical cancer (CC) risk phenotypes from self-reported lifestyle questionnaires and screening data. In general, computational phenotype discovery aims to find subgroups among individuals that share distinctive characteristics by analyzing electronic health records (EHR). This can benefit the understanding of a disease as well as uncover risk factors and provide possibilities for preventive action. The features in the women ( n = 6359 ) by questionnaire features ( p=29 ) matrix with missing data are of different statistical data types (e.g., binary or ordinal data). We use so-called generalized low-rank models (GLRM) that can address this challenge via different statistical-data-type-dependent loss functions. We show that these models can uncover phenotypes related to cervical cancer risk factors from large-scale questionnaire data.
With the rising trend of consumers being offered by start-up companies portable devices and applications for checking quality of purchased products, it appears of paramount importance to assess the reliability of miniaturized sensors embedded in such devices. Here, eight sensors were assessed for food fraud applications in skimmed milk powder. The performance was evaluated with dry- and wet-blended powders mimicking adulterated materials by addition of either ammonium sulfate, semicarbazide, or cornstarch in the range 0.5–10% of profit. The quality of the spectra was assessed for an adequate identification of the outliers prior to a deep assessment of performance for both non-targeted (soft independent modelling of class analogy, SIMCA) and targeted analyses (partial least square regression with orthogonal signal correction, OPLS). Here, we show that the sensors have generally difficulties in detecting adulterants at ca. 5% supplementation, and often fail in achieving adequate specificity and detection capability. This is a concern as they may mislead future users, particularly consumers, if they are intended to be developed for handheld devices available publicly in smartphone-based applications.
Hyperspectral sensor systems play a key role in the automation of work processes in the farming industry. Non-invasive measurements of plants allow for an assessment of the vitality and health state and can also be used to classify weeds or infected parts of a plant. However, one major downside of hyperspectral cameras is that they are not very cost-effective. In this paper, we show, that for specific tasks, multispectral systems with only a fraction of the wavelength bands and costs of a hyperspectral system can lead to promising results for regression and classification tasks. We conclude that for the ongoing automation efforts in the context of cognitive agriculture reduced multispectral systems are a viable alternative.
The consequences of food adulteration can be far reaching. In the past, inexpensive adulterants were used to inflate different products, leading to severe health issues. Contamination of food has many causes and can be physical (plant stems in tea), chemical (melamine in infant formula), or biological (bacterial contamination). Employing suitable sensor systems along the production process is a requirement for food safety. In this article, different approaches to food inspection are illustrated, and exemplary scenarios outline the potential of different sensor systems along the spectrum.
There are real world data sets where a linear approximation like the principalcomponents might not capture the intrinsic characteristics of the data. Nonlineardimensionality reduction ormanifoldlearning uses a graph-based approach tomodel the local structure of the data. Manifold learning algorithms assumethat the data resides on a low-dimensional manifold that is embedded in ahigher-dimensional space. For real world data sets this assumption might not beevident. However, using manifold learning for a classification task can reveal abetter performance than using a corresponding procedure that uses the principalcomponents of the data. We show that this is the case for our hyperspectral dataset using the two manifold learning algorithms Laplacian eigenmaps and locallylinear embedding.
Sensor-based sorting provides state-of-the-art solutions for sorting of cohesive, granular materials. Systems are tailored to a task at hand, for instance by means of sensors and implementation of data analysis. Conventional systems utilize scanning sensors which do not allow for extraction of motionrelated information of objects contained in a material feed. Recently, usage of area-scan cameras to overcome this disadvantage has been proposed. Multitarget tracking can then be used in order to accurately estimate the point in time and position at which any object will reach the separation stage. In this paper, utilizing motion information of objects which can be retrieved from multitarget tracking for the purpose of classification is proposed. Results show that corresponding features can significantly increase classification performance and eventually decrease the detection error of a sorting system.
Sensor-based sorting provides state-of-the-art solutions for sorting cohesive, granular materials. Typically, involved sensors, illumination, implementation of data analysis and other components are designed and chosen according to the sorting task at hand. A common property of conventional systems is the utilization of scanning sensors. However, the usage of area-scan cameras has recently been proposed. When observing objects at multiple time points, the corresponding paths can be reconstructed by using multiobject tracking. This in turn allows to accurately estimate the point in time and position at which any object will reach the separation stage of the optical sorter and hence contributes to decreasing the error in physical separation. In this paper, it is proposed to further exploit motion information for the purpose of material characterization. By deriving suitable features from the motion information, we show that high classification performance is obtained for an exemplary classification task. The approach therefore contributes towards decreasing the detection error of sorting systems.