Some generic modelling- and interpretation- problems in multivariate calibration of multichannel instruments (chemometric “machine learning”) are addressed wrt extrapolation, interpolation and interpretation. Multi-wavelength high-speed diffuse spectrophotometry in e.g. the Near InfraRed (NIR) wavelength range are information rich but require “mathematical cleanup”: Whether they are measured as transmittance, interactance or reflectance, such high-speed real-world data are affected by several different types of variation. These combine to offer certain data modelling challenges, like multicollinearity, mixed additive-and-multiplicative effects and response curvature. Through a progression of linear and bilinear preprocessing and calibration steps, multivariate data modelling is shown to pick up and correct for these challenges, with a strong focus on statistical validation wrt overfitting and on linear extrapolation power. A particular focus is on how to handle curvature, implicitly and explicitly. The lure of curvature is that mathematically useful, but causally meaningless linear modelling effects of nonlinear curvature may be interpreted as “new and interesting spectral details”. This is illustrated with high-precision NIR transmittance spectra of powder mixtures, from an experiment especially designed to reveal the “dirty” effects of light absorption and scattering. A simulation example demonstrates the lure of curvature explicitly. Finally, a successful linear approximation of nonlinear curvature is illustrated conceptually, by the approximation of a curved (nonlinear) 3D banana by a flat (linear) 2D boomerang.
ABSTRACT More rational, open‐minded use of quantitative Big Data in Science and Technology is required for better real‐world problem solving as well as for the stabilization of shared belief structures in society. Modern instrumentation gives informative but overwhelming data streams. A thermal video camera with suitable spatiotemporal subspace modeling allows us to detect surface temperature changes of, for example, engines, that can reveal something going on inside. An RGB video camera responds to both motions and color changes in nature, often with spatiotemporal change patterns that we can discover and describe mathematically, validate statistically, interpret graphically, and then use for sensible things. A hyperspectral Vis./NIR satellite camera with hundreds of wavelengths reveals changes in clouds and at each earth location, again and again. Today we know how to decode such overwhelming streams of high‐dimensional data into physical and chemical causalities by minimalistic hybrid multivariate subspace models. We thereby combine prior knowledge with the ability to discover new, reliable variation patterns. Minimalistic subspace models handle such data. These “open‐ended” multivariate linear hybrid models are computationally fast, statistically safe, and graphically understandable. The minimalistic subspace models are therefore suitable for both data modeling (based on multivariate measurements) and metamodeling (based on input–output simulation results for nonlinear mechanistic models' behavioral repertoire). That makes it easier to combine high‐dimensional streams of real‐world measurements and complicated, slow mechanistic models. Implemented as minimalistic foundation models with hierarchies of extended subspace models, this can form a basis for faster discovery and problem solving in Natural Science & Technology.
Extended Multiplicative Signal Correction (EMSC) is a multivariate linear modelling technique for multi-channel measurements that can identify and correct for different types of systematic variation patterns, known or unknown. It is typically used for pre-processing to separate light absorbance spectra, obtained by diffuse reflectance of intact samples, into three main sources of variation: additive variations due to chemical composition ( ≈ Beer’s law), mixed multiplicative and additive variations due to physical light scattering (≈ Lambert’s law) and more or less random measurement noise. The present work evaluates the use of EMSC to pre-process near infrared spectra obtained by hyperspectral imaging of Scots pine sapwood, inoculated with two different basidiomycete fungi and at various degradation stages. The spectral changes due to fungal decay and resulting mass loss are assessed by interpretation of the EMSC parameters and the partial least squares regression (PLSR) results. Including a cellulose (analyte) or bound water (interferent) spectral profile in the EMSC pre-processing model generally improves the predictive performance of the PLS modelling, but it can also make it worse. The inclusion of the additional polynomial baselines does not necessarily lead to a better separation of the physical and chemical effects present in the spectra. The estimated EMSC parameters provide insight into the differences in decay mechanisms. A detailed analysis of the EMSC results highlights advantages and disadvantages of using a complex pre-processing model.
Modern instruments generate BIG DATA that require information extraction before they can be used. A hybrid modelling framework for that is presented and illustrated. Its purpose is to convert meaningless data to meaningful information and to contribute to a theoretical, practical, and democratic basis for tomorrow's handling of BIG DATA in science and technology.
Hyperspectral imaging has recently gained increasing attention from academic and industrial world due to its capability of providing both spatial and physico-chemical information about the investigated objects. While this analytical approach is experiencing a substantial success and diffusion in very disparate scenarios, far less exploited is the possibility of collecting sequences of hyperspectral images over time for monitoring dynamic scenes. This trend is mainly justified by the fact that these so-called hyperspectral videos usually result in BIG DATA sets, requiring TBs of computer memory to be both stored and processed. Clearly, standard chemometric techniques do need to be somehow adapted or expanded to be capable of dealing with such massive amounts of information. In addition, hyperspectral video data are often affected by many different sources of variations in sample chemistry (for example, light absorption effects) and sample physics (light scattering effects) as well as by systematic errors (associated, e.g., to fluctuations in the behaviour of the light source and/or of the camera). Therefore, identifying, disentangling and interpreting all these distinct sources of information represents undoubtedly a challenging task. In view of all these aspects, the present work describes a multivariate hybrid modelling framework for the analysis of hyperspectral videos, which involves spatial, spectral and temporal parametrisations of both known and unknown chemical and physical phenomena underlying complex real-world systems. Such a framework encompasses three different computational steps: 1) motions ongoing within the inspected scene are estimated by optical flow analysis and compensated through IDLE modelling; 2) chemical variations are quantified and separated from physical variations by means of Extended Multiplicative Signal Correction (EMSC); 3) the resulting light scattering and light absorption data are subjected to the On-The-Fly Processing and summarised spectrally, spatially and over time. The developed methodology was here tested on a near-infrared hyperspectral video of a piece of wood undergoing drying. It led to a significant reduction of the size of the original measurements recorded and, at the same time, provided valuable information about systematic variations generated by the phenomena behind the monitored process.
NIR process monitoring and NIR hyperspectral video generates a deluge of non-selective spectral data, information-rich but per se useless. This paper demonstrates how interpretable data modelling can lead to simpler and better use of such NIR Big Data: A set of simple powder mixtures of the main constituents in wheat flour were measured by NIR transmission under different measurement conditions. Their absorbance spectra were submitted to multivariate calibration for predicting the protein content, by standard chemometric calibration by PLS regression. A reasonable calibration model was obtained, but it was unexpectedly complex and not robust. However, closer inspection the PLS regression subspace showed a surprising structure. This allowed us to identify the problem: Non-additive, strongly overlapping light scattering and light absorption effects in the NIR absorbance spectra. Based on this insight, a pragmatic, but causal preprocessing model was set up and iteratively optimized for predictive ability. This nonlinear optimized extended signal correction (OEMSC) separated and quantified the main physical and chemical sources of variation in the spectra. The preprocessing greatly simplified the NIR spectra and their quantitative calibration and prediction.
Hyperspectral cameras provide high spectral resolution data, but their usual low spatial resolution when compared to color (RGB) instruments is still a limitation for more detailed studies. This article presents a simple yet powerful method for fusing co-registered high spatial and low spectral resolution image data – e.g. RGB – with low spatial and high spectral resolution data – Hyperspectral. The proposed method exploits the overlap in observed phenomena by the two cameras to create a model through least square projections. This yields two images: 1) A high-resolution image spatially correlated with the input RGB image but with more spectral information than just the 3 RGB bands. 2) A low-resolution image showing the spectral information what is spatially uncorrelated with the RGB image. We show results for semi-artificial benchmark datasets and a real-world application. Performance metrics indicate the method is well suited for data enhancement.
Hyperspectral videos—multi-wavelength imaging of objects over time—generate a lot of informative data. But such diffuse spectroscopy measurements are usually non-selective, i.e., they respond to many different phenomena at the same time. To become quantitative, reliable and understandable, they require efficient mathematical data modeling. This article concerns how to model both known and unknown variation types in hyperspectral video data.
Over the latest decades, there has been a rapid development of laboratory techniques used for analyzing the pathways from the genes to the final physiology of an organism. The study of all genes in an organism is defined as genomics. In analog to this, studies of all the features along the path from the genes to the final phenotype is called transcriptomics, proteomics, metabolomics, lipidomics etc., with omics as a common name that covers them all. The comprehensive information that is now available on all the features in the cell, opens insight into biological mechanisms controlling growth and development of an organism at a level that was never before possible. This can be utilized, for example, in health care for improved diagnostics, improved medical treatment of diseases, and for personalized measures and preventive advice. In agricultural and food science, it opens opportunities for tailoring plants and animals with desirable properties, and for better microbiological control to improve food quality and safety. The data that are generated are comprehensive. Both from a biological and from a data analytical point of view, scientists are facing major challenges in the research area of analyzing the functionality that the omics data may uncover. The field of functional omics covers this. For each platform of molecular, chemical and biochemical analysis that is applied, special attentions needs to be drawn on pre-processing the data, which is beyond the scope of the present book chapter. This chapter aims to address some of the generic data analytical challenges in analyzing functional omics, which takes into realization the complexity of the comprehensive information that is uncovered by detailed studies on the molecular processes in the cells. We presents guidelines to some useful approaches for data analysis, focusing on building bridges between those that have insight into the biochemistry and biology of the cell and those that have insight into data modeling. This article points to important challenges to be considered in the analysis of omics data and it goes through practical analysis of several data sets to illustrate different useful multivariate methodologies that builds on a chemometric mindset of gaining insight into the phenomenon under study.
A new method is presented for extending a dynamic model of a six degrees of freedom robotic manipulator. A non-linear multivariate calibration of input-output training data from several typical motion trajectories is carried out with the aim of predicting the model systematic output error at time (t + 1) from known input reference up till and including time (t). A new partial least squares regression (PLSR) based method, nominal PLSR with interactions was developed and used to handle, unmodelled non-linearities. The performance of the new method is compared with least squares (LS). Different cross-validation schemes were compared in order to assess the sampling of the state space based on conventional trajectories. The method developed in the paper can be used as fault monitoring mechanism and early warning system for sensor failure. The results show that the suggested methods improves trajectory tracking performance of the robotic manipulator by extending the initial dynamic model of the manipulator.
Airborne hyperspectral imaging is a powerful technique for high-resolution classification of large areas of ground, applied today in fields like agriculture and environmental monitoring. Even though many classification algorithms are capable of handling shadows without a decrease in performance, visual inspection can be made easier if shadows are removed. In this paper we present a method for separating the effect of shadows (de-shadowing) and other partially known lighting condition changes from the effects due to the physical, chemical or biological properties of the ground, which are of interest. An example application is shown with good results.
This article shows how multivariate data modeling and multivariate metamodeling bridge the math‐gap in many sciences, e.g. in the biosciences. Biomedical science is constituted by two traditionally diverging lines of research, one being explorative, inductive, and holistic and the other being confirmative, deductive, and reductionistic. We argue in this article that both cultures are in dire need of more cross‐disciplinarity and openness to the knowledge and opportunities that lie on the other side of what we here call the ‘Math‐Gap in Bioscience’: The former tradition, which we choose to call ‘The House of Bio’, needs to expand to more powerful data‐driven statistical modeling and assessments in order to make sense out of their deluge of measurement data. The latter, which we here call ‘The House of Math’, needs to develop more and better mechanistic models to make possible the quantitative combination and effective utilization of today's enormous growth in biomedical knowledge and data. However, the use of mathematics and statistics in biomedical fields is hampered by lack of contact, curiosity, and respect between these two research cultures. To bridge this ‘Math‐Gap’ between them, we here propose a simple but powerful toolbox: soft multivariate modeling based on graphically interpreted, cross‐validated bilinear analysis. Nine different examples of data modeling and metamodeling show how this is successfully used in widely different fields of biomedical research – and from both research traditions.
4,627,014 12/1986 Lo et al. ............................ 364/571.01 4,642,778 2/1987 Hieftje et al............................ 364/498 4,744,657 5/1988 Aralis et al. ... 364/571.04X 4,782,456 l/1988 Poussier et al. ........................ 364/574 4,802,102 1/1989 Lacey ............ ... 364/497 4,884,213 1/1989 Iwata et al. .... 64,571.01 X 4,916,645 4/1990 Wuest et al........ 364,571.01 X 4,975,581 12/1990 Robinson et al. ...................... 250/339 5,046,846 9/1991 Ray et al. ........................... 364/498 X 5,081,597 1/1992 Kowalski. . 364/57.02 X 5,083,283 1/1992 Imai et al. ..... 364/571.02 X 5,369,578 11/1994 Roscoe et al. .......................... 364,422
Additional file 2: Table S2. Results for calibration of FTIR-spectra against GC-FID reference data, correlations between each FA and total fat percentage, and estimates of variance components with standard errors.
s of the lectures A New Infrared Spectroscopic Point-of-care Diagnostic for the Detection and Quantification of Pathogens in Red Blood Cells Bayden R. Wood, Phil Heraud, Anja Rüther, Brian Cooke, David Perez-Guaita Centre for Biospectroscopy, Monash University, Wellington Rd. Clayton Attenuated Total Reflection Fourier Transform Infrared (ATR-FTIR) spectroscopy, in combination with advanced computational modelling, offers tremendous potential for simultaneous, point-of-care diagnosis of multiple infectious diseases [1]. We have demonstrated the potential of the technique to detect parasite concentrations of 1/100,000 from packed red blood cells [2]. More recently we demonstrated how the approach can be used to simultaneously quantify malaria parasitemia, glucose and urea levels from a dried whole blood spot on a glass slide [3]. The ATR-FTIR approach is robust making it an ideal technology for screening of humans and other animals in any setting from remote regions of the developing world to hospital pathology labs or to mass screening in the modern clinical laboratory. ATR spectroscopy relies on detecting the molecular phenotype of the pathogen directly, and can discern malaria and Babesia parasites in red blood cells. Like malaria, Babesia is an Apicomplexan parasite that causes the disease known as Babesiosis, which severely affects cattle. We have developed a methodology that utilises lysis and concentration to isolate and purify malaria and babesia pathogens prior to ATR-FTIR analysis. The talk will focus on the application of the technology to malaria diagnosis in the field and provide preliminary results on the detection and quantification of pathogens that cause Babesiosis in cattle.
There was a mistake on Page 567, Section 2.4, Column 1, Line 19, where words “three in” were incorrectly included.1 The sentence “No more than three in 0.0003% of the mobile users are expected to develop brain tumor because of their phone use” should read: “No more than 0.0003% of the mobile users are expected to develop brain tumor because of their phone use.”
The aim of this paper is to achieve a model for prediction of cerebral palsy based on motion data of young infants. The prediction is formulated as a classification problem to assign each of the infants to one of the healthy or with cerebral palsy groups. Unlike formerly proposed features that are mostly defined in the time domain, this study proposes a set of features derived from frequency analysis of infants' motions. Since cerebral palsy affects the variability of the motions, and frequency analysis is an intuitive way of studying variability, suggested features are suitable and consistent with the nature of the condition. In the current application, a well-known problem, few subjects and many features, was initially encountered. In such a case, most classifiers get trapped in a suboptimal model and, consequently, fail to provide sufficient prediction accuracy. To solve this problem, a feature selection method that determines features with significant predictive ability is proposed. The feature selection method decreases the risk of false discovery and, therefore, the prediction model is more likely to be valid and generalizable for future use. A detailed study is performed on the proposed features and the feature selection method: the classification results confirm their applicability. Achieved sensitivity of 86%, specificity of 92% and accuracy of 91% are comparable with state-of-the-art clinical and expert-based methods for predicting cerebral palsy.
In this paper we aim at predicting cerebral palsy, the most serious and lifelong motor function disorder in children, at an early age by analysing infants' motion data. An essential step for doing so is to extract informative features with high class separability. We propose a set of features derived from frequency analysis of the motion data. Then, we evaluate the practicality of our features on one of the richest data sets collected to study this disease. In this data set, the motion data are extracted from both electromagnetic sensors as well as video camera. The proposed features are used for classifying both data sets. Using these features, we manage to achieve promising classification performance. Classification accuracy of 91% for the sensor data and 88% for the video-derived data show not only the advantage of employing these features for predicting cerebral palsy, but also that replacing electromagnetic sensors with a video camera is feasible.
This year we celebrate the 150th anniversary of the law of mass action. This law is often assumed to have been "there" forever, but it has its own history, background, and a definite starting point. The law has had an impact on chemistry, biochemistry, biomathematics, and systems biology that is difficult to overestimate. It is easily recognized that it is the direct basis for computational enzyme kinetics, ecological systems models, and models for the spread of diseases. The article reviews the explicit and implicit role of the law of mass action in systems biology and reveals how the original, more general formulation of the law emerged one hundred years later ab initio as a very general, canonical representation of biological processes.