Optical sensor arrays containing fluorescent solvatochromatic dyes immobilized in a plurality of polymers generate information-rich responses upon exposure to organic vapors. The response profiles are used to train a variety of computational networks such that subsequent exposure of the array to the vapors enables them to be classified and/or quantified. A number of strategies can be taken to enhance sensitivity and to increase sensor diversity.
The further development of a vapor-sensing device utilizing an array of broadly distributed optical sensors is detailed, Data from these optical sensors provided input to pattern-recognizing neural networks, which successfully identified and quantified a collection of 20 analyte vapors. The optical sensor array consisted of 19 optical fibers whose tips were coated with Nile Red immobilized in various polymer matrices. Responses consisted of the changes in fluorescence with time resulting from the presentation of a vapor to the sensor array. Numerical descriptors calculated from these responses were then used to highlight important temporal and spatial features. Learning vector quantization neural network models were constructed using subsets of these descriptors, and they accurately identified and quantified each of the presented analytes. Successful classification was achieved for both the training set data (89%) and for the external prediction set data (90%). Relative concentrations were correctly assigned for 90% of the prediction set data.
Computational neural networks have been developed to classify and quantify nine organic vapors. The neural network analyses used data that consisted of the change in fluorescence from a sensor array that consisted of 19 fiber optics with immobilized dye in polymer matrices. Plots of change in fluorescence intensity versus time were measured as pulses of analyte were presented to the sensor array. Descriptors were calculated from the intensity vs time plots, and they were used to build neural network models that accurately classified and quantified each of the nine analytes. Most of the data were used to train the neural networks (training set members), some were used to assist termination of training (cross-validation set members), and some were used to validate the models (prediction set members). Classification rates approaching 100% were achieved for the training set data, and 90% of the members in the prediction set were correctly classified. In addition, 97% of the prediction set observations were assigned a correct relative concentration.
The retention indices (RIs) of a set of alkylbenzenes on a polar gas chromatographic column are predicted directly from their molecular structures. Numerical descriptors are calculated based on the structure of a group of 150 alkylbenzenes. The descriptors are of three types: topological, geometric, and electronic. Statistical methods are employed to find an informative subset of these descriptors that can accurately predict the gas chromatographic RIs. The Automated Data Analysis and pattern Recognition Toolkit (ADAPT) software system is used to construct a large pool of structurally derived numerical descriptors which are used to build quantitative structure-retention relationships (QSRRs). Multiple linear regression analysis and computational neural networks are used to map the descriptors to the RIs.
The primary goal of a quantitative structure-property relationship (QSPR) is to identify a set of structurally based numerical descriptors that can be mathematically linked to a property of interest. The types of descriptors fall into three categories: topological, electronic, and geometric. In this study, 140 organic compounds with diverse structures were split into a training set, a cross-validation set, and a prediction set. The training set was used to build multiple linear regression and computational neural network models, the cross-validation set was used to prevent overtraining of the neural network, and the prediction set was used to validate the mathematical models. A set of nine descriptors was found that effectively linked the aqueous solubility to each structure. However, the polychlorinated biphenyls (PCBs) had a large root-mean-square (rms) error associated with them. Therefore models were also built using a training set that contained no PCBs. A set of nine descriptors was found with a significant improvement of the rms error of the training set as well as the prediction set.
Quantitative structure-property relationships (QSPRs) are used to develop mathematical models that accurately predict the reduced ion mobility constants (K(0)) for a set of 168 organic compounds directly from molecular structure. The K(0) values are taken from an unpublished database collected by G. A. Eiceman, Chemistry Department, New Mexico State University. The data were collected using a Graseby Ionics environmental vapour monitor (EVM) gas chromatography/ion mobility spectrometer. Standardized conditions with controlled temperature, pressure, and humidity were used, and 2,4-lutidine was used as an internal standard. K(0) values were measured for all monomer peaks. The best model was found with a feature selection routine which couples the genetic algorithm with multiple linear regression analysis. The set of six descriptors was also analyzed with a fully connected, feed-forward neural network. The model contains six molecular structure descriptors and has a root-mean-square error of about 0.04 K(0) unit. The descriptors in the model lend insight into some of the important molecular features that influence ion mobility. The model can be utilized for prediction of K(0) values of compounds for which there are no empirical K(0) data.
The central steps in developing QSARs are generation and selection of molecular structure descriptors and development of the model. Recently, computational neural networks have been employed as nonlinear models for QSARs. Neural networks can be trained efficiently with a quasi-Newton method, but the results are dependent on the descriptors used and: the initial parameters of the network. Thus, two potential opportunities for optimization arise. The first optimization problem is the selection of the descriptors for use by the neural network. In this study, generalized simulated annealing (GSA) is employed to select an optimal set of descriptors. The cost function used to evaluate the: effectiveness of the descriptors is based on the performance of the neural network. The second optimization problem is selecting the starting weights and biases for the network. GSA is also used for this optimization. The result is an automated descriptor selection algorithm that is an optimization inside of an optimization. Application of the method to a QSAR problem shows that effective descriptor subsets are found, and they support models that are as good or better than those obtained using traditional linear regression methods.
In order to obtain quantitative information, it is often necessary for a chemist to employ regression methods. An algorithm is described for determining the optimal subset of variables which gives the best prediction in regression analysis. The procedure is based on the generalized simulated annealing method (GSA) of optimization. From a comparison study with standard methods of variable subset selection by forward selection and backward elimination, GSA is found to perform better. Three data sets are used for distinction purposes. Two of the data sets are relatively small, allowing comparison to the global variable subset obtained by computing all possible variable combinations.
This paper utilizes variable step size generalized simulated annealing (VSGSA) to design multicomponent calibration samples for spectroscopic data. VSGSA is an optimization procedure which is capable of converging to exact positions of global optima located on multidimensional continuous functions. On the basis of analysis sample response vectors, optimally designed calibration concentration matrices are obtained assuming knowledge of components present. The complexity of response surfaces established by the optimization criteria is described.
Principal components (PCs) for principal component regression (PCR) have historically been selected from the top down for a reliable predictive model. That is, the PCs are arranged in a list starting with the most informative (PC associated with the largest singular value) and proceeding to the least informative (PC associated with the smallest singular value). PCs are then chosen starting at the top of this list. This paper discusses an alternative procedure of treating PC selection as an optimization problem. Specifically, without any regard to the ordering, the optimal subset of PCs for an acceptable predictive model is desired. Five data sets are analyzed using the conventional and alternative approaches. Two data sets are spectroscopic in nature, two data sets deal with quantitative structure-activity relationships (QSARs) and one data set is concerned with modeling. All five data sets confirm that selection of a subset without consideration to order secures the best results with PCR. One data set is also compared using partial least squares 1.
A promising procedure for computerized library searching and identification of ultraviolet-visible spectra for one- and two-component mixtures is evaluated. The procedure is based on singular value decomposition and utilizes band position and shape. Searching more than one library spectrum at a time is possible. Library spectra which have similar spectra (collinear) with the sample and each other are indicated by two measures. First, a large condition index must be obtained, and second, two or more large variance-decomposition proportions in the same row need to be associated with the large condition index. The searching procedure has a significant degree of differentiation between the actual sample and potential candidates and is compared with the dot product, Euclidean distance and the correlation coefficient.
ADVERTISEMENT RETURN TO ISSUEPREVArticleNEXTGlobal optimization by simulated annealing with wavelength selection for ultraviolet-visible spectrophotometryJohn H. Kalivas, Nancy. Roberts, and Jon M. SutterCite this: Anal. Chem. 1989, 61, 18, 2024–2030Publication Date (Print):September 15, 1989Publication History Published online1 May 2002Published inissue 15 September 1989https://doi.org/10.1021/ac00193a006RIGHTS & PERMISSIONSArticle Views421Altmetric-Citations181LEARN ABOUT THESE METRICSArticle Views are the COUNTER-compliant sum of full text article downloads since November 2008 (both PDF and HTML) across all institutions and individuals. These metrics are regularly updated to reflect usage leading up to the last few days.Citations are the number of other articles citing this article, calculated by Crossref and updated daily. Find more information about Crossref citation counts.The Altmetric Attention Score is a quantitative measure of the attention that a research article has received online. Clicking on the donut icon will load a page at altmetric.com with additional details about the score and the social media presence for the given article. Find more information on the Altmetric Attention Score and how the score is calculated. Share Add toView InAdd Full Text with ReferenceAdd Description ExportRISCitationCitation and abstractCitation and referencesMore Options Share onFacebookTwitterWechatLinked InReddit PDF (905 KB) Get e-Alerts Get e-Alerts