Grass glucuronoarabinoxylan (GAX) substitutions can inhibit enzymatic degradation and are involved in the interaction of xylan with cell wall cellulose and lignin, factors which contribute to the recalcitrance of biomass to saccharification. Therefore, identification of xylan characteristics central to biomass biorefining improvement is essential. However, the task of assessing biomass quality is complicated and is often hindered by the lack of a reference for a given crop. In this study, we created a reference library, expressed in glucose units, of Miscanthus sinensis GAX stem and leaf oligosaccharides, using DNA sequencer-Assisted Saccharide analysis in high throughput (DASH), supported by liquid chromatography (LC), nuclear magnetic resonance (NMR) spectroscopy and mass spectrometry (MS). Our analysis of a number of grass species highlighted variations in substitution type and frequency of stem and leaf GAX. In miscanthus, for example, the β-Xylp-(1 → 2)-α-Araf-(1 → 3) side chain is more abundant in leaf than stem. The reference library allows fast identification and comparison of GAX structures from different plants and tissues. Ultimately, this reference library can be used in directing biomass selection and improving biorefining.
The need for standard operating protocols (SOPs) has long been recognized in all branches of analytical chemistry and is especially useful in transferring a given assay from one laboratory to another and for the cross-comparison of results. However, the field of standardized protocols has received renewed interest following the completion of the human genome project and the birth of functional genomics. The use of analytical equipment such as NMR spectroscopy and mass spectrometry in metabolomics, proteomics, and related functional genomic approaches has led to an increased need to standardize the reporting of data acquisition so that results in one laboratory can be validated in another. Ultimately databases can be produced of experimental data that catalog a tier of cellular organization, and in this manner a systems biology description of the biological world is approached. Given that any true description of the proteome or metabolome must by very definition consider all the changes that occur to these dynamic systems, it becomes clear that a full description will only be achieved by community-led initiatives, and thus standardized protocols become a vital cornerstone of any such endeavor. This article surveys recent developments in the area of standard reporting of protocols in metabolomics. However, to do this, it is first necessary to set the context of standardization of protocols in biology in general and the field of functional genomics in particular.
Background Plant cell wall polysaccharide composition varies substantially between species, organs and genotypes. Knowledge of the structure and composition of these polysaccharides, accompanied by a suite of well characterised glycosyl hydrolases will be important for the success of lignocellulosic biofuels. Current methods used to characterise enzymatically released plant oligosaccharides are relatively slow. Results A method and software was developed allowing the use of a DNA sequencer to profile oligosaccharides derived from plant cell wall polysaccharides (DNA sequencer-Assisted Saccharide analysis in High throughput, DASH). An ABI 3730xl, which can analyse 96 samples simultaneously by capillary electrophoresis, was used to separate fluorophore derivatised reducing mono- and oligo-saccharides from plant cell walls. Using electrophoresis mobility markers, oligosaccharide mobilities were standardised between experiments to enable reproducible oligosaccharide identification. These mobility markers can be flexibly designed to span the mobilities of oligosaccharides under investigation, and they have a fluorescence emission that is distinct from that of the saccharide labelling. Methods for relative and absolute quantitation of oligosaccharides are described. Analysis of a large number of samples is facilitated by the DASHboard software which was developed in parallel. Use of this method was exemplified by comparing xylan structure and content in Arabidopsis thaliana mutants affected in xylan synthesis. The product profiles of specific xylanases were also compared in order to identify enzymes with unusual oligosaccharide products. Conclusions The DASH method and DASHboard software can be used to carry out large-scale analyses of the compositional variation of plant cell walls and biomass, to compare plants with mutations in plant cell wall synthesis pathways, and to characterise novel carbohydrate active enzymes.
The Golgi apparatus is the central organelle in the secretory pathway and plays key roles in glycosylation, protein sorting, and secretion in plants. Enzymes involved in the biosynthesis of complex polysaccharides, glycoproteins, and glycolipids are located in this organelle, but the majority of them remain uncharacterized. Here, we studied the Arabidopsis (Arabidopsis thaliana) membrane proteome with a focus on the Golgi apparatus using localization of organelle proteins by isotope tagging. By applying multivariate data analysis to a combined data set of two new and two previously published localization of organelle proteins by isotope tagging experiments, we identified the subcellular localization of 1,110 proteins with high confidence. These include 197 Golgi apparatus proteins, 79 of which have not been localized previously by a high-confidence method, as well as the localization of 304 endoplasmic reticulum and 208 plasma membrane proteins. Comparison of the hydrophobic domains of the localized proteins showed that the single-span transmembrane domains have unique properties in each organelle. Many of the novel Golgi-localized proteins belong to uncharacterized protein families. Structure-based homology analysis identified 12 putative Golgi glycosyltransferase (GT) families that have no functionally characterized members and, therefore, are not yet assigned to a Carbohydrate-Active Enzymes database GT family. The substantial numbers of these putative GTs lead us to estimate that the true number of plant Golgi GTs might be one-third above those currently annotated. Other newly identified proteins are likely to be involved in the transport and interconversion of nucleotide sugar substrates as well as polysaccharide and protein modification.
Nuclear magnetic resonance spectroscopy signals are modelled as a sum of decaying complex exponentials in noise. The spectral analysis of these signals allowing for their decomposition and the estimation of the parameters of the components is crucial to the study of biochemical samples. This paper presents a novel Gabor filterbank/notch filtering instantaneous frequency (IF) estimator, that enables the extraction of weaker and shorter lived exponentials. This new approach is an iterative procedure where a Gabor filterbank is first employed to obtain a reliable estimate of the IF of the strongest component present. The estimated strongest component is then notch filtered, which un-masks weaker components, and the procedure repeated. The performance of this method was evaluated using an artificial signal and compared to the short time Fourier transform, reassigned STFT, and the original Gabor filterbank approach. The results clearly demonstrate its superiority in uncovering weaker signals and resolving components that are very close to one another in frequency. Furthermore, the new method is shown to be more robust than the ITCMP technique at low signal to noise ratios.
High-resolution (1)H NMR spectroscopy is frequently used in the field of metabolomics to assess the metabolites found in biofluids or tissue extracts to define a metabolic profile that describes a given biological process. In this study, we aimed to increase the utility of NMR-based metabolomics by using advanced Bayesian modeling of the time-domain high-resolution 1D NMR free induction decay (FID). The improvement over traditional nonparametric binning is twofold and associated with enhanced resolution of the analysis and automation of the signal processing stage. The automation is achieved by using a Bayesian formalism for all parameters of the model including the number of components. The approach is illustrated with a study of early markers of acute exposure to different doses of a well-characterized nongenotoxic hepatocarcinogen, phenobarbital, in rats. The results demonstrate that Bayesian deconvolution produces a better model for the NMR spectra that allows the identification of subtle changes in metabolic concentrations and a decrease in the expected false discovery rate compared with approaches based on "binning". These properties suggest that Bayesian deconvolution could facilitate the biomarker discovery process and improve information extraction from high-resolution NMR spectra.
2‐DE is an important tool in quantitative proteomics. Here, we compare the deep purple (DP) system with DIGE using both a traditional and the SameSpots approach to gel analysis. Missing values in the traditional approach were found to be a significant issue for both systems. SameSpots attempts to address the missing value problem. SameSpots was found to increase the proportion of low volume data for DP but not for DIGE. For all the analysis methods applied in this study, the assumptions of parametric tests were met. Analysis of the same images gave significantly lower noise with SameSpots (over traditional) for DP, but no difference for DIGE. We propose that SameSpots gave lower noise with DP due to the stabilisation of the spot area by the common spot outline, but this was not seen with DIGE due to the co‐detection process which stabilises the area selected. For studies where measurement of small abundance changes is required, a cost–benefit analysis highlights that DIGE was significantly cheaper regardless of the analysis methods. For studies analysing large changes, DP with SameSpots could be an effective alternative to DIGE but this will be dependent on the biological noise of the system under investigation.
Muscle degeneration in the heart of 1-9 month-old mdx mice (a model for Duchenne muscular dystrophy) has been monitored using metabolomic and proteomic approaches. In both data sets, a pronounced aging trend was detected in control and mdx mice, and this trend was separate from the disease process. In addition, the characteristic increase in taurine associated with dystrophic tissue is correlated with proteins associated with oxidative phosphorylation and mitochondrial metabolism.
Type 2 diabetes mellitus is the result of a combination of impaired insulin secretion with reduced insulin sensitivity of target tissues. There are an estimated 150 million affected individuals worldwide, of whom a large proportion remains undiagnosed because of a lack of specific symptoms early in this disorder and inadequate diagnostics. In this study, NMR-based metabolomic analysis in conjunction with multivariate statistics was applied to examine the urinary metabolic changes in two rodent models of type 2 diabetes mellitus as well as unmedicated human sufferers. The db/db mouse and obese Zucker (fa/fa) rat have autosomal recessive defects in the leptin receptor gene, causing type 2 diabetes. 1H-NMR spectra of urine were used in conjunction with uni- and multivariate statistics to identify disease-related metabolic changes in these two animal models and human sufferers. This study demonstrates metabolic similarities between the three species examined, including metabolic responses associated with general systemic stress, changes in the TCA cycle, and perturbations in nucleotide metabolism and in methylamine metabolism. All three species demonstrated profound changes in nucleotide metabolism, including that of N-methylnicotinamide and N-methyl-2-pyridone-5-carboxamide, which may provide unique biomarkers for following type 2 diabetes mellitus progression.
The problem of model detection and parameter estimation for noisy signals arises in different areas of science and engineering including audio processing, seismology, electrical engineering, and NMR spectroscopy. We have adopted the Bayesian modeling framework to jointly detect and estimate signal resonances. This considers a model of the time-domain complex free induction decay (FID) signal as a sum of exponentially damped sinusoidal components. The number of model components and component parameters are considered unknown random variables to be estimated. A Reversible Jump Markov Chain Monte Carlo technique is used to draw samples from the joint posterior distribution on the subspaces of different dimensions. The proposed algorithm has been tested on synthetic data, the (1)H NMR FID of a standard of L-glutamic acid and a blood plasma sample. The detection and estimation performance is compared with Akaike information criterion (AIC), minimum description length (MDL) and the matrix pencil method. The results show the Bayesian algorithm superior in performance especially in difficult cases of detecting low-amplitude and strongly overlapping resonances in noisy signals.
In this article we present the activities of the Ontology Working Group (OWG) under the Metabolomics Standards Initiative (MSI) umbrella. Our endeavour aims to synergise the work of several communities, where independent activities are underway to develop terminologies and databases for metabolomics investigations. We have joined forces to rise to the challenges associated with interpreting and integrating experimental process and data across disparate sources (software and databases, private and public). Our focus is to support the activities of the other MSI working groups by developing a common semantic framework to enable metabolomics-user communities to consistently annotate the experimental process and to enable meaningful exchange of datasets. Our work is accessible via a public webpage and a draft ontology has been posted under the Open Biological Ontology umbrella. At the very outset, we have agreed to minimize duplications across omics domains through extensive liaisons with other communities under the OBO Foundry. This is work in progress and we welcome new participants willing to volunteer their time and expertise to this open effort.
Using an NMR based approach, employing both solution state and high resolution magic angle spinning (HR MAS) 1 H NMR spectroscopy, in conjunction with an array of statistical methods, we report cerebral metabolic deficits in a mouse model of Batten disease ( Cln3 null mutant mice). Batten disease is the most common progressive neurodegenerative disorder of childhood and is caused by mutations in the Cln3 gene. In particular, brain tissue from Cln3 mice was characterised by increased concentrations of glutamine, myo-inositol, scyllo-inositol, aspartate and lactate, alongside decreased concentrations of N -acetyl- l -aspartate (NAA), N -acetyl- l -glutamate (NAG), γ-amino butyric acid (GABA), glutamate and creatine. Accompanying changes in lipid deposition were also detected in intact cortical tissue by HR MAS 1 H NMR spectroscopy. To realise the true potential of metabolomic datasets necessitates a comprehensive analysis of the data, such that useful biological information can be extracted and used to generate hypotheses which can be further tested and refined. We found that using a combination of univariate and multivariate analyses, a maximal number of metabolic deficits were successfully identified. In particular the complementary nature of the statistical approaches allowed the definition of changes which were relative, absolute or simply a change in variance, allowing a greater understanding of the disease processes detected.
With the increasing production of metabolomic data there is an awareness of a need for a standardised description of this data to aid assessment, exchange, storage and curation of information from metabolomic studies. In this manuscript the first draft of a minimum requirement for the description of the biological context of a metabolomic study involving mammalian subjects is described. This recommendation has been produced by the Metabolomics Standards Initiative–Mammalian Context Working Sub-Group (MSI-MCWSG) as part of the wider standardisation initiative led by the Metabolomics society. The experiments considered include functional genomic studies, drug toxicology, nutrigenomics, clinical trials, and other human studies. Two reporting requirements are described for pre-clinical (e.g. functional genomics, toxicology) and clinical (e.g. clinical trials, nutrigenomics) studies. It is planned that this will lead to the development of a tool for the description of metabolomic experiments that enables storage, retrieval and manipulation of large amounts of data. This will benefit the assessment and dissemination of metabolomic data from mammalian studies.
The amount of data generated by NMR-based metabolomic experiments is increasing rapidly. Furthermore, diverse techniques increase the need for informative and comprehensive meta-data. These factors present a challenge in the dissemination, interpretation, reviewing and comparison of experimental results using this technology. Thus, there is a strong case for unification and standardisation of the data representation for both academia and industry. Here, a systems analysis of an NMR-based metabolomics experiment is presented in order to reveal the reporting requirements. An in-depth analysis of the NMR component of a metabolomics experiment has been produced, and a first round of data standard development completed. This has focussed on both one- and two-dimensional 1H NMR experiments, but is also applicable to higher dimensions and other nuclei. We also report the modelling of this schema using Unified Modelling Language (UML), and have extended this to a proof-of-concept implementation of the standard as an XML schema.
This short report is discussing a new approach to the enhancement of existing technologies for creating and maintaining ontologies. The approach assumes application of unsupervised neural networks to help domain experts to fit related objects and classes into developing ontology.
The hybrid approaches to knowledge representation provides several important advantages. First, it allows the use of various kinds of expert knowledge inside the intelligent system. Second, it makes it possible to organize an interchange of knowledge between various parts of the intelligent system (including interchange between parts that use connectionist and symbolic paradigms for representation of experts' knowledge). This article describes hybrid systems whose knowledge base contains heterogeneous modules actively coupled by shared memory. Such framework is also called loosely coupled architecture.
This article introduces a framework for interchange of trained neural network models. An XML-based language (Neural Network Markup Language) for the neural network model description is offered. It allows to write down all components of neural network model, which are necessary for its reproduction. We propose to use XML notation for full description of neural models, including data dictionary, properties of training sample, preprocessing methods, details of network structure and parameters, method for network output interpretation.
This article introduces a framework for the interchange of trained neural network models. An XML-based language (Neural Network Markup Language) for the neural network model description is offered. It allows to write down all the components of neural network model which are necessary for its reproduction. We propose to use XML notation for the full description of neural models, including data dictionary, properties of training sample, preprocessing methods, details of network structure and parameters and methods for network output interpretation.
The article introduces a framework for the interchange of trained neural network models. The XML-based language (neural network markup language (NNML)) is presented for the neural network model description, that allows one to write down all components of the neural network model, necessary for its realization. We propose to use the XML notation for full description of neural models, including data dictionary, properties of training sample, pre-processing methods, details of network structure and parameters, method for network output interpretation. The NNML allows interchanging of neural models as well as their documentation, storing and manipulating them independently from the individual simulation system