The medical community as a whole is aware that patients show variable responses to drug treatments but physicians currently lack efficient tools to predict these drug responses and to select appropriate treatments. Coupling our knowledge of genetic polymorphisms with clinical-response data promises a bright future for rapid advances in personalized medicine.
To study the late beta-cell-specific function of the homeodomain protein IPF1/PDX1 we have generated mice in which the Ipf1/Pdx1 gene has been disrupted specifically in beta cells. These mice develop diabetes with age, and we show that IPF1/PDX1 is required for maintaining the beta cell identity by positively regulating insulin and islet amyloid polypeptide expression and by repressing glucagon expression. We also provide evidence that IPF1/PDX1 regulates the expression of Glut2 in a dosage-dependent manner suggesting that lowered IPF1/PDX1 activity may contribute to the development of type II diabetes by causing impaired expression of both Glut2 and insulin.
In this study 87 amino acids (AA.s) have been characterized by 26 physicochemical descriptor variables. These descriptor variables include experimentally determined retention values in seven thin-layer chromatography (TLC) systems, three nuclear magnetic resonance (NMR) shift variables, and 16 calculated variables, namely six semiempirical molecular orbital indices, total, polar, and nonpolar surface area, van der Waals volume of the side chain, log P, molecular weight, and four indicator variables describing hydrogen bond donor and acceptor properties, and side chain charge. In the present study, the data from a previous characterization of 55 AA.s from our laboratory have been extended with data for 32 additional AA.s and 14 new descriptor variables. The new 32 AA.s were selected to represent both intermediate and more extreme physicochemical properties, compared to the 20 coded AA.s. The new extended and updated principal property scales, the z-scales, were calculated and aligned to previously reported z(old)-scales. The appropriateness of the extended z-scales were validated by the use in quantitative sequence-activity modeling (QSAM) of 89 elastase substrate analogues and in a QSAM of 29 neurotensin analogues.
Twenty nucleosides occurring in transfer ribonucleic acid (tRNA) have been characterized using 21 experimentally determined (HPLC, TLC, NMR, etc.) and calculated (log P, van der Waals surface area, ionization potential, etc.) variables. Principal component analysis (PCA) was performed on the data set and four statistically significant components or principal properties (PPs) were extracted. The PPs described 68·4% of the variance in the data. The PP values are discussed in terms of similarity and dissimilarity among the nucleosides. The loading vectors from the PCA are used for an interpretation of the nature of the PP vectors. Application of the PPs in sequence-activity modelling is demonstrated with 25 DNA-promoter sequences originating from E. coli. © 1996 by John Wiley & Sons, Ltd.
The gastrointestinal tract is subdivided into regions with different roles in digestion and absorption. How this patterning is established is unknown. We now report that the pancreatic-duodenal homeobox 1 gene (pdx1) is also expressed in cells of the distal stomach. Positive cells include subpopulations of the three main endocrine (gastrin, somatostatin and serotonin) cell types of this region. Pdx1 deficient mice were virtually devoid of gastrin cells, had normal numbers of somatostatin cells and increased numbers of serotonin cells. Pdx1 is thus important for development of the gastrin cells of the antropyloric mucosa of the stomach and probably acts by controlling the fate of gastrin/serotonin precursor cells.
A multivariate quantitative physicochemical characterization of the five bases adenine (A), cytosine (C), guanine (G), thymine (T) and uracil (U), followed by principal component analysis, shows that the relative dissimilarities between the bases of DNA (A, C, G and T) are almost the same (i.e. balanced). In contrast, mRNA (containing U instead of T) has a considerably larger relative physicochemical similarity between C and U than between all other pairs of bases and is therefore inherently more unbalanced. These results provide a physicochemical explanation of the presence of thymine instead of uracil as an element of DNA. The principal component scores enable a quantitative description of nucleic acid sequence data to be made for structure-activity modelling purposes.
We have previously shown that mice carrying a null mutation in the homeobox gene ipf1, now renamed to pdx1, selectively lack a pancreas. To elucidate the level at which PDX1 is required during the development of the pancreas, we have in this study analyzed the early stages of pancreas ontogeny in PDX-/- mice. These analyses have revealed that the early inductive events leading to the formation of the pancreatic buds and the appearance of the early insulin and glucagon cells occur in the PDX1-deficient embryos. However, the subsequent morphogenesis of the pancreatic epithelium and the progression of differentiation of the endocrine cells are arrested in the pdx1-/- embryos. In contrast, the pancreatic mesenchyme grows and develops, both morphologically and functionally, independently of the epithelium. We also show that the pancreatic epithelium in the pdx1 mutants is unable to respond to the mesenchymal-derived signal(s) which normally promote pancreatic morphogenesis. Together these data provide evidence that PDX-1 acts cell autonomously and that the lack of a pancreas in the pdx1-/- mice is due to a defect in the pancreatic epithelium.
Insulin promoter factor 1 (IPF1), is a homeodomain protein which, in the adult mouse pancreas, is selectively expressed in beta-cells, and which binds to, and transactivates, the insulin promoter via the P1 element. In mouse embryos, IPF1 expression is initiated when the foregut endoderm commits to a pancreatic fate, i.e. prior to both morphogenesis and hormone specific gene expression. At later stages of development the expression is restricted to the dorsal and ventral walls of the primitive foregut at the positions where the pancreases will form. Mice homozygous for a targeted mutation in the Ipf1 gene selectively lack the pancreas. The mutant pups develop to term and are born alive, but die after a few days. The gastrointestinal tract with its associated organs show no obvious malformations. No pancreatic tissue and no ectopic expression of insulin or pancreatic amylase could be detected in this region in mutant neonates or embryos. These findings demonstrate that IPF1 is needed for the formation of the pancreas, and suggest that IPF1 acts to determine the fate of common pancreatic precursor cells and/or to regulate their propagation. The lack of a pancreas in the Ipf1-deficient mutants, the pattern of IPF1 expression and its ability to stimulate insulin gene transcription, strongly suggest that IPF1 functions both in the early specification of the primitive gut to a pancreatic fate and in the maturation of the pancreatic beta-cell.
Multivariate quantitative structure-biodegradability relationships (QSBRs) were developed for a series of 20 halogenated aliphatic hydrocarbons investigated for their microbial biodehalogenation (expressed as half-lives). These QSBRs are based on a battery of quantum-chemical descriptors and the use of the multivariate data analytical techniques principal component analysis (PCA) and partial least-squares projections to latent structures (PLS). The developed models are introduced and discussed from a multivariate data analytical perspective, and much emphasis is placed on the elucidation of the predictive significance of the developed relationships. It is shown that multivariate QSBRs can be formulated for three sets of microorganisms and compounds. These QSBRs can be used to predict biodehalogenation properties for yet untested compounds.
THE mammalian pancreas is a mixed exocrine and endocrine gland that, in most species, arises from ventral and dorsal buds which subsequently merge to form the pancreas. In both mouse and rat the first histological sign of morphogenesis of the dorsal pancreas is a dorsal evagination of the duodenum ai the level of the liver at around the 22-25-somite stage, and shortly thereafter a ventral evagination appears as a derivative of the liver diverticulum(1-3). Low levels of insulin gene transcripts are already present and restricted to the dorsal foregut endoderm at 20 somites, suggesting that pancreas- or insulin gene-specific transcriptional factors are present in this region before the onset of morphogenesis(4). Insulin-promoter-factor 1 (IPF1) is a homeodomain protein which, in the adult mouse pancreas, is selectively expressed in the beta-cells and binds to and transactivates the insulin promoter(5). In mouse embryos, IPF1 expression is restricted to the developing pancreatic anlagen and is initiated when the foregut endoderm is committed to a pancreatic fate(5). We now show that mice homozygous for a targeted mutation in the Ipf1 gene selectively lack a pancreas. The mutant pups survive fetal development but die within a few days after birth. The gastrointestinal part and all other internal organs were normal in appearance. No pancreatic tissue and no ectopic expression of insulin or pancreatic amylase could be detected in mutant embryos and neonates. These findings show that IPF1 is needed for the formation of the pancreas and suggest that it acts to determine the fate of common pancreatic precursor cells and/or to regulate their propagation.
Mycoplasmas are small, cell wall-deficient bacteria. The metabolic regulation of the lipid composition in the membrane of the species Acholeplasma laidlawii, strains A-EF22 and B-JU, is governed mainly by the balance between the potential formation of lamellar and nonlamellar phase structures. However, the regulatory features have not been consistently observed in the B-PG9 strain. A comparison has been performed between the membrane lipid composition for strains A-EF22 and B-PG9, simultaneously changing eight experimental conditions known to affect the regulation and packing properties of the A-EF22 lipids. Multiple regression and partial least-square discriminant analyses of many variables showed: (i) quantitative differences in membrane lipid and protein composition, and in membrane protein molecular masses of the two strains; (ii) different molar fractions of the major polar lipids monoglucosyldiacylglycerol (nonlamellar) and diglucosyldiacylglycerol (lamellar), which were caused by differences in lipid acyl chain length and unsaturation inherent in the strains and by the type of growth medium used; and (iii) similar regulatory mechanisms for changes in the lipid composition under most conditions, responding to the experimentally varied bilayer and nonbilayer properties of the lipid matrix. These regulatory principles are probably valid in other bacteria as well.
Theoretical descriptors for nucleic acids are generated by applying principal component analysis to 3-D field data of non-bonded and charge-charge interaction type. These descriptors are subsequently used to develop quantitative sequence-property models with good predictability for E. coli transcriptional DNA-promoter sequences using the method of partial least squares projections to latent structures (PLS). The resulting 3-D contour maps can be used to investigate requirements for new nucleic acids to be synthesized in order to obtain sequences with altered activities.
A multivariate characterization of the 136 tetra- to octa-chloro-substituted dibenzo-p-dioxins (PCDDs) and dibenzofurans (PCDFs) is reported. By principal component analysis (PCA) of a battery of physicochemical descriptors, an overview of similarities and differences between congener groups and substitution patterns is obtained. Available data for two biochemical end points, the induction of the enzymes aryl hydrocarbon hydroxylase (AHH) and ethoxy resorufin O-deethylase (EROD), were used to construct quantitative structure-activity relationships (QSARs) using partial least-squares projections with latent variables (PLS). By applying the principles of statistical experimental design to the orthogonal scores from the PCA, a set of PCDD and PCDF congeners is proposed for further biological and toxicological evaluations.
The validation of the predictive capability of a quantitative structure-activity relationship (QSAR) is a significant step toward the construction of a reliable model. This point is discussed and illustrated with data for a class of halogenated aliphatic hydrocarbons. For this class of compounds, a QSAR concerning their acute toxicity toward rat was recently published. This QSAR is verified in this study by selecting and testing an external validation set comprising six compounds. The QSAR is also used for predicting the acute toxicity of 38 nontested members of this class.
Biopolymer sequences (e.g., DNA, RNA, proteins and polysaccharides) and chemical processes (e.g., a batch or continuous polymer synthesis run in a chemical plant) have close similarities from the modelling point of view. When a set of sequences or processes is characterized by multivariate data, a three-way data matrix is obtained. With sequences the position and with processes the time is one direction in this matrix. The multivariate modelling of this matrix by principal component analysis (PCA) or partial least-squares (PLS) methods for the following purposes is discussed: classification of sequences; quantitative relationships between sequence and biological activity or chemical properties; optimizing a sequence with respect to selected properties; process diagnostics; and quantitative relationships between process variables and product quality variables. To obtain good models, a number of problems have to be adequately dealt with: appropriate characterization of the sequence or process; experimental design (selecting sequences or process settings); transforming the three-way into a two-way matrix; and appropriate modelling and validation (modelling interactions, periodicities, ''time series'' structures and ''neighbour effects''). A multivariate approach to sequence and process modelling using PCA and PLS projections to latent structures is discussed and illustrated with several sets of peptide and DNA promoter data.
Models have been developed that allow the biological activity of a DNA segment to be altered in a desired direction. Partial least squares projections to latent structures (PLS) was used to establish a quantitative model between a numerical description of 68 bp fragments of 25 E.coli promoters and their corresponding quantitative measure of in vivo strength. This quantitative sequence-activity model (QSAM) was used to generate two 68 bp fragments predicted to be more potent promoters than any of those on which the model originally was based. The optimized structures were experimentally verified to be strong promoters in vivo.
The effects of combined toxicity were studied, using marine periphyton communities exposed to mixtures of tri-n-butyl tin (TBT) and diuron (3-(3,4-dichlorophenyl)-1,1-dimethylurea, DCMU) in indoor aquaria during four weeks. The experimental design of the study followed a central composite design (CCD) and utilized dose-response surface methodology for evaluation of the results. The detection of pollution-induced community tolerance (PICT) was accomplished by short-term (1 h) tests on inhibition of photosynthesis. Both single-toxicant and two-toxicant short-term tests were used. Two tentative measures of tolerance are proposed to achieve convenient comparisons of the tolerances from the two-toxicant tests. With the detection of PICT, effects of the long-term exposure were recorded on diatom species richness, chlorophyll a accumulation and copepod abundance. The decrease of diatom species richness was accompanied by an increased tolerance (PICT), which was detectable by all tolerance measures used. Primary effects on microalgae were recorded as a decrease in chlorophyll a at higher toxicant concentrations. At lower concentrations, primary effects on copepods were found, which resulted in reduced grazing and increased chlorophyll a content.
A strategy for the systematic analysis and priority ranking of environmental chemicals has been applied to a class of 58 halogenated aliphatic hydrocarbons. A training set of ten compounds representing this class, was selected by statistical design. The training set compounds were then subjected to biological testing in the Salmonella typhimurium reverse mutation assay (Ames test). The measured biological data, recorded as dose-response curves, were analyzed to determine the mutagenic potency (slope of the initial portion) and the mutagen dose (MD 50) required to increase the number of revertants above the background by 50%. For each compound, four mutagenic potency estimates and four MD 50 values were determined, all originating from the tester strains TA 100 and TA 1535 with and without metabolic activation. The obtained responses were analyzed with multivariate techniques to give QSAR models relating the mutagenic potency data to the physico-chemical properties of the compounds. Finally, the derived QSARs were used to predict the mutagenic potencies and the MD 50S for the non-tested compounds in the class.
A new way to represent and analyze DNA sequence data is described. This approach complements methods currently used, in that it allows the systematic part of the variation between different sequences to be modeled. This can prove as informative as absence of variation (homology), which is the most widely used criterion for comparing sequence data. A multivariate sequence-activity model (SAM), for DNA-promoter sequences is presented, by which the relative promoter strength is modeled in terms of the primary DNA-sequence. The model is shown to have a good predictive capability. The coefficients from the model are interpreted, and used to design new structures predicted to be strong promoters in the system investigated. The approach described is also applicable to other kinds of sequence data, e.g. RNAs, proteins or peptides.