The ionization state of drugs influences many pharmaceutical properties such as their solubility, permeability, and biological activity. It is therefore important to understand the structure property relationship for the acid-base dissociation constant pKa during the lead optimization process to make better-informed design decisions. Computational approaches, such as implemented in MoKa, can help with this; however, they often predict with too large error especially for proprietary compounds. In this contribution, we look at how retraining helps to greatly improve prediction error. Using a longitudinal study with data measured over 15 years in a drug discovery environment, we assess the impact of model training on prediction accuracy and look at model degradation over time. Using the MoKa software, we will demonstrate that regular retraining is required to address changes in chemical space leading to model degradation over six to nine months.
This analysis elucidates the impact of small molecule architecture on common in vitro ADME assays. In vitro parameters considered in this analysis included Caco-2 permeability/efflux, CYP3A4 inhibition, hERG inhibition, and rat microsomal extraction ratio (ER). The statistical significance and practical meaningfulness of chirality were determined by comparison of the distribution of enantiomers with the experimental variation distribution observed from duplicate measurements. Statistical tools were applied to characterize the role of molecular architecture on the outcome of a given in vitro assay. We found that CYP3A4 inhibition, hERG inhibition, Caco-2 permeability, and efflux are unlikely to be modulated by chirality. However, rat microsomal ER provides a statistically significant, and quantitatively meaningful, chance of being influenced by chirality.
Mass spectrometry in three dimensions (MS3D) is a newly developed method for the determination of protein structures involving intramolecular chemical crosslinking of proteins, proteolytic digestion of the resulting adducts, identification of crosslinks by mass spectrometry (MS), peak assignment using theoretical mass lists, and computational reduction of crosslinks to a structure by distance geometry methods. To facilitate the unambiguous identification of crosslinked peptides from proteolytic digestion mixtures of crosslinked proteins by MS, we introduced double 18O isotopic labels into the crosslinking reagent to provide the crosslinked peptides with a characteristic isotope pattern. The presence of doublets separated by 4 Da in the mass spectra of these materials allowed ready discrimination between crosslinked and modified peptides, and uncrosslinked peptides using automated intelligent data acquisition (IDA) of MS/MS data. This should allow ready automation of the method for application to whole expressible proteomes.
Most preclinical leads exhibit poor ADME/PK properties and require optimizing to increase the likelihood of becoming successful pharmaceuticals. As a means of accelerating the evaluation of these leads in vivo, we assessed the use of LC/MS with the chemiluminescent-nitrogen detector (CLND) and a stable isotope to identify and quantify in vivo metabolites and to measure excretion. A 14C-labeled preclinical lead that also contained two chlorine atoms was administered orally to rats, and samples of bile, urine, and plasma were collected and analyzed by LC with radiodetection and by LC/MS-CLND with the chlorine atoms used as tracers. Both methods identified seven metabolites in bile and two metabolites in urine. The amount and abundance of each metabolite was measured, and the results were equivalent for the two methods. Material balance was measured by liquid scintillation counting of the starting samples, by LC/radiodetection, and by LC/MS-CLND. All three methods yielded the same results and showed that the primary route of clearance was metabolism followed by immediate excretion. This study demonstrates that LC/MS-CLND with a stable isotope is a method that can efficiently track and accurately quantify metabolites, making it possible to rapidly study ADME/PK in vivo without radiolabeling.
This paper presented the development of an automated HPLC small-scale purification method for single bead compounds derived from combinatorial libraries. The method was found to produce higher and more consistent recoveries of purified compounds as compared to conventional manual HPLC purification. Using the manual method, the average percentage recovery of one synthetic compound was determined to be 24% and the coefficients of variation (C.V.%) of recovery were found to be greater than 38%. Using the automated system, the average percentages recovery of a standard compound at 600 and 1000 micromol l(-1) were determined to be 72.63+/-10.17% and 81.34+/-4.39%, respectively. This represented an approximate 3-folds increase in percentage recovery compared to that of the manual small-scale purification process. It was also found that the C.V.% of recovery were less than 15% at both concentration levels. The development of this automated method was found to be straightforward. The importance and implications of this study were discussed.
We have used intramolecular cross-linking, MS, and sequence threading to rapidly identify the fold of a model protein, bovine basic fibroblast growth factor (FGF)-2. Its tertiary structure was probed with a lysine-specific cross-linking agent, bis(sulfosuccinimidyl) suberate (BS(3)). Sites of cross-linking were determined by tryptic peptide mapping by using time-of-flight MS. Eighteen unique intramolecular lysine (Lys-Lys) cross-links were identified. The assignments for eight cross-linked peptides were confirmed by using post source decay MS. The interatomic distance constraints were all consistent with the tertiary structure of FGF-2. These relatively few constraints, in conjunction with threading, correctly identified FGF-2 as a member of the beta-trefoil fold family. To further demonstrate utility, we used the top-scoring homolog, IL-1beta, to build an FGF-2 homology model with a backbone error of 4.8 A (rms deviation). This method is fast, is general, uses small amounts of material, and is amenable to automation.
In the ongoing quest for ever more rapid techniques to quantify small organic molecules, we have evaluated a chemiluminescent nitrogen detector (CLND) as a universal quantitation tool for nitrogen-containing molecules. By how injection analysis (FIA) and in conjunction with reversed-phase (RP) chromatography using gradient elution, the CLND produced a linear response from 25 to 6400 pmol of nitrogen that was equivalent for a set of chemically and structurally diverse compounds. Over the entire linear range, the absolute response exhibited an average error of approximately +/-10% among the compounds. In addition, the response was independent of mobile-phase composition. These results demonstrate that the CLND can be used with FIA or on-line with RP-HPLC for rapid and accurate quantitation down to low-picomole levels, using a single external standard. We also used the CLND in combination with a UV detector and a mass spectrometer (MS) during RP-HPLC (LC/UV/N/MS) to characterize several samples containing small organic compounds synthesized by both standard and combinatorial methods, The identity, quantity, and purity of compounds of interest:were assessed from a single HPLC injection of each sample. These results show this technique (LC/UV/N/MS) to be a widely applicable, generic method for the pharmaceutical industry to rapidly identify, quantify, and determine the purity of small organic compounds.
The screening of diverse libraries of small molecules created by combinatorial synthetic methods is a recent development which has the potential to accelerate the identification of lead compounds in drug discovery. We have developed a direct and rapid method to identify lead compounds in libraries involving affinity selection and mass spectrometry. In our strategy, the receptor or target molecule of interest is used to isolate the active components from the library physically, followed by direct structural identification of the active compounds bound to the target molecule by mass spectrometry. In a drug design strategy, structurally diverse libraries can be used for the initial identification of lead compounds. Once lead compounds have been identified, libraries containing compounds chemically similar to the lead compound can be generated and used to optimize the binding characteristics. These strategies have also been adopted for more detailed studies of protein–ligand interactions.
Most biopolymer molecules are much smaller than the wavelength of light used in classical light-scattering experiments (ca. 500 nm), and thus the simple Rayleigh equation and a 90° light-scattering photometer are sufficient to determine their molecular weight. In combination with high-performance liquid chromatography (HPLC), it is demonstrated that a simple HPLC fluorimeter can be used as a 90° light-scattering detector for biopolymer molecular weight determinations. To simplify data handling, only relative molecular weights are measured. Three mathematical assumptions are adopted, and their validity for proteins is shown. To place the work in perspective, the relative advantages and limitations of this 90° light-scattering detector are compared with the more commonly used low-angle laser light-scattering detector. Two examples of protein molecular weight determinations are given to illustrate the broad utility of 90° classical light scattering to the study of biopolymer structures and interactions.
An high-performance liquid chromatographic method has been developed which simultaneously determines three critical physical properties of polyethylene glycol (PEG)-modified proteins: molecular size, polymer distribution and weight composition. With both UV and refractive index (RI) detectors in series, size-exclusion chromatography (SEC) is used to separate the PEG-protein species according to size. The size analysis of these PEG-proteins is predicted to be accurately calibrated with the viscosity radius (universal calibration), which compensates for the shape differences between PEG and protein structures. The heterogeneity of the PEG-protein grafted copolymer is represented by the polymeric term “polydispersity”, which describes the size distribution. Separate SEC calibrations of the PEG and the protein used for conjugation allow a determination of the weight composition of the PEG-protein (weight PEG/weight protein) by combining UV and RI chromatograms of a PEG-protein sample. This compositional analysis is validated through independent and direct measurement of the PEG on a PEG-protein via acid hydrolysis and quantitative SEC. Comparisons of compositional analysis of PEG-protein with sodium dodecyl sulfate polyacrylamide gel electrophoresis densitometry demonstrate that gel analysis of some proteins is misleading.