Dear Editor, The identification of target combinations with synergistic effects on cancer is at the leading edge of modern cancer research, especially for the development of combined anticancer therapies.1 However, at present, the basis for selection of beneficial target combinations commonly relies on expert opinion without any systematic rationale. The development of high-throughput technologies has led to the availability of large-scale clinical gene expression data sets.2, 3, 4 Mining of these data sets for identification of gene combinations with synergetic effects on survival outcome in cancer could provide a systematic rationale for the identification of target combinations with potential therapeutic synergy. Multiple online tools have been developed recently to assess the relationship between expression of a single gene and clinical outcome across a variety of cancers.5, 6 Here we present SynTarget, the first tool able to test the cumulative effect of two genes on survival outcome and, therefore, can identify gene pairs with synergistic effects. At present, SynTarget is based on 15 large-scale gene expression data sets covering eight different cancers with the possibility to select clinically important subtypes (i.e., triple-negative, p53 mutated cancers, K-ras mutated cancers etc.). On submission, the user selects a specific cancer data set and subtype and specifies two genes. Patients in the selected data set (subtype) are split into four groups with respect to the expression of the specified genes (high−high, high−low, low−high and low−low). Next, each group is tested versus other samples to find any statistical differences in survival outcome (i.e., high−high versus others, high−low versus others and so on). This information is accompanied by individual-gene survival effects in order to understand the degree of gene synergy. Among other drug classes, immunotherapeutic agents have enormous potential for synergistic combinations.1 Triple-negative breast cancer is the subtype with the worst prognosis among all breast cancer subtypes, with currently no known molecular targets.7 We used SynTarget to search cell surface genes with the synergistic potential on survival, whose high expression leads to significantly negative prognosis. For example, ADAM9 is a membrane-anchored protein and has been implicated in a variety of biological processes, as well as being involved in cancer metastasis. RC3H2 is a membrane-associated nucleic acid-binding protein. High expression of both genes individually was slightly negatively associated (P-values ~0.07 and 0.04) with survival in triple-negative patients from the METABRIC data set.3 The subgroup of triple-negative patients where both genes are highly expressed has a significantly negative shift (P-value<6e−05) in survival, in comparison with other patients (see Supplementary Materials for details). Therefore, SynTarget provides statistical evidence that high expression of both RC3H2 and ADAM9 synergistically affects survival of patients with triple-negative breast cancer. In summary, SynTarget supports the need of biomedical researchers to estimate the synergy of gene expression on survival of cancer patients. To our knowledge this is a first tool of this kind and, as shown by several examples (see Supplementary Materials), SynTarget can be used for fast validation of the clinical synergy for two genes. SynTarget is incorporated into BioProfiling.de, an analytical portal for high-throughput cell biology,8 and is freely available at http://www.bioprofiling.de/synergy2G. The authors declare no conflict of interest. This work was supported by the UK Medical Research Council (MRC) and fundamental research program of the Russian State Academies of Sciences. Supplementary Information accompanies this paper on Cell Death and Differentiation website
p53MutaGene is the first online tool for statistical validation of hypotheses regarding the effect of p53 mutational status on gene regulation in cancer. This tool is based on several large-scale clinical gene expression data sets and currently covers breast, colon and lung cancers. The tool detects differential co-expression patterns in expression data between p53 mutated versus p53 normal samples for the user-specified genes. Statistically significant differential co-expression for a gene pair is indicative that regulation of two genes is sensitive to the presence of p53 mutations. p53MutaGene can be used in 'single mode' where the user can test a specific pair of genes or in 'discovery mode' designed for analysis of several genes. Using several examples, we demonstrate that p53MutaGene is a useful tool for fast statistical validation in clinical data of p53-dependent gene regulation patterns. The tool is freely available at http://www.bioprofiling.de/tp53.
Motivation: Protein-protein and protein-ligands interactions play a central role in biochemical reactions, and understanding these processes is an important step in several fields of biomedical science and drug discovery.Results: Our research is conducted on a number of protein-protein interactions. We attempt in this report to show interactive links between virtual and experimental approaches in total pipeline "From gene to drug and using modem SPR technology (optical biosensor) for assessing the strengths of protein-protein and protein-ligand interactions.Availability: Preprint of this paper is available on request from the authors.
In the present study proteomes of liver samples were analyzed after administration of phenobarbital (PB) or 3-methylcholantrene (3-MC) to mice. Liver cell homogenates were subfractionated by differential ultracentrifugation into cytosol and microsomes, which were subjected to 2-DE to generate the proteomic maps of these fractions. 2-DE yielded 1100 and 800 protein spots for microsomes and cytosol, respectively. General trends of the fraction-specific alterations after 3-MC or PB treatment were evaluated using the Student's t-test and the principal component analysis (PCA). According to the PCA-derived data, the microsomal changes after 3-MC and PB treatment were quite similar. However, in the case of the cytosol data, the specificities of 3-MC- and PB-induced responses could be clearly distinguished from each other. Protein spots, whose expression levels differed from control, were identified by MALDI-TOF PMF. Proteomic studies such as those reported herein can be useful in identifying the molecular-based toxicity of lead drug candidates.
Proteomics of membrane-bound proteins requires new technological approaches powered by the appropriate bioinformatic tools. Off-line molecular scanning is one of such approaches, that fully exploits the capacities provided by modern instruments and robotized lines. The SDS-PAGE separated liver microsomes are scanned by picking up the 1.5 mm spots closely one after another. Each picked spot is digested and peptide mass fingerprint is collected. The series of fingerprints is then processed to assemble the consistent 1D proteomic map.
Cytochrome P450 Knowledgebase (CPK) is the formalized data repository containing the information about both the structural and functional features of the enzymes of this superfamily. The protein structure is represented by its sequence (either DNA or amino acid one) and, also, by 3D structure when available. The functional description is provided through specifying the list of substrates, inducers and inhibitors for each enzymatic form. Such an approach enables to capture the relevant pieces of information even from the text of the abstract. Having collected enough data it is possible to produce simple SAR models for prediction of the P450s most likely interacting with the given substance. The access to CPK is publicly available on the Internet at http*/ /cl2d.ibmh.msk.su (http://195.178.207.138).
For the past few decades, cytochromes P450 (CYPs) have been the subject of extensive research, owing to the ability of these enzymes to serve as drug targets as well as to their active participation in drug metabolism. Other varieties and functions of CYPs have been discovered and this superfamily currently comprises over 2000 different protein species. In the present study, the protein sequences of CYPs were submitted to computer analysis for elucidation of the structural basis of their pronounced functional diversity. The basic local alignment search tool (BLAST) was used to demonstrate that CYP protein sequences share a certain general similarity; at the same time, it was shown that the CYP superfamily may be split into a number of groups of intimately related proteins. These groups, the families, were revealed by means of cluster analysis, which demonstrated a strong hierarchy among the animal, bacterial and fungal P450s, and the lack of such a hierarchy for plant enzymes. Multiple alignment and consensus sequence analysis were the approaches taken to find out which structural peculiarities of P450s are responsible for the deviations from the random picture. Proteins within each family were aligned and collapsed to the corresponding consensus sequences, the alignment of which produced the consensus for the whole superfamily. The latter consensus yielded a number of unity motifs (most of which being related to the heme-fixing assembly), while the cross-family comparison of consensus sequences enabled the retrieval of some diversity motifs. Three consensus sequences (for the CYP51 and CYP2 families and for the superfamily) were compared, to line up the unity and diversity motifs with the appropriate X-ray data.
The prominent diversity of cytochromes P450 requires the automatic means for their annotation. To tackle the annotation problem, we propose the motif-based approach. The core idea-is to distribute the protein sequences among a number of clusters in such a way that the members of each cluster would share the similar structural-functional motifs. Thus we sequentially performed: (1) cluster analysis, (2) multiple alignment of sequences-within each cluster; (3) construction of consensus for each cluster; (4) unraveling of the motifs within the consensus. Out of all the possible clusters, only those are selected which share the larger number of statistically distinguishable motifs. From our studies we conclude that there is no unified definition of the P450 family in terms of sequence similarity because family clusters arise at different cutoffs of agglomeration. All the motifs revealed could be divided into two major groups: the first group involves the unity motifs that occur in every cytochrome P450, while the second involves the diversity motifs that vary across the families.
Cytochromes P450 (P450s) form the wide superfamily of heme-containing enzymes. They play a key role in oxidative modification of various xenobiotics and endogenous compounds in majority of living organisms. Modem genomic researches significantly increase the number of known P450s sequences (up to 3000), while only 13 crystal structures of P450s are available from PDB (on 05.2003). The difference between the number of known sequences and crystallized structures will increase, as the majority of P450s are membrane bound proteins, which are extremely difficult to crystallize. Deficiency of known structures forces the researchers to use computer 3D modelling. Paper is devoted to basic approaches, problems and prospects of homology modelling of D450s, as well as for criteria of successful modelling, model refinement and verification of its correctness and reliability.
The approach for the automated systematic collection of information on the structure and function of different P450 isozymes constitutes the core of the Cytochrome P450 Knowledgebase. First, the BLAST is routinely used to scan the SwissProt/TrEMBL for the presence of new P450s. On getting this new one, the pair-wise alignment is used for the rough annotation of the new sequence, and currently over 2000 primary structures of cytochromes P450 are available. The second task is to extract the knowledge about the P450 functioning from the published papers. For this purpose we have downloaded the full collection of the PubMed abstracts, containing the term "cytochrome P450". Downloaded texts were sorted by their relevance to the topic of P450 functioning. While arranging the abstracts we have accumulated the vocabularies of special terms, and currently the vocabulary of names of low-molecular weight compounds encounters about 10 000 unique entries. Combining the created vocabularies with the semantic patterns we have programmed the algorithm for converting the abstract into a number of facts. The fact enters the database after being confirmed by an expert.
Article Common motifs in microsomal cytochrome P450 N-terminal membrane fragments was published on July 1, 1992 in the journal Journal of Basic and Clinical Physiology and Pharmacology (volume 3, issue Supplement).
Alexei Kopylov合作论文数Computer Science Department, Cornell University1
Kirill Degtyarenko合作论文数European Bioinformatics Institute|Wellcome Trust Genome Campus1