![IEEE International Conference on Systems Biology : [proceedings]](https://originalfileserver.aminer.cn/sys/aminer/magazine.png)
Rapid advancement of next-generation sequencing (NGS) technologies has facilitated the search for genetic susceptibility factors that influence disease risk in the field of human genetics. In particular whole genome sequencing (WGS) has been used to obtain the most comprehensive genetic variation of an individual and perform detailed evaluation of all genetic variation. To this end, sophisticated methods to accurately call high-quality variants and genotypes simultaneously on a cohort of individuals from raw sequence data are required. On chromosome 22 of 818 WGS data from the Alzheimer's Disease Neuroimaging Initiative (ADNI), which is the largest WGS related to a single disease, we compared two multi-sample variant calling methods for the detection of single nucleotide variants (SNVs) and short insertions and deletions (indels) in WGS: (1) reduce the analysis-ready reads (BAM) file to a manageable size by keeping only essential information for variant calling (“REDUCE”) and (2) call variants individually on each sample and then perform a joint genotyping analysis of the variant files produced for all samples in a cohort (“JOINT”). JOINT identified 515,210 SNVs and 60,042 indels, while REDUCE identified 358,303 SNVs and 52,855 indels. JOINT identified many more SNVs and indels compared to REDUCE. Both methods had concordance rate of 99.60% for SNVs and 99.06% for indels. For SNVs, evaluation with HumanOmni 2.5M genotyping arrays revealed a concordance rate of 99.68% for JOINT and 99.50% for REDUCE. REDUCE needed more computational time and memory compared to JOINT. Our findings indicate that the multi-sample variant calling method using the JOINT process is a promising strategy for the variant detection, which should facilitate our understanding of the underlying pathogenesis of human diseases.
Quality improvement (QI) requires systematic and continuous efforts to enhance healthcare services. A healthcare provider might wish to compare local statistics with those from other institutions in order to identify problems and develop intervention to improve the quality of care. However, the sharing of institution information may be deterred by institutional privacy as publicizing such statistics could lead to embarrassment and even financial damage. In this article, we propose a PRivacy-prEserving Cloud-assisted quality Improvement Service in hEalthcare (PRECISE), which aims at enabling cross-institution comparison of healthcare statistics while protecting privacy. The proposed framework relies on a set of state-of-the-art cryptographic protocols including homomorphic encryption and Yao's garbled circuit schemes. By securely pooling data from different institutions, PRECISE can rank the encrypted statistics to facilitate QI among participating institutes. We conducted experiments using MIMIC II database and demonstrated the feasibility of the proposed PRECISE framework.
Phosphorylation is a post-translational modification process mediated by kinases through the addition of a covalently bound phosphate group, which plays important roles in a wide range of cellular progresses, such as signaling cascades and development. Over the past years, despite many phosphorylation sites have been determined with mass spectrometry techniques, it is not clear which kinase phosphorylates which proteins. Under the circumstance, we propose a new probabilistic model to identify the substrates phosphorylated by certain kinases. Furthermore, we construct three tissue-specific phosphorylation networks based on protein expression data. Investigating the constructed tissue-specific networks, we find they are functionally consistent with the corresponding tissues, implying the effectiveness and biological significance of our proposed approach.
Comprehensive detection and identification of copy number and LOH of chromosomal aberration is required to provide an accurate therapy of human cancer. As a cost-saving and high-throughput tool, SNP arrays facilitate analysis of chromosomal aberration throughout the whole genome. The performance of previous approaches has been limited to several critical issues such as normal cell contamination, aneuploidy and tumor heterogeneity. For these reasons we present a Hidden Markov Model (HMM) based approach called TH-HMM (Tumor Heterogeneity HMM), for simultaneous detection of copy number and LOH in heterogeneous tumor samples using data from Illumina SNP arrays. Through adopting an efficient EM algorithm, our method can correctly detect chromosomal aberration events in tumor subclones. Evaluation on simulated data series indicated that TH-HMM could accurately estimate both normal cell and subclone proportions, and finally recovery the aberration profiles for each clones.
Effects of the relationship between species and environment on an eco-epidemiological system are investigated. And periodic variation is also added to the disharmony parameter. The dynamic behaviors of the system are simulated numerically. A variety of complex population dynamics including stable state, periodic resonance and chaos are obtained. The most important result is that harmony relationship between prey species and environment is benefit for the controlling of disease. Our result reinforces the conjecture that the relationship between species and environment is crucial to transmission of infectious disease.
Brassica rapa L. ssp. chinensis (L.) Hanelt is an important vegetable in eastern Asia. Its leaf area is one of main factor of influencing on the yield of individual plant. To reveal the mechanism of the leaf expansion, the joint segregation analysis of multiple generations (P1, P2, F1, B1, B2 and F2) was used to analyze the genetics of leaf expansion trait in combination SW-13×L-118 of Brassica rapa L. ssp. chinensis (L.) Hanelt. In this paper, the results showed that the leaf expansion trait was controlled by one pair of additive-dominant major genes plus additive-dominant-epistatic polygene (D model). The additive effect of the major gene was -2.3061, the dominant effect was -1.4525. The heritability of major gene of leaf expansion was 2.56% in B1 generation, 1.55% in B2 generation and 12.12% in F2 generation. The heritabilites of the polygene were 54.04% in B1 generation, 25.11% in B2 generation, and 62.60% in F2 generation, respectively. It appeared that the polygene effects should be given prior to the improvement of leaf expansion trait of combination SW-13×L-118, and the selection was effective in later generation, and the negative additive effect and negative dominant effect of major gene were noticed.
Recent researchers suggested Dimethyl sulphide (DMS) flux emission in Arctic Ocean plays an important role for the global warming. A Genetic Algorithm (GA) method was developed and used in calibrating the DMS model parameters in Barents Sea in Arctic Ocean (70-80N, 30-35E). Two-step GA calibrations were performed. First step was to calibrate the most sensitive parameters based on Chlorophyll_a (CHL) satellite SeaWIFS 8-day data. DMS model was then calibrated for another 5 most sensitive parameters. The best fitness was as good as -0.76 for CHL calibration in 1998-2002. The GA proved an efficient tool in the multiple-parameter calibration task. Model simulations indicate significant inter-annual variation in the CHL amount leading to significant inter-annual variability in the observed and modeled production of DMS and DMS flux in the study region in Arctic Ocean.
DNA binding sequence motifs are becoming increasingly important in the analysis of gene regulation, disease diagnosis and drug design. Although so far there are amount of tools available to discover these kinds of motifs, little was done to identify the biological functions, especially in tissue or cell type specific contributions, of those motifs. In this paper we used an integrated pipeline to discover sequences motifs for the promoter regions of human genes. Then we distinguished two types of motifs: tissue rich motifs (TRM) and tissue even motifs (TEM), using hypotheses test approaches including Bayesian hypothesis, Binomial distribution and traditional z-test. We finally got 233 overlapped TRMs and 56 TEMs. Most of those motifs are validated against JASPAR databases.
Discretization serves as an important preprocessing step for analyzing gene expression data and many algorithms have been proposed. However, most of the discretization methods were designed for microarrays. As a new technology, digital gene expression (DGE) profiles can overcome the limitation of microarrays and were applied in a widely range. In this paper, we proposed a novel discretization method for DGE data and the validations in a time-series gene expression dataset proved the efficiency of our method.
Chinese herbs always have activity on multiple targets. For the identification of potential anti-inflammatory compounds from Chinese herbs, six targets, which are mostly associated with inflammatory, were selected as following: Cox-2 (cyclooxygenase 2), PDE4B (phosphodiesterase 4B), p38α MAPK (p38α mitogen-activated protein kinase), JNK3 (c-Jun N-terminal kinases 3), ICE (interlenkin-1β converting enzyme) and iNOS (inducible type of nitric oxide synthase). Structure-based pharmacophore models of the inhibitors of each target were generated by LigandScout based on complexes from the PDB (Protein Data Bank). Based on the screening results of MDDR (MDL Drug Data Report database) and a new metric CAI (comprehensive appraisal index), the best models for each target were defined and used to identify the potential anti-inflammatory compounds from Chinese herbs. Six compounds and sixteen herbs were obtained that can act on multiple targets. The traditional function of the most hit herbs was `heat-clearing and detoxifying', which has been experimentally demonstrated to have anti-inflammatory activity.
Mutual-inhibition motif is frequently-occuring motif in transcriptional regulatory networks for cell lineage commitment. Stable attractors represent cell commitment state. But how progenitor-specific transcription factors stabilize progenitor cells and commit them to different cell fates remains unexplained. In this paper we represent the motif for cell commitment, composed of mutual-inhibition motif and progenitor-specific transcription factor, and develop associated mathematical model. In the analysis of bifurcation and dynamical simulation, the model could exhibit multiple steady stable states and transition between them, coo responding to progenitor, committed cell state and different commitment processes. Furthermore, we demonstrate that different commitment patterns, for example that of hematopoitic stem cell and neural stem cell, could be represented with different bifurcation features.
Rheumatoid arthritis (RA) is a chronic disease that affects the joints, often those in a person's wrists, fingers, and feet. In contrast to FDA-approved anti-RA drugs, Tripterygium wilfordii Hook F (TwHF), a traditional Chinese medicine (TCM), featured as multi-targeting, have been acknowledged with notable anti-RA effects although the pharmacology is unclear. In this work, we investigated the therapeutic mechanisms of TwHF at protein network level. First, RA-associated genes, the protein targets of FDA approved anti-RA drugs and TwHF were collected. Then we mapped the protein targets of TwHF on the drug-target network of FDA approved anti-RA drugs and KEGG RA pathway, based on these information and resources. Furthermore, we quantitatively analyzed the anti-rheumatic effect of TwHF and compared it with those of FDA approved anti-RA drugs by a network based anti-rheumatic effect score. Our study suggests that TwHF may function as a combination of disease-modifying anti-rheumatic drug and non-steroidal anti-inflammatory drug and its anti-rheumatic power could be comparable with that of anti-inflammatory agents. This study may facilitate our understanding of the RA treatment by TwHF from the perspective of network systems and it may suggest new approach for the study of TCM pharmacology.
In this work, based on the ACF model and the SVM classifier, succeeded on trials mining information that it's more effective to analyze the subcellular localization prediction of apoptosis proteins when adopting hydrophobicity property. This information is obtained in three benchmark datasets by using the ACF model and SVM to scan the AAindex database, which contains 544 kinds of amino acids. The contribution of this work is that it first did a comprehensive research on the effectiveness of the amino acid index for the subcellular localization of apoptosis proteins.
Modeling is an important direction in systems biology. The target towards kinetic modeling for metabolic network is to develop a practical computational method which can handle incomplete parameters. In principle, we could start with a set of randomly chosen parameters; calculating fluxes and metabolites concentration and comparing with experiments; iterating until the best parameters are found. But the large parametric space may require billions of times of iterations. In order to overcome such a difficulty, we develop a method to obtain the structure of parametric space. We are able to discover the correlation between parameters and variables, which is helpful for us to estimate the possible value of parameters. Differ from previous method, the implicit relationship between parameter and variable are also provided directly by our method, which provides a potential for us to analyze the feature of metabolic network.
Biochemical systems can be described by biochemical reactions. Biochemical reactions can be investigated through mathematical modeling and stochastic simulations. Deterministic and stochastic models are two basic categories of models for biochemical reactions. Due to transmembrane transportation of biochemical species and delayed degradation, time delays are ubiquitous in coupled biochemical systems. Therefore, models for biochemical reactions can be further classified into delayed and un-delayed ones. For biochemical systems without delays, researchers have established the connections between deterministic models and stochastic models directly from the deterministic ones. For delayed biochemical systems, researchers have proposed some stochastic simulation methods to cope with biochemical reactions with time delays. However, the existing delayed stochastic simulation algorithms (SSA) are all incapable of realizing the comparison between highly nonlinear deterministic delayed models and stochastic models directly from the the deterministic ones. In this paper, we proposed a delayed SSA, which can realize the comparison between deterministic models and its stochastic counterparts. Furthermore, one can also use the algorithm to investigate intrinsic noise-induced behaviors, and the effect of system volumes. Several numerical examples show the effectiveness and correctness of our algorithm.
In microbial communities, the taxonomic structure and functional capability are highly related. We proposed a method by considering the combination of taxa and functional categories to explore the ecological mechanisms of microbial communities. Using GOS metagenomic samples, we tested this idea and its effectiveness. The combination of taxonomies and functional groups could reflect the difference between habitats and may help to explain the combination adaptability of microbes to environment.
The karyotype and chromosomal characteristics of the vulnerable species Onychostoma lini (Wu 1939) from China were studied by examining metaphase chromosome spreads obtained from kidney. According to the 200 metaphase spreads from 10 specimen of Onychostoma lini, captured from the Duliu river (located in the Pearl River system), China, the chromosome formula in the species might be described as 2n=50=12M+8SM+4ST+26T and FN=70. The mean values of chromosome lengths in Onychostoma lini ranged from 7.975 to 14.270 μm, and the haploid chromosome length of the species was 289.111±27.767 μm. This study provides first knowledge on karyotypes in Onychostoma lini which may facilitate aquaculture, conservation practices of the species. Also, the evolutionary level of Onychostoma lini is preliminarily analyzed based on the karyotype of the species in this paper.
Mathematical models have been used to understand the factors that govern infectious disease progression in viral infections. In this paper, based on the standard mass action incidence, an anti-HBV therapy model with time-delayed immune response is set up. The time-delay is used to describe the period of time for antigenic stimulation to generate CTLs. The globally asymptotically stable analysis of the infection-free equilibrium is given in the paper. Some conditions for Hopf bifurcation around endemic equilibrium to occur are also obtained by using the time delay as a bifurcation parameter.
With the development of next-generation sequencing and metagenomic technologies, the number of metagenomic samples of microbial communities is increasing with exponential speed. The comparison among metagenomic samples could facilitate the data mining of the valuable yet hidden biological information held in the massive metagenomic data. However, current methods for metagenomic comparison are limited by their ability to process very large number of samples each with large data size.In this work, we have developed an optimized GPU-based metagenomic comparison algorithm, GPU-Meta-Storms, to evaluate the quantitative phylogenetic similarity among massive metagenomic samples, and implemented it using CUDA (Compute Unified Device Architecture) and C++ programming. The GPU-Meta-Storms program is optimized for CUDA with non-recursive transform, register recycle, memory alignment and so on. Our results have shown that with the optimization of the phylogenetic comparison algorithm, memory accessing strategy and parallelization mechanism on many-core hardware architecture, GPU-Meta-Storms could compute the pair-wise similarity matrix for 1920 metagenomic samples in 4 minutes, which gained a speed-up of more than 1000 times compared to CPU version Meta-Storms on single-core CPU, and more than 100 times on 16-core CPU. Therefore, the high-performance of GPU-Meta-Storms in comparison with massive metagenomic samples could thus enable in-depth data mining from massive metagenomic data, and make the real-time analysis and monitoring of constantly-changing metagenomic samples possible.
Identifying regulatory genes partaking in disease development is important to medical advances. Since gene expression data of multiple experiments exist, combining results from multiple gene regulatory network discoveries offers higher sensitivity and specificity. However, data for multiple experiments on the same problem may not possess the same set of genes, and hence many existing combining methods are not applicable. In this paper, we approach this problem using a number of meta-analysis methods and compare their performances. Simulation results show that vote counting is outperformed by methods belonging to the Fisher's chi-square (FCS) family, of which FCS test is the best. Applying FCS test to the real human HeLa cell-cycle dataset, degree distributions of the combined network is obtained and compared with previous works. Consulting the BioGRID database reveals the biological relevance of gene regulatory networks discovered using the proposed method.