
Comparative sequence analysis is a powerful approach to identify functional elements for functional genomics. For gene prediction, this would mean whether given genomic sequences exhibit significant similarity to some arbitrary region in the genome sequence of an evolutionary related organism or not. Three gene prediction methods were examined to see how they perform on randomly chosen sequences. The focus was on examining the methods rather than data analysis. Thus, the dataset was selected in a flexible manner and may be biased. Our analysis of methods indicated that there are not many differences between recent version of GENEMARK.HMM and GENSCAN in terms of the algorithm used. Both programs use duration HMM and can predict genes and exons on both DNA strands simultaneously. MZEF, on the other hand, uses completely different approach. It employs quadratic discriminant function to distinguish between coding and noncoding regions. The analysis showed that the new generation of programs has substantially better results than the programs analyzed in previous studies. Specifically, GENEMARK.HMM and GENSCAN performed relatively better than MZEF.
E-Health systems have an increasing role in the everyday health care delivery. It facilitates providing health care including consultations to individuals in rural areas or in areas with unavailability of physicians. In this paper we highlight the study conducted on using mobile devices as extension to e-health. We also discuss usability study conducted on our system by a group of health care professionals.
Heart disease remains primary cause of premature mortality in the western world. To date, there are no cures-just heuristics for reducing the risk(s) of contracting the disease through diet and exercise. In this study, we examine two typical heart disease datasets that provide quantitative information regarding possible risk factors associated with heart disease. We wish to extract rules that relate these risk factors to the occurrence of heart disease in a form that is easily interpreted by both medical professionals and laymen alike. Our approach based on rough sets and the notion of approximate reducts, provides a rigorous and accurate classification scheme, which produces results that rival or excel more conventional methods. During the process of classification, rough sets removes redundant information and generates rules that can be used to describe the relationship(s) between the attributes and the final decision.
Breast cancer remains the dominant form of cancer in women in most of the western world. The best current diagnostic scheme entails a collaborative approach between cytologist, radiologist, and the opinion of the examining surgeon. Due to the invasive nature of a radical mastectomy, most decisions err on the side of caution, resulting in very high specificity - approaching 100% in most cases. Generally, high specificity is accompanied by much lower sensitivity. In this paper, we present a medical decision support system that provides very high values for specificity without the concomitant reduction in :sensitivity. The DSS we develop is derived from a study of 692 Breast cancer patient with 12 attributes including the decision). Our results produce an average specificity of 94% and a sensitivity of 94%, with a positive predictive value of 96%. Furthermore, our study examined the role of the relative importance of the individual attributes towards the decision process. Elimination of unnecessary attributes reduces the burden on medical personnel and may increase the sensitivity of the diagnostic procedure.
We develop mathematical models to estimate the number of HIV positive People in U.S.A., in U.K., and in Canada in 2001. In all the three cases, our estimates are very close to the official ones. It is emphasized that our estimates depend upon the models and will change as we change the models and/or the assumptions therein.
Forest Ecosystem Modeling is a difficult problem, because 1) It's generally hard to isolate the factors that have influence on a tree species growth; 2) with interplay of multiple plant species, it's essentially an NP complete problem. The above characters make this problem an ideal candidate for distributed simulations. A variable resolution scheme of individual modeling is proposed. This paper suggests a scalable parallel solution to the forest ecosystem modeling problem, which not only helps our understanding of such complex ecology system, but also makes large scale forest ecosystem management a reachable goal.
A modification of the mathematical model of the respiratory system of a human organism is offered. The model is used to study the consequences of pharmacological correction of hypoxic states of the organism in case of ischemia and atherosclerosis diseases. It was assumed that the reason of the development of ischemia is partial or full occlusion or the microcirculatory vessels of myocardium. The computing experiments with the model allowed to calculate the decrease of arterial blood pressure and to define the period of stabilization of oxygen and carbon dioxide regimes of the organism. The model of partial occlusion of coronary arteries is a convenient tool to investigate the dynamics of regimes of respiratory gases in the heart by ischemia disease and to anticipate the changes made by pharmacological correction.
In the present work there is considered one non-local in time initial boundary value problem for one non-linear nonstationary, parabolic equation originated from the mathematical modelling of certain bio-chemical process. In the above mentioned problem the linear member is a coercive operator and nonlinear one is monotone bounded operator. Non-local in tinge conditions are given in different forms (e.g. in integral form). There are investigated the existence and uniqueness of solution for stated problems. To solve these problems some iteration processes are constructed. dome a priori estimations are obtained and convergence of the iteration process is proved. In certain assumptions positiveness of solution is demonstrated.
While significant efforts have been devoted to the development of sophisticated search engines for a wide variety of data sources, how to record and share knowledge derived from data retrieved by search engines is still a major challenge.In this article, we discuss the problem of the lack of visual representations of data and concepts in current search engines. We then present a conceptual relationship diagramming system we have created to deal this problem. Compare to existing diagramming/drawing tools, our solution enables the sharing and modification of conceptual diagrams by researchers in different geographic locations. We believe such a system will expedite extraction and sharing of knowledge and ideas from various data sources.BCDE application can be downloaded at http://arrayanalysis.mhri.ined.uinich.edit/draw/BCD Editor windows beta.exe. BCDE applet and the BCDE database can both be accessed at http://arrayanalysis.mhri.med.umich.edu/draw/. Please use public/public as username/password to login to the network.
Prediction of protein structures attracts many researchers. Given a protein main chain conformation, building side chains is an important part to predict protein structures. Most of the methods that predict side chain conformations use statistical data generated from known protein structures. It is a computationally intractable problem for building side chains from all possible rotamers simultaneously using information of known protein structures. Reducing the number of possibility is a main issue to predict side chain conformation. This paper analyzes the all-atom distances from high-resolution protein structures to find the conformational preferences of amino acid residues. The results of our analysis indicate that unrelated side chain rotamers can be eliminated dramatically. Consequently, searching for suitable residue side chains from possible side chain conformations becomes computationally possible.
This paper describes in detail the design of a 32-channel mixed-signal full-customized CMOS integrated biopotential sensor chip for extracellular recording of neural signal. Each analog channel has three major sub-circuit macros, which includes contact sensors, unity gain buffer and variable gain amplifier. Each channel has three working modes: test mode, run mode and zero modes Analog channels are grouped and addressed by eight 4-1 decoders which obtain control signals from a 16-bit serial-in parallel-out shift register. This device was fabricated by MOSIS using AMI's 1.5 mu m, double poly, double level metal CMOS technology.
RNA plays a crucial role in post-transcriptional regulation. Similar to transcriptional regulation, post-transcriptional regulation is often accomplished by the binding of proteins to specific motifs in mRNA molecules. Unlike DNA binding proteins, which recognize motifs composed of conserved sequences, RNA protein binding sites are more conserved in structures than in sequences. A lot of works have been done for RNA structure prediction; however, most of them focus on single RNA structure prediction instead of finding characteristic structure motifs within a RNA family. Though some current approaches can now identify common structure motifs from a set of RNAs, they typically assume the given set forms a single family, which is not necessarily correct. We propose a new adaptive method that conducts structure prediction and clustering simultaneously. Its performance is demonstrated on several real RNA families.
Cyanobacteria and the cyanophage that infect these bacteria are abundant throughout fresh water and marine ecosystems. Unfortunately, the knowledge of genomic interaction between these cyanophages and their hosts are limited due to the lack of integration of the cyanogroup information. To remedy this deficiency, a Cyanogroup Genomics Knowledge Base (CGKB) for cyanophages and cyanobacteria has been initiated: http://www.cyanogroup.com/. This knowledge base includes: (1) the literature database of freshwater and marine cyanogroups; (2) the genomic database of these cyanogroups including NCBI GenBank data and sequence data from both our laboratory and that of other researchers; (3) data analysis - DNA, protein sequences and phylogenetic analysis; (4) relevant links for the cyanogroups. Regular updates of the local database will be performed to maintain the high accuracy of the analysis and BLAST searches. CGKB also provides sequence and phylogenetic analysis tools for users to search and query any related information stored in the database.
A new method for classification of protein-domain primary structures into protein-domain families is presented which is based on correlations between hydrophobicity measures. A series of hydrophobicity profiles is generated from a training set for each protein-domain family. A protein domain of unknown family classification is then compared with the profiles for the family using correlations to generate a classification score. The Pfam database "top-twenty" seed families are used to test the method with good results.
We present a framework for the semantic data mining of nursing care cases, where the semantic layer consists of taxonomy classes for nursing diagnoses, interventions, and outcomes. The semantic layer is expressed in XML documents and semantic queries are implemented via mediator documents and XPath expressions. The primary purpose of the data mining framework is to allow the clinician to perform complex queries on nursing care data, across nursing taxonomy levels, in order to assess the effectiveness of nursing interventions in producing desirable outcomes for various nursing diagnoses.
Signal compression is an important element encountered in data storage applications. Over the years various techniques for data reductions have been proposed. In this paper we introduce an effective method for compressing semi-periodic signals. Although the approach is applicable to any semi-periodic signal, our attention is focused on the compression of electrocardiogram (ECG) signals. An ECG signal is composed of many similar beats which makes it to behave semi-periodic. This paper deals with beat variation periods and exploits the correlation between cycles (inter-beat) and correlation within each cycle (intra-beat) for compression. For efficient compression a 2-dimensional array is constructed from the one dimensional ECG signal. Since reasonable results in image compression have been achieved by means of set partitioning in hierarchical trees (SPIHT) algorithm, we use SPIHT algorithm to code the 2-D wavelet transform of the ECG signals.Experimental results on selected records of ECG from the MIT-BIH arrhythmia database indicate that the proposed algorithm is significantly more efficient than those schemes previously proposed for ECG compression.
Calmodulin is a calcium-sensing signaling protein, and its conformation can be largely changed by binding with Ca2+ and peptide. However, which part of the calmodulin, and how does it contribute to the conformational change are unclear. In this study, a 4-nanosecond molecular dynamics simulation was performed to investigate the conformational change of whole calmodulin molecular starting from a bend structure caused by Ca2+ and peptide, when the Ca2+ and peptide were removed. To discover the main driver of the conformational change, Four 500-picosecond simulations of partial calmodulin molecule were also carried out. In the simulation of the whole calmodulin molecule, it was found that calmodulin changed to a more collapsed conformation from its initial crystal form. Of the total 8 helices and 7 loops, 6 helices and 4 loops moved toward the center-of-mass of the whole molecule, and 2 helices and 3 loops moved away from the center. The conformation collapse time of calmodulin is about I nanosecond. In the simulation of the partial calmodulin molecule, the central linker became a perfect hairpin-shape helix. This indicates that the central linker is the main driver of the contractive conformational change. We think the results are very useful for revealing the disorganization of ligand-binding proteins under the unliganded condition and for the drug design.
This paper describes an efficient software tool ProkProbePicker (PPP) for the design of arrays of 16S rRNA-targeted probes that can be used to determine the genetic affinity of unknown bacteria. Unlike most currently available applications that have been developed to design probes for a single species or bacterial grouping, PPP was designed to extract probes for all major groupings in a known phylogenetic tree. Arrays designed by PPP will allow determination of the genetic affinity of an unknown organism relative to the groups included in the array. PPP employs the Karp-Rabin algorithm for string comparisons and the AVL data structure for a balanced binary search and then fore is more efficient than other tools. As an auxiliary part of PPP, the program OrgMap was designed to help analyze hybridization data based on probes designed by PPP. To test the capability of mapping unknown organisms based on PPP-designed probes, we conducted in silico hybridization experiments on two recently identified organisms. The results showed that a PPP designed array could correctly determine the genetic affinity of these organisms to the gents level in one case and the species level in the other. It is concluded that PPP is a useful tool for the design of arrays that can determine genetic affinity of unknown organisms.
Delivering an injection correctly into the small diameter lumen of a tubular biological structure placed under the skin and not visible to the eye presents difficulties in medical practice. Suitable assist devices are currently not available. A new concept has been formulated to combine four modes of nonradiological imaging: scanning electrical impedance plethysmography; ultrasound; transillumination; and Moire topography to localize the subcutaneous biological structure. Image data is converted to a signal which controls an electro-mechanical device which assists the medical personnel's hand holding the injection syringe to approach the biological structure accurately. Force sensors are positioned at the base of the injection needle coupled to signal conditioning systems and piezoelectric tactile vibrators. Thereby enhanced perception of forces felt as each biological tissue layer is penetrated is obtained and the common problems of excessive penetration overcome. Design focuses on the specific application of delivering an injection into the vas deferens in the male scrotum for reversible contraception. The system may be adapted for injection into blood vessels.
This paper presents a robot-kinematic analysis of relative pose (position and orientation) distribution among amino acids in proteins. Unlike the Ramachandran plot that provides information of protein conformations using two angles, our method offers six-dimensional relative pose information. For each amino acid, a coordinate frame is affixed to the R group (side chain) of the amino acid. By computing the relative position vector and orientation matrix between two coordinate frames, the relative pose between two amino acids can be obtained. In addition, the effect of helix length is also discussed. This method is applied to 1230 protein structures in Protein Data Bank and results are discussed. This approach can be a useful tool for generating new models of polypeptide chains.