
Akdag, Sa Al-Rbeawi, Salam Alizamir, M Amirian, E Amodu, Omowunmi Ao, Songjian Badoga, Sandeep Balusamy, Saravanan Banerjee, Sayandeep Bao, Zhidong Benjamin, I. Nelson Binyet, E Bu, Quan Çalık, Ahmet Cano, Antonio Cao, Zhe Chandra Babu, Jakka Sarat Chang, Tao Chang, Xiangchun Chen, Anqing Chen, Liang Chen, Shuyuan Chen, Weizhong Chen, Z.H. Cheng, Liyuan Cheng, Rui Chong, W. W. F. Cui, Huiying Dai, Cheng Devi, B. L. A. Prabhavathi Ding, Rui Ding, Xiujian Du, Shanghai Du, Shuheng Du, Xuebin El-Seesy, Ahmed I. Fan, Chaojun Feng, Bo Feng, Wenjie Fiket, Zeljka Font, Xavier Fu, Hanliang Gan, Huajun Ge, Zhaolong Ge, Shilong Genc, Mustafa Geng, Shuai Gong, Yanjie González, B. M. Gozgor, G Guler, O Guo, Chen Guo, Tiankui Guris, Burak Hamza, M He, Baojie He, Kun Hou, Dy Hou, Mingcai Hou, Xiaowei Hu, Guang Hu, Jia Hu, Jie Hu, Tao Hu, Wenxuan Hu, Xiancai Huang, Bingxiang Huang, Haiping Hudisteanu, S Huo, Aidi Hussain, Furqan Jakovljevi c, Ivan Jeong, Daein Ji, Hancheng Jiang, Bo Jiang, Guangzheng Jiang, Wei Jiang, Zaixing Jiang, Zhenxue Jiao, Kun Jirasek, J Ju, Wei Kang, Xun Kaplan, Ya Kaya, Mustafa KeRdzior, S. Kedzior, Slawomir Khan, Mi Khan, Salman Kolak, Jonathan J. Kuszewski, Hubert Lai, Jin Lee, Kyung Jae Lee, Youngsoo Levendis, Yiannis A. Li, Bo Li, Changzhu Li, Dong Li, Guiqiang Li, Guorong Li, Hao Li, J Li, Meijun Li, Qing Li, Shengli Energy Exploration & Exploitation 2019, Vol. 37(2) 884–886 ! The Author(s) 2019 DOI: 10.1177/0144598719836561 journals.sagepub.com/home/eea
CHARGE syndrome is an autosomal dominant developmental disorder associated with a constellation of traits involving almost every organ and sensory system, in particular congenital anomalies, including choanal atresia and malformations of the heart, inner ear, and retina. Variants in CHD7 have been shown to cause CHARGE syndrome. Here, we report the identification of a novel de novo p.Asp2119_Pro2120ins6 duplication variant in a conserved region of CHD7 in a severely affected boy presenting with 3 and 5 of the CHARGE cardinal major and minor signs, respectively, combined with congenital umbilical hernia, congenital hernia at the linea alba, mildly hypoplastic inferior vermis, slight dilatation of the lateral ventricles, prominent metopic ridge, and hypoglycemic episodes.
The high mortality rate of neonatal sepsis is directly connected with time-consuming diagnostic methods that have low sensitivity and specificity. The need of the hour is to develop novel diagnostic techniques that are rapid and more specific. In this study, we estimated the expression levels of circulating microRNAs (miRNAs) that are involved in regulating immune response genes and underlying inflammatory responses, which may be used for sepsis diagnosis. The total circulating miRNA was isolated and the candidate miRNAs (miR-132, miR-146a, miR-155, and miR-223) were quantified by real-time polymerase chain reaction technique. Statistical analysis revealed that miR-132 (P < .01) and miR-223 (P < .05) were downregulated in septic newborns compared with healthy babies. The decrease in expression of miR-132 and miR-223 may be associated with increased expression of immune-related genes involved in TLR (Toll-like receptor) signaling pathway. Further case-control studies with large sample size are required to identify the potential of miRNAs in neonatal sepsis diagnosis.
A great deal of ambiguity exists in the development of guidelines for genomic applications used in clinical practice. The GRADE (Grading of Recommendations Assessment, Development and Evaluation) approach has the potential to be applied in the guidelines and recommendations development process in genomics. Here, we discuss whether and how GRADE can be applied to address the challenges posed by the evidence-based guidelines and recommendations development process in genomics. To see how GRADE can complement to the current guidelines development in genomics, we compare and contrast GRADE with other approaches. GRADE differed from other methods by incorporating patient values and preferences and balance of consequences. We conclude that the groups trying to implement genomics into practice may gleam more information from applying the GRADE framework. However, it is not clear yet whether GRADE can address the issue of timeliness in terms of the differences between the time required for guidelines development and the rapid pace of genomics.
Ganoderma lucidum (lingzhi) has been used for the general promotion of health in Asia for many centuries. The common method of consumption is to boil lingzhi in water and then drink the liquid. In this study, we examined the potential anticancer activities of G. lucidum submerged in two commonly consumed forms of alcohol in East Asia: malt whiskey and rice wine. The anticancer effect of G. lucidum, using whiskey and rice wine-based extraction methods, has not been previously reported. The growth inhibition of G. lucidum whiskey and rice wine extracts on the prostate cancer cell lines, PC3 and DU145, was determined. Using Affymetrix gene expression assays, several biologically active pathways associated with the anticancer activities of G. lucidum extracts were identified. Using gene expression analysis (real-time polymerase chain reaction [RT-PCR]) and protein analysis (Western blotting), we confirmed the expression of key genes and their associated proteins that were initially identified with Affymetrix gene expression analysis.
The issue of multiple testing, also termed multiplicity, is ubiquitous in studies where multiple hypotheses are tested simultaneously. Genome-wide association study (GWAS), a type of genetic association study that has gained popularity in the past decade, is most susceptible to the issue of multiple testing. Different methodologies have been employed to address the issue of multiple testing in GWAS. The purpose of the review is to examine the methodologies employed in dealing with multiple testing in the context of gene discovery using GWAS in sickle cell disease complications.
High-density linkage maps are vital to supporting the correct placement of scaffolds and gene sequences on chromosomes and fundamental to contemporary organismal research and scientific approaches to genetic improvement, especially in paleopolyploids with exceptionally complex genomes, eg, upland cotton (Gossypium hirsutum L., "2n = 52"). Three independently developed intraspecific upland mapping populations were analyzed to generate 3 high-density genetic linkage single-nucleotide polymorphism (SNP) maps and a consensus map using the CottonSNP63K array. The populations consisted of a previously reported F-2, a recombinant inbred line (RIL), and reciprocal RIL population, from "Phytogen 72" and "Stoneville 474" cultivars. The cluster file provided 7417 genotyped SNP markers, resulting in 26 linkage groups corresponding to the 26 chromosomes (c) of the allotetraploid upland cotton (AD) 1 arisen from the merging of 2 genomes ("A" Old World and "D" New World). Patterns of chromosome-specific recombination were largely consistent across mapping populations. The high-density genetic consensus map included 7244 SNP markers that spanned 3538 cM and comprised 3824 SNP bins, of which 1783 and 2041 were in the A(t) and D-t subgenomes with 1825 and 1713 cM map lengths, respectively. Subgenome average distances were nearly identical, indicating that subgenomic differences in bin number arose due to the high numbers of SNPs on the D-t subgenome. Examination of expected recombination frequency or crossovers (COs) on the chromosomes within each population of the 2 subgenomes revealed that COs were also not affected by the SNPs or SNP bin number in these subgenomes. Comparative alignment analyses identified historical ancestral A(t)-subgenomic translocations of c02 and c03, as well as of c04 and c05. The consensus map SNP sequences aligned with high congruency to the NBI assembly of Gossypium hirsutum. However, the genomic comparisons revealed evidence of additional unconfirmed possible duplications, inversions and translocations, and unbalance SNP sequence homology or SNP sequence/loci genomic dominance, or homeolog loci bias of the upland tetraploid At and Dt subgenomes. The alignments indicated that 364 SNP-associated previously unintegrated scaffolds can be placed in pseudochromosomes of the NBI G hirsutum assembly. This is the first intraspecific SNP genetic linkage consensus map assembled in G hirsutum with a core of reproducible mendelian SNP markers assayed on different populations and it provides further knowledge of chromosome arrangement of genic and nongenic SNPs. Together, the consensus map and RIL populations provide a synergistically useful platform for localizing and identifying agronomically important loci for improvement of the cotton crop.
Long noncoding RNAs (lncRNAs) which were initially dismissed as "transcriptional noise" have become a vital area of study after their roles in biological regulation were discovered. Long noncoding RNAs have been implicated in various developmental processes and diseases. Here, we perform exon mapping of human lncRNA sequences (taken from National Center for Biotechnology Information GenBank) using digital filters. Antinotch digital filters are used to map out the exons of the lncRNA sequences analyzed. The period 3 property which is an established indicator for locating exons in genes is used here. Discrete wavelet transform filter bank is used to fine-tune the exon plots by selectively removing the spectral noise. The exon locations conform to the ranges specified in GenBank. In addition to exon prediction, G-C concentrations of lncRNA sequences are found, and the sequences are searched for START and STOP codons as these are indicators of coding potential.
In mammals, extracellular miRNAs circulate in biofluids as stable entities that are secreted by normal and diseased tissues, and can enter cells and regulate gene expression. Drosophila melanogaster is a proven system for the study of human diseases. They have an open circulatory system in which hemolymph (HL) circulates in direct contact with all internal organs, in a manner analogous to vertebrate blood plasma. Here, we show using deep sequencing that Drosophila HL contains RNase-resistant circulating miRNAs (HL-miRNAs). Limited subsets of body tissue miRNAs (BT-miRNAs) accumulated in HL, suggesting that they may be specifically released from cells or particularly stable in HL. Alternatively, they might arise from specific cells, such as hemocytes, that are in intimate contact with HL. Young and old flies accumulated unique populations of HL-miRNAs, suggesting that their accumulation is responsive to the physiological status of the fly. These HL-miRNAs in flies may function similar to the miRNAs circulating in mammalian biofluids. The discovery of these HL-miRNAs will provide a new venue for health and disease-related research in Drosophila.
The objective of this study was to explore the known narrow genetic diversity and discover single-nucleotide polymorphic (SNP) markers for marker-assisted breeding within Pima cotton (Gossypium barbadense L.) leaf transcriptomes. cDNA from 25-day plants of three diverse cotton genotypes [Pima S6 (PS6), Pima S7 (PS7), and Pima 3-79 (P3-79)] was sequenced on Illumina sequencing platform. A total of 28.9 million reads (average read length of 138 bp) were generated by sequencing cDNA libraries of these three genotypes. The de novo assembly of reads generated transcriptome sets of 26,369 contigs for PS6, 25,870 contigs for PS7, and 24,796 contigs for P3-79. A Pima leaf reference transcriptome was generated consisting of 42,695 contigs. More than 10,000 single-nucleotide polymorphisms (SNPs) were identified between the genotypes, with 100% SNP frequency and a minimum of eight sequencing reads. The most prevalent SNP substitutions were C-T and A-G in these cotton genotypes. The putative SNPs identified can be utilized for characterizing genetic diversity, genotyping, and eventually in Pima cotton breeding through marker-assisted selection.
With the increasing number of sequenced genomes and their comparisons, the detection of orthologs is crucial for reliable functional annotation and evolutionary analyses of genes and species. Yet, the dynamic remodeling of genome content through gain, loss, transfer of genes, and segmental and whole-genome duplication hinders reliable orthology detection. Moreover, the lack of direct functional evidence and the questionable quality of some available genome sequences and annotations present additional difficulties to assess orthology. This article reviews the existing computational methods and their potential accuracy in the high-throughput era of genome sequencing and anticipates open questions in terms of methodology, reliability, and computation. Appropriate taxon sampling together with combination of methods based on similarity, phylogeny, synteny, and evolutionary knowledge that may help detecting speciation events appears to be the most accurate strategy. This review also raises perspectives on the potential determination of orthology throughout the whole species phylogeny.
Genomic studies have become noncoding RNA (ncRNA) centric after the study of different genomes provided enormous information on ncRNA over the past decades. The function of ncRNA is decided by its secondary structure, and across organisms, the secondary structure is more conserved than the sequence itself. In this study, the optimal secondary structure or the minimum free energy (MFE) structure of ncRNA was found based on the thermodynamic nearest neighbor model. MFE of over 2600 ncRNA sequences was analyzed in view of its signal properties. Mathematical models linking MFE to the signal properties were found for each of the four classes of ncRNA analyzed. MFE values computed with the proposed models were in concordance with those obtained with the standard web servers. A total of 95% of the sequences analyzed had deviation of MFE values within ±15% relative to those obtained from standard web servers.
Conventional rigid docking algorithms have been unsatisfactory in their computational results, largely due to the fact that protein structures are flexible in live environments. In response, we propose to introduce the side-chain flexibility in protein motif into the docking. First, the Morse theory is applied to curvature labeling and surface region growing, for segmentation of the protein surface into smaller patches. Then, the protein is described by an ensemble of conformations that incorporate the flexibility of interface side chains and are sampled using rotamers. Next, a 3D rotation invariant shape descriptor is proposed to deal with the flexible motifs and surface patches; thus, pairwise complementarity matching is needed only between the convex patches of ligand and the concave patches of receptor. The iterative closest point (ICP) algorithm is implemented for geometric alignment of the two 3D protein surface patches. Compared with the fast Fourier transform-based global geometric matching algorithm and other methods, our FlexDock system generates much less false-positive docking results, which benefits identification of the complementary candidates. Our computational experiments show the advantages of the proposed flexible docking algorithm over its counterparts.
Conventional protein docking methods use rigid docking algorithms but the results have been unsatisfactory, because protein structures are flexible in live environments. In this study, the side-chain flexibility is introduced into protein docking. First, the Morse theory is applied to curvature labelling and region-growing for segmenting the protein surface into smaller patches. The surface critical points are then extracted for each surface region which has been obtained in the segmentation, and a list of more precise and smaller surface patches can be obtained. Secondly, based on the segmented surface, we propose to describe the protein with an ensemble of conformations which incorporate the flexibility of interface side-chains by sampling their conformations using rotamers. We define a 3D rotation invariant shape descriptor based on the Spherical Harmonics Descriptor to deal with the flexible protein surface patches. The signature for each surface patch can be more accurately described and easily compared for the similarity between the signatures. The pairwise complementarity matching is needed only between the convex patches of the ligand with concave patches of the receptor and vice versa, thus the computational cost is greatly reduced. Finally, we are able to use the iterative closest point (ICP) algorithm, thanks to the advantage of our matching method that is able to pre-compute and store the transformation invariant, fast Fourier transform (FFT) docking correlation based representation of each flexible surface patch. Compared with the global geometric matching algorithm and other methods, our FlexDock system generates much less false positive docking candidates, it shows promising advantages in identification of the complementary candidates.
Aiming at generating a comprehensive genomic database on Elaeis spp., our group is leading several R&D initiatives with Elaeis guineensis (African oil palm) and Elaeis oleifera (American oil palm), including the whole-genome sequencing of the last. Genome size estimates currently available for this genus are controversial, as they indicate that American oil palm genome is about half the size of the African oil palm genome and that the genome of the interspecific hybrid is bigger than both the parental species genomes. We estimated the genome size of three E. guineensis genotypes, five E. oleifera genotypes, and two interspecific hybrids genotypes. On average, the genome size of E. guineensis is 4.32 ± 0.173 pg, while that of E. oleifera is 4.43 ± 0.018 pg. This indicates that both genomes are similar in size, even though E. oleifera is in fact bigger. As expected, the hybrid genome size is around the average of the two genomes, 4.40 ± 0.016 pg. Additionally, we demonstrate that both species present around 38% of GC content. As our results contradict the currently available data on Elaeis spp. genome sizes, we propose that the actual genome size of the Elaeis species is around 4 pg and that American oil palm possesses a larger genome than African oil palm.
In the present study, recurrent copy number variations (CNVs) from non-tumor blood cell DNAs of Caucasian non-cancer subjects and glioma, myeloma, and colorectal cancer-patients, and Korean non-cancer subjects and hepatocellular carcinoma, gastric cancer, and colorectal cancer patients, were found to reveal for each of the two ethnic cohorts highly significant differences between cancer patients and controls with respect to the number of CN-losses and size-distribution of CN-gains, suggesting the existence of recurrent constitutional CNV-features useful for prediction of predisposition to cancer. Upon identification by machine learning, such CNV-features could extensively discriminate between cancer-patient and control DNAs. When the CNV-features selected from a learning-group of Caucasian or Korean mixed DNAs consisting of both cancer-patient and control DNAs were employed to make predictions on the cancer predisposition of an unseen test group of mixed DNAs, the average prediction accuracy was 93.6% for the Caucasian cohort and 86.5% for the Korean cohort.
Most molecular biological concepts derive from physical chemical assumptions about the genetic code that are basically more than 40 years old. Additionally, systems biology, another quantitative approach, investigates the sum of interrelations to obtain a more holistic picture of nucleotide sequence order. Recent empirical data on genetic code compositions and rearrangements by mobile genetic elements and noncoding RNAs, together with results of virus research and their role in evolution, does not really fit into these concepts and compel a reexamination. In this review, we try to find an alternate hypothesis. It seems plausible now that if we look at the abundance of regulatory RNAs and persistent viruses in host genomes, we will find more and more evidence that the key players that edit the genetic codes of host genomes are consortia of RNA agents and viruses that drive evolutionary novelty and regulation of cellular processes in all steps of development. This agent-based approach may lead to a qualitative RNA sociology that investigates and identifies relevant behavioral motifs of cooperative RNA consortia. In addition to molecular biological perspectives, this may lead to a better understanding of genetic code evolution and dynamics.
Most molecular biological concepts derive from physical chemical assumptions about the genetic code that are basically more than 40 years old. Additionally, systems biology, another quantitative approach, investigates the sum of interrelations to obtain a more holistic picture of nucleotide sequence order. Recent empirical data on genetic code compositions and rearrangements by mobile genetic elements and noncoding RNAs, together with results of virus research and their role in evolution, does not really fit into these concepts and compel a reexamination. In this review, we try to find an alternate hypothesis. It seems plausible now that if we look at the abundance of regulatory RNAs and persistent viruses in host genomes, we will find more and more evidence that the key players that edit the genetic codes of host genomes are consortia of RNA agents and viruses that drive evolutionary novelty and regulation of cellular processes in all steps of development. This agent-based approach may lead to a qualitative RNA sociology that investigates and identifies relevant behavioral motifs of cooperative RNA consortia. In addition to molecular biological perspectives, this may lead to a better understanding of genetic code evolution and dynamics.
Ribonucleic acids (RNA) are hypothesized to have preceded their derivatives, deoxyribonucleic acids (DNA), as the molecular media of genetic information when life emerged on earth. Molecular biologists are accustomed to the dramatic effects a subtle variation in the ribose moiety composition between RNA and DNA can have on the stability of these molecules. While DNA is very stable after extraction from biological samples and subsequent treatment, RNA is notoriously labile. The short half-life property, inherent to RNA, benefits cells that do not need to express their entire repertoire of proteins. The cellular machinery turns off the production of a given protein by shutting down the transcription of its cognate coding gene and by either actively degrading the remaining mRNA or allowing it to decay on its own. The steady-state level of each mRNA in a given cell varies continuously and is specified by changing kinetics of synthesis and degradation. Because it is technically possible to simultaneously measure thousands of nucleic acid molecules, these quantities have been studied by the life sciences community to investigate a range of biological problems. Since the RNA abundance can change according to a wide range of perturbations, this makes it the molecule of choice for exploring biological systems; its instability, on the other hand, could be an underestimated source of technical variability. We found that a large fraction of the RNA abundance originally present in the biological system prior to extraction was masked by the RNA labeling and measurement procedure. The method used to extract RNA molecules from cells and to label them prior to hybridization operations on DNA arrays affects the original distribution of RNA. Only if RNA measurements are performed according to the same procedure can biological information be inferred from the assay read out.
With the advent of high throughput sequencing platforms and relevant analytical tools, the rate of microbial genome sequencing has accelerated which has in turn led to better understanding of microbial molecular biology and genetics. The complete genome sequences of important industrial organisms provide opportunities for human health, industry, and the environment. Bacillus species are the dominant workhorses in industrial fermentations. Today, genome sequences of several Bacillus species are available, and comparative genomics of this genus helps in understanding their physiology, biochemistry, and genetics. The genomes of these bacterial species are the sources of many industrially important enzymes and antibiotics and, therefore, provide an opportunity to tailor enzymes with desired properties to suit a wide range of applications. A comparative account of strengths and weaknesses of the different sequencing platforms are also highlighted in the review.