There is a misconception that intrinsic disorder in proteins is equivalent to darkness. The present study aims to establish, in the scope of the Swiss-Prot and Dark Proteome databases, the relationship between disorder and darkness. Three distinct predictors were used to calculate the disorder of Swiss-Prot proteins. The analysis of the results obtained with the used predictors and visualization paradigms resulted in the same conclusion that was reached before: disorder is mostly unrelated to darkness.
The dark proteome, as we define it, is the part of the proteome where 3D structure has not been observed either by homology modeling or by experimental characterization in the protein universe. From the 550.116 proteins available in Swiss-Prot (as of July 2016), 43.2% of the eukarya universe and 49.2% of the virus universe are part of the dark proteome. In bacteria and archaea, the percentage of the dark proteome presence is significantly less, at 12.6% and 13.3% respectively. In this work, we present a necessary step to complete the dark proteome picture by introducing the map of the dark proteome in the human and in other model organisms of special importance to mankind. The most significant result is that around 40% to 50% of the proteome of these organisms are still in the dark, where the higher percentages belong to higher eukaryotes (mouse and human organisms). Due to the amount of darkness present in the human organism being more than 50%, deeper studies were made, including the identification of 'dark' genes that are responsible for the production of so-called dark proteins, as well as the identification of the 'dark' tissues where dark proteins are over represented, namely, the heart, cervical mucosa, and natural killer cells. This is a step forward in the direction of gaining a deeper knowledge of the human dark proteome.
Recently we surveyed the dark-proteome, i.e., regions of proteins never observed by experimental structure determination and inaccessible to homology modelling. Surprisingly, we found that most of the dark proteome could not be accounted for by conventional explanations (e.g., intrinsic disorder, transmembrane domains, and compositional bias), and that nearly half of the dark proteome comprised dark proteins, in which the entire sequence lacked similarity to any known structure. In this paper we will present the Dark Proteome Database (DPD) and associated web services that provide access to updated information about the dark proteome.
During the last decades, the volume of genomic data has increased enormously, as has the complexity of related gene annotation data. As a result, there is an ever-increasing need for methods that can help biologists compare sets of genes efficiently. While several standard tools exist that support such analyses, to date there is no broadly accepted approach for the visual comparison of gene sets, and most tools generate as output flat files with lists of annotations and their associated significance scores. More advanced techniques are required by the visual analytics community to help exploration of results efficiently. Here, we report our preliminary work on adaptable, zoomable, and interactive tag-clouds as one such potential improvement.
Significance A key remaining frontier in our understanding of biological systems is the “dark proteome”—that is, the regions of proteins where molecular conformation is completely unknown. We systematically surveyed these regions, finding that nearly half of the proteome in eukaryotes is dark and that, surprisingly, most of the darkness cannot be accounted for. We also found that the dark proteome has unexpected features, including an association with secretory tissues, disulfide bonding, low evolutionary conservation, and very few known interactions with other proteins. This work will help future research shed light on the remaining dark proteome, thus revealing molecular processes of life that are currently unknown.
As the amount of genome information increases rapidly, there is a correspondingly greater need for methods that provide accurate and automated annotation of gene function. For example, many high-throughput technologies - e.g., next-generation sequencing - are being used today to generate lists of genes associated with specific conditions. However, their functional interpretation remains a challenge and many tools exist trying to characterize the function of gene-lists. Such systems rely typically in enrichment analysis and aim to give a quick insight into the underlying biology by presenting it in a form of a summary-report. While the load of annotation may be alleviated by such computational approaches, the main challenge in modern annotation remains to develop a systems form of analysis in which a pipeline can effectively analyze gene-lists quickly and identify aggregated annotations through computerized resources. In this article we survey some of the many such tools and methods that have been developed to automatically interpret the biological functions underlying gene-lists. We overview current functional annotation aspects from the perspective of their epistemology (i.e., the underlying theories used to organize information about gene function into a body of verified and documented knowledge) and find that most of the currently used functional annotation methods fall broadly into one of two categories: they are based either on 'known' formally-structured ontology annotations created by 'experts' (e.g., the GO terms used to describe the function of Entrez Gene entries), or - perhaps more adventurously - on annotations inferred from literature (e.g., many text-mining methods use computer-aided reasoning to acquire knowledge represented in natural languages). Overall however, deriving detailed and accurate insight from such gene lists remains a challenging task, and improved methods are called for. In particular, future methods need to (1) provide more holistic insight into the underlying molecular systems; (2) provide better follow-up experimental testing and treatment options, and (3) better manage gene lists derived from organisms that are not well-studied. We discuss some promising approaches that may help achieve these advances, especially the use of extended dictionaries of biomedical concepts and molecular mechanisms, as well as greater use of annotation benchmarks. (C) 2014 Elsevier Inc. All rights reserved.
Neste artigo reveem-se algumas tecnicas de processamento da informacao representada em imagens medicas digitais 2D, para obter malhas volumicas de orgaos do corpo humano para fins computacionais. Os processos descritos neste trabalho sao: segmentacao de estruturas representadas em imagens medicas, geracao de malhas superficiais e posterior pos-processamento (simplificacao e suavizacao), e ainda a geracao de malhas volumicas a partir das malhas superficiais anteriores. Sao considerados dois exemplos praticos utilizando software nao comercial. No primeiro exemplo comparam-se dois metodos de segmentacao para o cerebro e no segundo reconstroem-se os ossos e os musculos do antebraco.
Este artigo lida com o tratamento de informacao digital proveniente de imagens medicas (CT, MRI, etc), com o objectivo de obter malhas tridimensionais de orgaos humanos para propositos computacionais. Segmentacao de imagens, geracao de malhas superficiais, ajustamentos de malhas superficiais e ainda geracao de malha volumica, sao conceitos que sao referidos neste artigo. A aplicacao destes metodos e ilustrada aqui com um exemplo.