The 'canonical' protein sets distributed by UniProt are widely used for similarity searching, and functional and structural annotation. For many investigators, canonical sequences are the only version of a protein examined. However, higher eukaryotes often encode multiple isoforms of a protein from a single gene. For unreviewed (UniProtKB/TrEMBL) protein sequences, the longest sequence in a Gene-Centric group is chosen as canonical. This choice can create inconsistencies, selecting >95% identical orthologs with dramatically different lengths, which is biologically unlikely. We describe the ortho2tree pipeline, which examines Reference Proteome canonical and isoform sequences from sets of orthologous proteins, builds multiple alignments, constructs gap-distance trees, and identifies low-cost clades of isoforms with similar lengths. After examining 140 000 proteins from eight mammals in UniProtKB release 2022_05, ortho2tree proposed 7804 canonical changes for release 2023_01, while confirming 53 434 canonicals. Gap distributions for isoforms selected by ortho2tree are similar to those in bacterial and yeast alignments, organisms unaffected by isoform selection, suggesting ortho2tree canonicals more accurately reflect genuine biological variation. 82% of ortho2tree proposed-changes agreed with MANE; for confirmed canonicals, 92% agreed with MANE. Ortho2tree can improve canonical assignment among orthologous sequences that are >60% identical, a group that includes vertebrates and higher plants.
The purpose of the meeting described in this review was to decide how best to ensure the sustainability of the Network for Integrating Bioinformatics into Life Science Education (NIBLSE; pronounced “nibbles”). Biology research today generates large and complex datasets, and the analysis of these datasets is becoming increasingly critical to progress in the field. The long-term goal of NIBLSE is to address this need and achieve the full integration of bioinformatics into undergraduate life sciences education. Meeting participants supported several next steps for NIBLSE, including further development and dissemination of bioinformatics learning resources through our novel incubators and Faculty Mentoring Networks, vigorously pursuing assessment strategies for our learning resources, connecting learning resources with open educational resource (OER) textbooks, learning more about barriers to bioinformatics implementation for underrepresented groups, and developing future workshops and meetings. About half the participants at the meeting were newcomers to NIBLSE, a positive sign for the future. NIBLSE has many exciting opportunities available, and we welcome life science educators with any level of bioinformatics expertise as new members.
The integration of mitochondrial genome fragments into the nuclear genome is well documented, and the transfer of these mitochondrial nuclear pseudogenes (numts) is thought to be an ongoing evolutionary process. With the increasing number of eukaryotic genomes available, genome-wide distributions of numts are often surveyed. However, inconsistencies in genome quality can reduce the accuracy of numt estimates, and methods used for identification can be complicated by the diverse sizes and ages of numts. Numts have been previously characterized in rodent genomes and it was postulated that they might be more prevalent in a group of voles with rapidly evolving karyotypes. Here, we examine 37 rodent genomes, and an additional 26 vertebrate genomes, while also considering numt detection methods. We identify numts using DNA:DNA and protein:translated-DNA similarity searches and compare numt distributions among rodent and vertebrate taxa to assess whether some groups are more susceptible to transfer. A combination of protein sequence comparisons (protein:translated-DNA) and BLASTN genomic DNA searches detect 50% more numts than genomic DNA:DNA searches alone. In addition, higher-quality RefSeq genomes produce lower estimates of numts than GenBank genomes, suggesting that lower quality genome assemblies can overestimate numts abundance. Phylogenetic analysis shows that mitochondrial transfers are not associated with karyotypic diversity among rodents. Surprisingly, we did not find a strong correlation between numt counts and genome size. Estimates using DNA: DNA analyses can underestimate the amount of mitochondrial DNA that is transferred to the nucleus.
Similarity searching for homologs, typically using the BLAST programs, is the most powerful and widely used strategy for characterizing newly sequenced genomes. Similarity searching is both sensitive – homologs that last shared a common ancestor more than 2 billion years ago are readily identified using protein sequences – and reliable – similarity statistics accurately predict the frequency of false positives. Homology inference is asymmetrical. While homology, or common ancestry, can be reliably inferred from statistically significant excess sequence or structural similarity, non-homology cannot be deduced from a lack significant similarity. Because sequences change faster than structures, homologous proteins with very similar structures often lack excess similarity.
Bioinformatics, a discipline that combines aspects of biology, statistics, mathematics, and computer science, is becoming increasingly important for biological research. However, bioinformatics instruction is not yet generally integrated into undergraduate life sciences curricula. To understand why we studied how bioinformatics is being included in biology education in the US by conducting a nationwide survey of faculty at two- and four-year institutions. The survey asked several open-ended questions that probed barriers to integration, the answers to which were analyzed using a mixed-methods approach. The barrier most frequently reported by the 1,260 respondents was lack of faculty expertise/training, but other deterrents—lack of student interest, overly-full curricula, and lack of student preparation—were also common. Interestingly, the barriers faculty face depended strongly on whether they are members of an underrepresented group and on the Carnegie Classification of their home institution. We were surprised to discover that the cohort of faculty who were awarded their terminal degree most recently reported the most preparation in bioinformatics but teach it at the lowest rate.
Massive amounts of metagenomics data are currently being produced, and in all such projects a sizeable fraction of the resulting data shows no or little homology to known sequences.It is likely that this fraction contains novel viruses, but identification is challenging since they frequently lack homology to known viruses.To overcome this problem, we developed a strategy to detect ORFan protein families in shotgun metagenomics data, using similarity-based clustering and a set of filters to extract bona fide protein families.We applied this method to 17 virus-enriched libraries originating from human nasopharyngeal aspirates, serum, feces, and cerebrospinal fluid samples.This resulted in 32 predicted putative novel gene families.Some families showed detectable homology to sequences in metagenomics datasets and protein databases after reannotation.Notably, one predicted family matches an ORF from the highly variable Torque Teno virus (TTV).Furthermore, follow-up from a predicted ORFan resulted in the complete reconstruction of a novel circular genome.Its organisation suggests that it most likely corresponds to a novel bacteriophage in the microviridae family, hence it was named bacteriophage HFM.Characterization of the human virome is crucial for our understanding of the role of the microbiome in health and disease.The shift from culture-based methods to metagenomics in recent years, combined with the development of virus particle enrichment protocols, has made it possible to efficiently study the entire flora of human viruses and bacteriophages associated with the human microbiome.These methods have led to the discovery of numerous human viruses 1 and human-resident bacteriophages 2 , and have made it possible to characterize the virus content of entire collections of clinical samples 3,4 .Traditional characterization of virome datasets has largely relied on homology-based approaches [5][6][7][8] .These methods can accurately identify sequences from characterized virus families and distant relatives, but they are unable to annotate viral sequences that have little or no sequence similarity to known viruses.Therefore, a substantial fraction of microbiome datasets cannot be classified, despite the recent rapid increase of sequence information in public databases.For instance, the recent identification of the crAssphage 9 illustrates how sequence homology-based methods have failed to recognize a bacteriophage genome constituting 1.7% of all available fecal metagenomic data.In viral-enriched metagenomics datasets, we expect that a fraction of these unclassifiable sequences originate from protein coding segments from unknown viruses and other kinds of "biological dark matter" 9 .Consequently, the detection of coding sequences with no homologs, or ORFans 10 , in such datasets can be a first step towards the discovery of novel viral species, since novel protein sequences can be used as anchors for the characterization of entire viral genomes.
Although bioinformatics is becoming increasingly central to research in the life sciences, bioinformatics skills and knowledge are not well integrated into undergraduate biology education. This curricular gap prevents biology students from harnessing the full potential of their education, limiting their career opportunities and slowing research innovation. To advance the integration of bioinformatics into life sciences education, a framework of core bioinformatics competencies is needed. To that end, we here report the results of a survey of biology faculty in the United States about teaching bioinformatics to undergraduate life scientists. Responses were received from 1,260 faculty representing institutions in all fifty states with a combined capacity to educate hundreds of thousands of students every year. Results indicate strong, widespread agreement that bioinformatics knowledge and skills are critical for undergraduate life scientists as well as considerable agreement about which skills are necessary. Perceptions of the importance of some skills varied with the respondent's degree of training, time since degree earned, and/or the Carnegie Classification of the respondent's institution. To assess which skills are currently being taught, we analyzed syllabi of courses with bioinformatics content submitted by survey respondents. Finally, we used the survey results, the analysis of the syllabi, and our collective research and teaching expertise to develop a set of bioinformatics core competencies for undergraduate biology students. These core competencies are intended to serve as a guide for institutions as they work to integrate bioinformatics into their life sciences curricula.
Relational databases can integrate diverse types of information and manage large sets of similarity search results, greatly simplifying genome-scale analyses. By focusing on taxonomic subsets of sequences, relational databases can reduce the size and redundancy of sequence libraries and improve the statistical significance of homologs. In addition, by loading similarity search results into a relational database, it becomes possible to explore and summarize the relationships between all of the proteins in an organism and those in other biological kingdoms. This unit describes how to use relational databases to improve the efficiency of sequence similarity searching and demonstrates various large-scale genomic analyses of homology-related data. It also describes the installation and use of a simple protein sequence database, seqdb_demo, which is used as a basis for the other protocols. The unit also introduces search_demo, a database that stores sequence similarity search results. The search_demo database is then used to explore the evolutionary relationships between E. coli proteins and proteins in other organisms in a large-scale comparative genomic analysis. © 2017 by John Wiley & Sons, Inc.
Bioinformatics, a discipline that combines aspects of biology, statistics, and computer science, is increasingly important for biological research. However, bioinformatics instruction is rarely integrated into life sciences curricula at the undergraduate level. To understand why, the Network for Integrating Bioinformatics into Life Sciences Education (NIBLSE, “nibbles”) recently undertook an extensive survey of life sciences faculty in the United States. The survey responses to open-ended questions about barriers to integration were subjected to keyword analysis. The barrier most frequently reported by the ~1,260 respondents was lack of faculty training. Faculty at associate’s-granting institutions report the least training in bioinformatics and the least integration of bioinformatics into their teaching. Faculty from underrepresented minority groups (URMs) in STEM reported training barriers at a higher rate than others, although the number of URM respondents was small. Interestingly, the cohort of faculty with the most recently awarded PhD degrees reported the most training but were teaching bioinformatics at a lower rate than faculty who earned their degrees in previous decades. Other barriers reported included lack of student interest in bioinformatics; lack of student preparation in mathematics, statistics, and computer science; already overly full curricula; and limited access to resources, including hardware, software, and vetted teaching materials. The results of the survey, the largest to date on bioinformatics education, will guide efforts to further integrate bioinformatics instruction into undergraduate life sciences education.
Iterative similarity search programs, like psiblast, jackhmmer, and psisearch, are much more sensitive than pairwise similarity search methods like blast and ssearch because they build a position specific scoring model (a PSSM or HMM) that captures the pattern of sequence conservation characteristic to a protein family. But models are subject to contamination; once an unrelated sequence has been added to the model, homologs of the unrelated sequence will also produce high scores, and the model can diverge from the original protein family. Examination of alignment errors during psiblast PSSM contamination suggested a simple strategy for dramatically reducing PSSM contamination. psiblast PSSMs are built from the query-based multiple sequence alignment (MSA) implied by the pairwise alignments between the query model (PSSM, HMM) and the subject sequences in the library. When the original query sequence residues are inserted into gapped positions in the aligned subject sequence, the resulting PSSM rarely produces alignment over-extensions or alignments to unrelated sequences. This simple step, which tends to anchor the PSSM to the original query sequence and slightly increase target percent identity, can reduce the frequency of false-positive alignments more than 20-fold compared with psiblast and jackhmmer, with little loss in search sensitivity.
The FASTA package provides a comprehensive set of similarity searching programs, similar to those provided by the BLAST package, and some additional programs that are not provided by BLAST for searching with short peptides and oligonucleotides. The FASTA programs work with a wide variety of database formats, including mySQL sequence databases. FASTA provides very accurate statistical significance estimates, and is more sensitive than BLASTN when comparing DNA sequences. These protocols describe how to use the FASTA programs to characterize protein and DNA sequences, using protein:protein, protein:DNA, and DNA:DNA comparisons.
The FASTA programs provide a comprehensive set of rapid similarity searching tools ( fasta36 , fastx36 , tfastx36 , fasty36 , tfasty36 ), similar to those provided by the BLAST package, as well as programs for slower, optimal, local, and global similarity searches ( ssearch36 , ggsearch36 ), and for searching with short peptides and oligonucleotides ( fasts36 , fastm36 ). The FASTA programs use an empirical strategy for estimating statistical significance that accommodates a range of similarity scoring matrices and gap penalties, improving alignment boundary accuracy and search sensitivity. The FASTA programs can produce “BLAST‐like” alignment and tabular output, for ease of integration into existing analysis pipelines, and can search small, representative databases, and then report results for a larger set of sequences, using links from the smaller dataset. The FASTA programs work with a wide variety of database formats, including mySQL and postgreSQL databases. The programs also provide a strategy for integrating domain and active site annotations into alignments and highlighting the mutational state of functionally critical residues. These protocols describe how to use the FASTA programs to characterize protein and DNA sequences, using protein:protein, protein:DNA, and DNA:DNA comparisons. © 2016 by John Wiley & Sons, Inc.
The Curriculum Task Force (CTF) of ISCB’s Education Committee seeks to define curricular guidelines for those who educate or train bioinformatics professionals at all career stages. A recent report of the CTF [1] presented a draft set of bioinformatics core competencies, derived from the results of surveys of (1) core facility directors, (2) career opportunities, and (3) existing curricula. Since the publication of its 2014 report, the CTF has focused on the application of the guidelines in varied contexts to identify areas where refinement is needed. As a first step, the task force held an open meeting at the ISMB conference in July 2014. The ideas discussed at the meeting spawned four working groups (WGs), which focus on (i) defining core competencies for specific types and levels of bioinformatics training, (ii) mapping the curriculum guidelines and competencies to existing materials in order to identify the need for development of new materials, and (iii) identifying where revision of the guidelines may be valuable. The CTF is engaging the ISCB community through open WG meetings at ISCB’s official conferences. Thus far, the WGs have convened at the ISCB Great Lakes Bioinformatics Conference (Purdue University, May 2015) and at the ISMB/ECCB Conference (Dublin, Ireland, July 2015). Additionally, the CTF held a workshop at the Annual General Meeting of the Global Organization of Bioinformatics Learning, Education and Training (Cape Town, South Africa, November 2015). Specifically, the draft competencies have been employed in a wide range of activities and contexts (see Table 1 and [2–11]), including the development of new curricula, the analysis of existing curricula, and the creation of new roles involving bioinformatics. These activities have resulted in the identification of several areas where refinement would be useful: Table 1 Summary of the activities of the ISCB Curriculum Task Force. Identify different levels or phases of competency. It would be helpful to define different phases of competency development, or different levels of competency appropriate for distinct roles. Define competency profiles for disciplines that don’t fit into our current silos. Bioengineering provides an illustrative example of a discipline that requires core competency in bioinformatics but does not fit into our current categories. There are almost certainly others. It would be helpful if we could provide some guidance on how to produce ‘hybrid’ competency profiles, perhaps borrowing some competencies from the TF’s core set and others from different disciplines. The LifeTrain initiative (www.lifetrain.eu) [2, 3] is collecting competency profiles for a range of disciplines of relevance to the biomedical sciences and may provide a useful resource kit for this. Broaden the scope of the competency profiles in response to cutting-edge and emerging research. Current areas requiring improvement include incorporating competencies that capture a fundamental understanding of the biological principles central to analyzing biomolecular data, and broadening the user WG to include applications beyond medicine. Provide guidance on the evidence required to assess whether someone has acquired each competency. For undergraduate, Master’s and PhD programs, learning outcomes for each competency, perhaps with examples of appropriate means of assessment, would be valuable. For established professionals who need to assimilate competencies into their working lives, a different approach may be required (such as keeping a portfolio to capture evidence of competency); the CTF should seek guidance from relevant professional bodies, especially in regulated professions such as healthcare. Provide indicative course content or examples of programs that map to the competency requirements. We do not wish to prescribe what course providers should teach or how they should teach it; however, if a course provider is designing a course to meet a specific competency requirement, it may be helpful to find examples of other programs that do this successfully. One way of achieving this is by mapping existing training content to the TF’s competencies. Another way might be to provide an indication, perhaps based on several courses, of the course content that would meet the competency requirements. This would give course providers the freedom to build their own course syllabi without having to reinvent the wheel. Initiatives to collect examples of Creative Commons (or otherwise reusable) course materials will provide an extremely valuable bank of training materials that could be mapped to the core competencies.
The characterization of new genomes based on their protein sets has been revolutionized by new sequencing technologies, but biologists seeking to exploit new sequence information are often frustrated by the challenges associated with accurately assigning biological functions to newly identified proteins. Here, we highlight some of the challenges in functional inference from sequence similarity. Investigators can improve the accuracy of function prediction by (1) being conservative about the evolutionary distance to a protein of known function; (2) considering the ambiguous meaning of "functional similarity," and (3) being aware of the limitations of annotations in functional databases. Protein function prediction does not offer "one-size-fits-all" solutions. Prediction strategies work better when the idiosyncrasies of function and functional annotation are better understood.
BACKGROUND:Protein domains are commonly used to assess the functional roles and evolutionary relationships of proteins and protein families. Here, we use the Pfam protein family database to examine a set of candidate partial domains. Pfam protein domains are often thought of as evolutionarily indivisible, structurally compact, units from which larger functional proteins are assembled; however, almost 4% of Pfam27 PfamA domains are shorter than 50% of their family model length, suggesting that more than half of the domain is missing at those locations. To better understand the structural nature of partial domains in proteins, we examined 30,961 partial domain regions from 136 domain families contained in a representative subset of PfamA domains (RefProtDom2 or RPD2).RESULTS:We characterized three types of apparent partial domains: split domains, bounded partials, and unbounded partials. We find that bounded partial domains are over-represented in eukaryotes and in lower quality protein predictions, suggesting that they often result from inaccurate genome assemblies or gene models. We also find that a large percentage of unbounded partial domains produce long alignments, which suggests that their annotation as a partial is an alignment artifact; yet some can be found as partials in other sequence contexts.CONCLUSIONS:Partial domains are largely the result of alignment and annotation artifacts and should be viewed with caution. The presence of partial domain annotations in proteins should raise the concern that the prediction of the protein's gene may be incomplete. In general, protein domains can be considered the structural building blocks of proteins.