Bioinformatics is now intrinsic to life science research, but the past decade has witnessed a continuing deficiency in this essential expertise. Basic data stewardship is still taught relatively rarely in life science education programmes, creating a chasm between theory and practice, and fuelling demand for bioinformatics training across all educational levels and career roles. Concerned by this, surveys have been conducted in recent years to monitor bioinformatics and computational training needs worldwide. This article briefly reviews the principal findings of a number of these studies. We see that there is still a strong appetite for short courses to improve expertise and confidence in data analysis and interpretation; strikingly, however, the most urgent appeal is for bioinformatics to be woven into the fabric of life science degree programmes. Satisfying the relentless training needs of current and future generations of life scientists will require a concerted response from stakeholders across the globe, who need to deliver sustainable solutions capable of both transforming education curricula and cultivating a new cadre of trainer scientists.
Information and communication technologies (ICTs) new data storage infrastructure, broadband Internet, high speed computing and analytical software tools are radically modifying the way science is conducted and the way the results of research are disseminated. Whilst openness has always been one of the accepted norms of scientific practice, a new paradigm of 'Open Science' is emerging. This encompasses a more collaborative scientific enterprise, open access to scientific data, open access to scientific journals and greater engagement of civil society including industry. In parallel, the availability and scale of data that is available for, and produced by, science has massively increased as has our ability to interrogate and analyse that data. 'Big data' and data driven research is now ubiquitous across all scientific disciplines and is opening up exciting new possibilities for addressing previously inaccessible scientific challenges. Meanwhile, the ability to link data from different sources and fields is providing new insights into the complex global societal challenges.
Demand for training life scientists in bioinformatics methods, tools and resources and computational approaches is urgent and growing. To meet this demand, new trainers must be prepared with effective teaching practices for delivering short hands-on training sessions—a specific type of education that is not typically part of professional preparation of life scientists in many countries. A new Train-the-Trainer (TtT) programme was created by adapting existing models, using input from experienced trainers and experts in bioinformatics, and from educational and cognitive sciences. This programme was piloted across Europe from May 2016 to January 2017. Preparation included drafting the training materials, organizing sessions to pilot them and studying this paradigm for its potential to support the development and delivery of future bioinformatics training by participants. Seven pilot TtT sessions were carried out, and this manuscript describes the results of the pilot year. Lessons learned include (i) support is required for logistics, so that new instructors can focus on their teaching; (ii) institutions must provide incentives to include training opportunities for those who want/need to become new or better instructors; (iii) formal evaluation of the TtT materials is now a priority; (iv) a strategy is needed to recruit, train and certify new instructor trainers (faculty); and (v) future evaluations must assess utility. Additionally, defining a flexible but rigorous and reliable process of TtT 'certification' may incentivize participants and will be considered in future.
EMBL Australia Bioinformatics Resource (EMBL-ABR) is a developing national research infrastructure, providing bioinformatics resources and support to life science and biomedical researchers in Australia. EMBL-ABR comprises 10 geographically distributed national nodes with one coordinating hub, with current funding provided through Bioplatforms Australia and the University of Melbourne for its initial 2-year development phase. The EMBL-ABR mission is to: (1) increase Australia's capacity in bioinformatics and data sciences; (2) contribute to the development of training in bioinformatics skills; (3) showcase Australian data sets at an international level and (4) enable engagement in international programs. The activities of EMBL-ABR are focussed in six key areas, aligning with comparable international initiatives such as ELIXIR, CyVerse and NIH Commons. These key areas-Tools, Data, Standards, Platforms, Compute and Training-are described in this article.
This poster briefly presents the goals and work of GOBLET's Standards Committee, with particular emphasis on the outcomes of a recent survey of standards awareness amongst life scientists in Australia.
Data-driven science is generating biological data at unprecedented rates. To make the information sequestered in the accumulating data accessible to the community, and to represent it such that it can be translated to knowledge, requires the development of rigorous data-stewardship practices and deployment of training. Data science is an emerging field, and training programmes are just being created. Best practices in data curation and in data-science training need to be established. This poster introduces the GOBLET Standards committee members, the institutional members involved and the activities the committee is currently working on.
In the last decade, network approaches have transformed our understanding of biological systems. Network analyses and visualizations have allowed us to identify essential molecules and modules in biological systems, and improved our understanding of how changes in cellular processes can lead to complex diseases, such as cancer, infectious and neurodegenerative diseases. "Network medicine" involves unbiased large-scale network-based analyses of diverse data describing interactions between genes, diseases, phenotypes, drug targets, drug transport, drug side-effects, disease trajectories and more. In terms of drug discovery, network medicine exploits our understanding of the network connectivity and signaling system dynamics to help identify optimal, often novel, drug targets. Contrary to initial expectations, however, network approaches have not yet delivered a revolution in molecular medicine. In this review, we propose that a key reason for the limited impact, so far, of network medicine is a lack of quantitative multi-disciplinary studies involving scientists from different backgrounds. To support this argument, we present existing approaches from structural biology, 'omics' technologies (e.g., genomics, proteomics, lipidomics) and computational modeling that point towards how multi-disciplinary efforts allow for important new insights. We also highlight some breakthrough studies as examples of the potential of these approaches, and suggest ways to make greater use of the power of interdisciplinarity. This review reflects discussions held at an interdisciplinary signaling workshop which facilitated knowledge exchange from experts from several different fields, including in silico modelers, computational biologists, biochemists, geneticists, molecular and cell biologists as well as cancer biologists and pharmacologists.
Throughout history, the life sciences have been revolutionised by technological advances; in our era this is manifested by advances in instrumentation for data generation, and consequently researchers now routinely handle large amounts of heterogeneous data in digital formats. The simultaneous transitions towards biology as a data science and towards a ‘life cycle’ view of research data pose new challenges. Researchers face a bewildering landscape of data management requirements, recommendations and regulations, without necessarily being able to access data management training or possessing a clear understanding of practical approaches that can assist in data management in their particular research domain. Here we provide an overview of best practice data life cycle approaches for researchers in the life sciences/bioinformatics space with a particular focus on ‘omics’ datasets and computer-based data processing and analysis. We discuss the different stages of the data life cycle and provide practical suggestions for useful tools and resources to improve data management practices.
Scientific research relies on computer software, yet software is not always developed following practices that ensure its quality and sustainability. This manuscript does not aim to propose new software development best practices, but rather to provide simple recommendations that encourage the adoption of existing best practices. Software development best practices promote better quality software, and better quality software improves the reproducibility and reusability of research. These recommendations are designed around Open Source values, and provide practical suggestions that contribute to making research software and its source code more discoverable, reusable and transparent. This manuscript is aimed at developers, but also at organisations, projects, journals and funders that can increase the quality and sustainability of research software by encouraging the adoption of these recommendations.
Where does your research data go once you’ve published your paper? Can you do better? Good data management spans all stages of the data life cycle: finding, collecting, integrating, processing, visualising, analysing, publishing, sharing and reusing data and metadata. EMBL Australia Bioinformatics Resource (EMBL-ABR) aims to increase Australia’s capacity to deal with the large heterogeneous data sets now part of modern life science and biomedical research, in line with FAIR principles, so that data are Findable, Accessible, Interoperable and Reusable. In this poster we present recent efforts and international engagement in data life cycle best-practice for Australian bioscience and medical research.
The Global Organisation for Bioinformatics Learning, Education and Training (GOBLET: http://mygoblet.org) was established to provide a global, sustainable support structure to foster international communities of bioinformatics trainers and trainees. The activities of GOBLET are carried out via committees, which have independent but overlapping focus areas. The Learning, Education and Training Committee (LET) primarily focuses on providing resources for bioinformatics trainers. Here we describe some of the recent activities and resources developed by the LET Committee: i) consensus descriptors for training materials to ensure FAIR (Findable, Accessible, Interoperable, Reusable)-compliance. ii) core competencies and guidelines on their use. iii) charting the bioinformatics e-learning landscape, to discover existing resources and develop new materials. For these activities, we partnered with other networks and organisations with similar goals.
Several efforts in postgraduate skilling up across a variety of bioinformatics skills for life scientists have emerged and continue to increase across the globe. Australia has also seen a steady increase in postgraduate bioinformatics training. Here we provide an overview of bioinformatics training in Australia and around the globe and what are the needs in Australia. Do these differ from the rest of the globe? Are there specific challenges when it comes to Australia and if so how can we overcome these?
GOBLET is a global organisation that coordinates, shares and supports bioinformatics training activities worldwide, aiming to plug critical skills gaps, ultimately to facilitate the advancement of health- and life-science research. The focus of GOBLET’s Train-the-Trainer initiative is on setting up effective training courses to help plug known skills gaps, especially in the area of NGS data analysis. This initiative will help to share bioinformatics training expertise, experience and resources; train bioinformatics and life-science specialists; support life-science research; promote collaborations among scientists worldwide; build capacity in developing and developed countries. The programme will consist of workshops, which will take place on different continents (e.g., South America, Africa, Asia) and are expected to co-locate as satellite events to major conferences. Each workshop is organised around two main topics: 1) how to exploit NGS data, and 2) how to set up and deliver excellent training courses. Trainers (members of GOBLET, expert in the field) will teach on a volunteer basis. Workshop participants will commit to replicate the workshops at least once, driving an exponential effect. To deliver this ambitious project, GOBLET is seeking partners and sponsors interested in increasing the provision of bioinformatics training in the area of NGS, either as GOBLET collaborators to customise the programme to specific communities or to fund workshops at given locations. To learn more about this project, see http://www.mygoblet.org/content/fund-raising, contact frc@mygoblet.org, talk to us in GOBLET’s booth and visit this poster!
In recent years, high-throughput technologies have brought big data to the life sciences. The march of progress has been rapid, leaving in its wake a demand for courses in data analysis, data stewardship, computing fundamentals, etc., a need that universities have not yet been able to satisfy-paradoxically, many are actually closing "niche" bioinformatics courses at a time of critical need. The impact of this is being felt across continents, as many students and early-stage researchers are being left without appropriate skills to manage, analyse, and interpret their data with confidence. This situation has galvanised a group of scientists to address the problems on an international scale. For the first time, bioinformatics educators and trainers across the globe have come together to address common needs, rising above institutional and international boundaries to cooperate in sharing bioinformatics training expertise, experience, and resources, aiming to put ad hoc training practices on a more professional footing for the benefit of all.
BioJS is an open source software project that develops visualization tools for different types of biological data. Here we report on the factors that influenced the growth of the BioJS user and developer community, and outline our strategy for building on this growth. The lessons we have learned on BioJS may also be relevant to other open source software projects.
One of the foundations of the scientific method is to be able to reproduce experiments and corroborate the results of research that has been done before. However, with the increasing complexities of new technologies and techniques, coupled with the specialisation of experiments, reproducing research findings has become a growing challenge. Clearly, scientific methods must be conveyed succinctly, and with clarity and rigour, in order for research to be reproducible. Here, we propose steps to help increase the transparency of the scientific method and the reproducibility of research results: specifically, we introduce a peer-review oath and accompanying manifesto. These have been designed to offer guidelines to enable reviewers (with the minimum friction or bias) to follow and apply open science principles, and support the ideas of transparency, reproducibility and ultimately greater societal impact. Introducing the oath and manifesto at the stage of peer review will help to check that the research being published includes everything that other researchers would need to successfully repeat the work. Peer review is the lynchpin of the publishing system: encouraging the community to consciously (and conscientiously) uphold these principles should help to improve published papers, increase confidence in the reproducibility of the work and, ultimately, provide strategic benefits to authors and their institutions.
Data sharing, integration and annotation are essential to ensure the reproducibility of the analysis and interpretation of the experimental findings. Often these activities are perceived as a role that bioinformaticians and computer scientists have to take with no or little input from the experimental biologist. On the contrary, biological researchers, being the producers and often the end users of such data, have a big role in enabling biological data integration. The quality and usefulness of data integration depend on the existence and adoption of standards, shared formats, and mechanisms that are suitable for biological researchers to submit and annotate the data, so it can be easily searchable, conveniently linked and consequently used for further biological analysis and discovery. Here, we provide background on what is data integration from a computational science point of view, how it has been applied to biological research, which key aspects contributed to its success and future directions.