The African BioGenome Project (AfricaBP) is a Pan-African initiative aimed at improving food systems and biodiversity conservation through genomics while ensuring equitable data sharing and benefits. The Open Institute is the knowledge exchange platform of the AfricaBP which aims to bridge local knowledge gaps in biodiversity genomics and bioinformatics and enable infrastructural developments. In 2024, the AfricaBP Open Institute advanced this mission by organising 31 workshops that attracted more than 3500 registered attendees and trained 380 African researchers in genomics, bioinformatics, molecular biology, sample collections and biobanking, and ethics, across all five African geographical regions involving 40 African and non-African organizations. These workshops provide current understanding on the applications of biodiversity genomics and bioinformatics to the African bioeconomy as well as providing practical and hands-on training in genomics, bioinformatics, molecular biology, gene editing, and sample collection and processing. Here, we provide the current understanding of the applications of biodiversity genomics and bioinformatics to the African bioeconomy through synthetic reviews and presentations, including descriptions of 31 workshops organised as well as three fellowship programs delivered or launched by the AfricaBP Open Institute in collaboration with African and international institutions and industry partners. We review the current national bioeconomy strategies across Africa and the economic impact of sequencing African genomes locally, illustrated by a case study on the proposed 1000 Moroccan Genome Project. Finally, we provide recommendations on how African countries could integrate biodiversity genomics and bioinformatics into national economic plans and bioeconomy strategies.
The African BioGenome Project (AfricaBP) is a Pan-African effort aimed at sequencing the genomes of 105,000 African endemic and indigenous species to support food systems, conservation, and ensure data-sharing and equitable benefits. This effort aligns with the Kunming-Montreal Global Biodiversity Framework (KMGBF), which aims to prevent or mitigate biodiversity loss while facilitating equitable access and benefit-sharing from genetic resources and Digital Sequence Information (DSI) and securing adequate technical and scientific cooperations. The AfricaBP Open Institute for Genomics and Bioinformatics (AfricaBP Open Institute) is the knowledge exchange programme of the AfricaBP which aims to overcome infrastructural barriers through the development of technology and infrastructure. A key component of AfricaBP Open Institute's vision is the establishment of the African Digital Sequence Information Data Bank for Biodiversity and Agriculture (African DSI Data Bank), a federated platform for storing, analyzing, visualizing and sharing genetic data across the African continent. The African DSI Data Bank will address the current fragmentation of DSI across African institutions by linking existing databases and resources while ensuring compliance with regional and global standards. It will use a federated model, leveraging existing (and new) infrastructures across Africa, that allow institutions and countries to retain data sovereignty while adhering to national, regional, and international access and benefit-sharing regulations. Through a proposed Global Access Point (GAP), researchers will be able to gain equitable access to sequence data and genomic metadata via a decentralized network. Furthermore, to understand the current landscape of biodiversity and agricultural DSI databases, analyses, visualization, and data sharing platforms, AfricaBP Open Institute conducted a survey across Africa, and recorded 161 responses. Although the majority of these participants shared common challenges such as limited infrastructure, funding, and capacity building, the overwhelming indication was that they support an African-based DSI platform through an inclusive governance model. Consequently, we describe the proposed roadmap for the creation of an African DSI Data Bank that includes African DSI federated database, visualization, analysis, and sharing platforms, as well as the ethical, legal, social, KMGBF, and sustainability considerations associated with such an infrastructure.
The African BioGenome Project (AfricaBP) Open Institute for Genomics and Bioinformatics aims to overcome barriers to capacity building through its distributed African regional workshops and prioritizes the exchange of grassroots knowledge and innovation in biodiversity genomics and bioinformatics. In 2023, we implemented 28 workshops on biodiversity genomics and bioinformatics, covering 11 African countries across the 5 African geographical regions. These regional workshops trained 408 African scientists in hands-on molecular biology, genomics and bioinformatics techniques as well as the ethical, legal and social issues associated with acquiring genetic resources. Here, we discuss the implementation of transformative strategies, such as expanding the regional workshop model of AfricaBP to involve multiple countries, institutions and partners, including the proposed creation of an African digital database with sequence information relating to both biodiversity and agriculture. This will ultimately help create a critical mass of skilled genomics and bioinformatics scientists across Africa. The African BioGenome Project (AfricaBP) Open Institute for Genomics and Bioinformatics established a series of regional workshops in 2023 to exchange knowledge and overcome barriers, which could serve as a model for other scientific communities.
The African BioGenome Project (AfricaBP) is a Pan-African initiative which aims to improve food systems and conservation through genomics, and ensure data sharing and benefits. The Kunming-Montreal Global Biodiversity Framework (KMGBF) is one of the frameworks of the Convention on Biological Diversity which seeks to reduce threats to biodiversity, ensure sustainable use of biodiversity as well as equitable sharing of benefits. AfricaBP’s objectives and activities are closely aligned with the goals of the KMGBF. However, implementing genomic research in the African context presents unique ethical, legal and social challenges and benefits. Here, we explore the alignment between the AfricaBP and the KMGBF, focusing on the potentials for genomics to drive biodiversity conservation and food security across Africa. We critically examine the ethical, legal, and social implications (ELSI) and related challenges associated with implementing the KMGBF. In response to these challenges, and to strengthen AfricaBP’s capacity to implement the KMGBF goals, we make specific recommendations such as, amongst others, the creation of clear policy and legal frameworks, implement transparent monitoring and reporting mechanisms, and ensure interoperability of key regulatory instruments in biodiversity conservation. We also discuss how AfricaBP integrates the theory of change in its activities to enhance the implementation of the KMGBF by strengthening biodiversity data infrastructure, creating awareness via communication and capacity-building whilst empowering local communities, promoting gender diversity in the African biodiversity genomics landscape, facilitating research and innovation by advancing ethical and legal frameworks, and understanding access and benefit-sharing and KMGBF through roundtable meetings, survey development and analysis.
In 2022, around 54 % of African students were denied student visas to study in the United States (US), compared to 36 % of Asian students and 9 % of European students, despite African immigrants in the US often being more highly educated than the US native-born population. This issue cannot be attributed solely to the dichotomy between the Global North and South in visa regimes, but it is also evident among African nations across regional economic blocs. The African BioGenome Project (AfricaBP) Open Institute for Genomics and Bioinformatics, which aims to overcome barriers to capacity building through its distributed African regional workshops, prioritizes grassroots knowledge exchange and innovation in biodiversity genomics and bioinformatics. In 2023, we orchestrated the implementation of 27 capacity building workshops on biodiversity genomics and bioinformatics, covering 10 African countries across 5 African geographical regions. The AfricaBP Open Institute regional workshops raised awareness of biodiversity genomics and bioinformatics among 3788 registered participants, and trained 408 African scientists in hands-on molecular biology, genomics, and bioinformatics techniques. Here, we discuss the implementation of transformative strategies by deploying the AfricaBP Open Institute multi-country, multi-institution, and multi-partner hybrid regional workshop model, including the proposed creation of an African digital database containing sequence information relating to biodiversity and agriculture.
The aim of the UniProt Knowledgebase is to provide users with a comprehensive, high-quality and freely accessible set of protein sequences annotated with functional information. In this publication we describe enhancements made to our data processing pipeline and to our website to adapt to an ever-increasing information content. The number of sequences in UniProtKB has risen to over 227 million and we are working towards including a reference proteome for each taxonomic group. We continue to extract detailed annotations from the literature to update or create reviewed entries, while unreviewed entries are supplemented with annotations provided by automated systems using a variety of machine-learning techniques. In addition, the scientific community continues their contributions of publications and annotations to UniProt entries of their interest. Finally, we describe our new website (https://www.uniprot.org/), designed to enhance our users' experience and make our data easily accessible to the research community. This interface includes access to AlphaFold structures for more than 85% of all entries as well as improved visualisations for subcellular localisation of proteins.
Africa, a continent of 1.3 billion people, had 326 researchers per one million people in 2018 (Schneegans, 2021; UNESCO, 2022), despite the global average for the number of researchers per million people being 1368 (Schneegans, 2021; UNESCO, 2022). Nevertheless, a strong research community is a requirement to advance scientific knowledge and innovation and drive economic growth (Agnew, et al., 2020; Sianes, et al., 2022). This low number of researchers extends to scientific research across Africa and finds resonance with genomic projects such as the African BioGenome Project (Ebenezer, et al., 2022).The African BioGenome project (AfricaBP) plans to sequence 100,000 endemic African species in 10 years (Ebenezer, et al., 2022) with an estimated 203,000 gigabases of DNA sequence. AfricaBP aims to generate these genomes on-the-ground in Africa. However, for AfricaBP to achieve its goals of on-the-ground sequencing and data analysis, there is a need to empower African scientists and institutions to obtain the required skill sets, capacity and infrastructure to generate, analyse, and utilise these sequenced genomes in-country. The Open Institute is the genomics and bioinformatics knowledge exchange programme for the AfricaBP (Figures 1 & 2). It consists of 10 participating institutions including the University of South Africa in South Africa and National Institute of Agricultural Research in Morocco. It aims to: develop biodiversity genomics and bioinformatics curricula targeted at African scientists, promote and develop genomics and bioinformatics tools that will address critical needs relevant to the African terrain such as limited internet access, and advance grassroot knowledge exchange through outreach and public engagement such as quarterly training and workshops.
The Open Institute of the African BioGenome Project empowers African scientists and institutions with the skill sets, capacity and infrastructure to advance scientific knowledge and innovation and drive economic growth.
Build a major genomics resource on the continent to help breeders and conservationists. Build a major genomics resource on the continent to help breeders and conservationists.
ABSTRACT Euglenoids (Euglenida) are unicellular flagellates possessing exceptionally wide geographical and ecological distribution. Euglenoids combine a biotechnological potential with a unique position in the eukaryotic tree of life. In large part these microbes owe this success to diverse genetics including secondary endosymbiosis and likely additional sources of genes. Multiple euglenoid species have translational applications and show great promise in production of biofuels, nutraceuticals, bioremediation, cancer treatments and more exotically as robotics design simulators. An absence of reference genomes currently limits these applications, including development of efficient tools for identification of critical factors in regulation, growth or optimization of metabolic pathways. The Euglena International Network (EIN) seeks to provide a forum to overcome these challenges. EIN has agreed specific goals, mobilized scientists, established a clear roadmap (Grand Challenges), connected academic and industry stakeholders and is currently formulating policy and partnership principles to propel these efforts in a coordinated and efficient manner.
Abstract teaserEuglenoids show great promise to benefit our world; as biofuels, environmental remediators, anti-cancer agents, robotics design simulators and food nutritional agents, but the absence of reference genomes currently limit realizing these benefits. The Euglena International Network (EIN) (https://euglenanetwork.org/) aims to address these challenges, and is currently seeking formative phase support and funding.Body startOf the nearly 1000 known species of euglenoids (Triemer and Zakryś, 2015), including Euglena gracilis and Rhabdomonas costata, fewer than 2 % have been explored for any level of translational potential in the past 20 years. The absence of reference genomes currently limits biotechnology applications, including the development of efficient tools for genetic manipulation in euglenoids.EIN aims to advance euglenoid science through a creative amalgam of academic institutions, national research institutes and biotechnology industry, to translate and exploit euglenoids through genome sequencing. EIN has defined goals, mobilized scientists, established a clear roadmap (Grand Challenges), connected academic and industry professionals and is currently formulating policy and partnership principles, driven by EIN Executive and Science committees. However, for EIN’s activities to be maintained and durable, long-term support is vital. We call on national and continental funding agencies and research councils, protists and algae communities, and biotechnology and pharmaceutical industries, to embrace, support and fund translational exploitation of these highly valuable organisms.
November 2020 marked 2 y since the launch of the Earth BioGenome Project (EBP), which aims to sequence all known eukaryotic species in a 10-y timeframe. Since then, significant progress has been made across all aspects of the EBP roadmap, as outlined in the 2018 article describing the project’s goals, strategies, and challenges (1). The launch phase has ended and the clock has started on reaching the EBP’s major milestones. This Special Feature explores the many facets of the EBP, including a review of progress, a description of major scientific goals, exemplar projects, ethical legal and social issues, and applications of biodiversity genomics. In this Introduction, we summarize the current status of the EBP, held virtually October 5 to 9, 2020, including recent updates through February 2021. References to the nine Perspective articles included in this Special Feature are cited to guide the reader toward deeper understanding of the goals and challenges facing the EBP. It is urgent that the EBP move forward. The year 2020 marked a global failure in meeting any of the 20 “Aichi goals” for the preservation of wildlife and ecosystems (2). The International Union for Conservation of Nature now counts more than 35,000 (28%) of all surveyed species of plants and animals as threatened with extinction (3). The Earth may lose 50% of its biodiversity by the end of this century if nothing is done to mitigate the anthropogenic factors that drive species to extinction and destroy the health of global ecosystems that sustain human existence (2). Degradation of aquatic and terrestrial ecosystems has continued unabated, and we may soon face the possibility of massive ecosystem collapse on a global scale. Such a collapse would have an enormous impact not only on biodiversity, but also on global political stability, and might ultimately affect the survival … [↵][1]1To whom correspondence may be addressed. Email: lewin{at}ucdavis.edu. [1]: #xref-corresp-1-1
The aim of the UniProt Knowledgebase is to provide users with a comprehensive, high-quality and freely accessible set of protein sequences annotated with functional information. In this article, we describe significant updates that we have made over the last two years to the resource. The number of sequences in UniProtKB has risen to approximately 190 million, despite continued work to reduce sequence redundancy at the proteome level. We have adopted new methods of assessing proteome completeness and quality. We continue to extract detailed annotations from the literature to add to reviewed entries and supplement these in unreviewed entries with annotations provided by automated systems such as the newly implemented Association-Rule-Based Annotator (ARBA). We have developed a credit-based publication submission interface to allow the community to contribute publications and annotations to UniProt entries. We describe how UniProtKB responded to the COVID-19 pandemic through expert curation of relevant entries that were rapidly made available to the research community through a dedicated portal. UniProt resources are available under a CC-BY (4.0) license via the web at https://www.uniprot.org/.
Africa plays a central importance role in the human origins, and disease susceptibility, agriculture and biodiversity conservation. Nigeria as the most populous and most diverse country in Africa, owing to its 250 ethnic groups and over 500 different native languages is imperative to any global genomic initiative. The newly inaugurated Nigerian Bioinformatics and Genomics Network (NBGN) becomes necessary to facilitate research collaborative activities and foster opportunities for skills' development amongst Nigerian bioinformatics and genomics investigators. NBGN aims to advance and sustain the fields of genomics and bioinformatics in Nigeria by serving as a vehicle to foster collaboration, provision of new opportunities for interactions between various interdisciplinary subfields of genomics, computational biology and bioinformatics as this will provide opportunities for early career researchers. To provide the foundation for sustainable collaborations, the network organises conferences, workshops, trainings and create opportunities for collaborative research studies and internships, recognise excellence, openly share information and create opportunities for more Nigerians to develop the necessary skills to exceed in genomics and bioinformatics. NBGN currently has attracted more than 650 members around the world. Research collaborations between Nigeria, Africa and the West will grow and all stakeholders, including funding partners, African scientists, researchers across the globe, physicians and patients will be the eventual winners. The exponential membership growth and diversity of research interests of NBGN just within weeks of its establishment and the unanticipated attendance of its activities suggest the significant importance of the network to bioinformatics and genomics research in Nigeria.
The human genome project, which was completed in 2003, ushered in a new era of scientific applications in medicine and bioscience, and also enhanced the generation of high-throughput data which required laboratory and computational analytical approaches in fields known as genomics and bioinformatics respectively. Internationally, specific advances have been achieved which involved the formation and emergence of strong scientific communities to sustain these technological advancements. On the African continent and regionally, the Human Hereditary and Health in Africa (H3Africa), Biosciences eastern and central Africa - International Livestock Research Institute (BecA - ILRI) Hub, and the Alliance for Accelerated Crop Improvements in Africa (ACACIA), are helping to push some of these advances in human health, biosciences, and agriculture respectively. In Nigeria, we believe that significant advances have also been made by various groups since the human genome project was completed. However, a scientific gathering platform to sustainably enable scientists discuss and update these progresses remained elusive. In this article, we report the First Nigerian Bioinformatics Conference (FNBC) hosted by the Nigerian Bioinformatics and Genomics Network (NBGN) in collaboration with the Nigerian Institute of Medical Research (NIMR). The conference was held from 24th - 26th June, 2019, with the theme: “Bioinformatics in the era of genomics in Africa”. Quantitatively, the conference recorded 195 online registered participants, and up to 186 actual participants; comprising of 8 keynote speakers, 6 invited speakers, 25 oral presenters, 83 poster presenters, and up to 73 non-presenting participants. Attendees with national (up to 179) and international (up to 16) affiliations also participated at the conference. Qualitatively, broad scope of bioinformatics, genomics and molecular biology presentations in biomedicine, health, and biosciences were featured at the conference. We discuss the conference structure and activities, lessons learned, and way forward for future bioinformatics conferences in Nigeria. We further discuss the relevance of the conference which presents an increased visibility for the Nigerian bioinformatics community, positions Nigeria as a dynamic community player within the African bioinformatics space, and provides a platform for national impact through the application and implementation of the benefits of bioinformatics.
Euglena spp. are phototrophic flagellates with considerable ecological presence and impact. Euglena gracilis harbours secondary green plastids, but an incompletely characterised proteome precludes accurate understanding of both plastid function and evolutionary history. Using subcellular fractionation, an improved sequence database and MS we determined the composition, evolutionary relationships and hence predicted functions of the E. gracilis plastid proteome. We confidently identified 1345 distinct plastid protein groups and found that at least 100 proteins represent horizontal acquisitions from organisms other than green algae or prokaryotes. Metabolic reconstruction confirmed previously studied/predicted enzymes/pathways and provided evidence for multiple unusual features, including uncoupling of carotenoid and phytol metabolism, a limited role in amino acid metabolism, and dual sets of the SUF pathway for FeS cluster assembly, one of which was acquired by lateral gene transfer from Chlamydiae. Plastid paralogues of trafficking-associated proteins potentially mediating fusion of transport vesicles with the outermost plastid membrane were identified, together with derlin-related proteins, potential translocases across the middle membrane, and an extremely simplified TIC complex. The Euglena plastid, as the product of many genomes, combines novel and conserved features of metabolism and transport.
Background Photosynthetic euglenids are major contributors to fresh water ecosystems. Euglena gracilis in particular has noted metabolic flexibility, reflected by an ability to thrive in a range of harsh environments. E. gracilis has been a popular model organism and of considerable biotechnological interest, but the absence of a gene catalogue has hampered both basic research and translational efforts. Results We report a detailed transcriptome and partial genome for E. gracilis Z1. The nuclear genome is estimated to be around 500 Mb in size, and the transcriptome encodes over 36,000 proteins and the genome possesses less than 1% coding sequence. Annotation of coding sequences indicates a highly sophisticated endomembrane system, RNA processing mechanisms and nuclear genome contributions from several photosynthetic lineages. Multiple gene families, including likely signal transduction components, have been massively expanded. Alterations in protein abundance are controlled post-transcriptionally between light and dark conditions, surprisingly similar to trypanosomatids. Conclusions Our data provide evidence that a range of photosynthetic eukaryotes contributed to the Euglena nuclear genome, evidence in support of the ‘shopping bag’ hypothesis for plastid acquisition. We also suggest that euglenids possess unique regulatory mechanisms for achieving extreme adaptability, through mechanisms of paralog expansion and gene acquisition.
Euglena gracilis is a well-studied biotechnologically exploitable phototrophic flagellate harbouring secondary green plastids. Here we describe its plastid proteome obtained by high-resolution proteomics. We identified 1,345 candidate plastid proteins and assigned functional annotations to 774 of them. More than 120 proteins are affiliated neither to the host lineage nor the plastid ancestor and may represent horizontal acquisitions from various algal and prokaryotic groups. Reconstruction of plastid metabolism confirms both the presence of previously studied/predicted enzymes/pathways and also provides direct evidence for unusual features of its metabolism including uncoupling of carotenoid and phytol metabolism, a limited role in amino acid metabolism and the presence of two sets of the SUF pathway for FeS cluster assembly. Most significantly, one of these was acquired by lateral gene transfer (LGT) from the chlamydiae. Plastidial paralogs of membrane trafficking-associated proteins likely mediating a poorly understood fusion of transport vesicles with the outermost plastid membrane were identified, as well as derlin-related proteins that potentially act as protein translocases of the middle membrane, supporting an extremely simplified TIC complex. The proposed innovations may be also linked to specific features of the transit peptide-like regions described here. Hence the Euglena plastid is demonstrated to be a product of several genomes and to combine novel and conserved metabolism and transport processes.
Photosynthetic euglenids are major components of aquatic ecosystems and relatives of trypanosomes. Euglena gracilis has considerable biotechnological potential and great adaptability, but exploitation remains hampered by the absence of a comprehensive gene catalogue. We address this by genome, RNA and protein sequencing: the E. gracilis genome is >2Gb, with 36,526 predicted proteins. Large lineage-specific paralog families are present, with evidence for flexibility in environmental monitoring, divergent mechanisms for metabolic control, and novel solutions for adaptation to extreme environments. Contributions from photosynthetic eukaryotes to the nuclear genome, consistent with the shopping bag model are found, together with transitions between kinetoplastid and canonical systems. Control of protein expression is almost exclusively post-transcriptional. These data are a major advance in understanding the nuclear genomes of euglenids and provide a platform for investigating the contributions of E. gracilis and its relatives to the biosphere.