Transformer models have recently become the dominant paradigm for learning representations of complex sequential data. However, their behavior under extreme spatio-temporal structure, multimodality, and data scarcity remains insufficiently understood. This paper presents a structured and critical review of transformer-based approaches for sign language processing, using sign language as a stress test for transformer learning in low-resource, multi-modal, and spatially structured settings. We synthesize and analyze existing work in sign language translation and production, with particular emphasis on sign language production as a generative learning problem that exposes limitations in current modeling assumptions. Our review examines how transformer architectures originally developed for natural language processing and computer vision are adapted to capture spatial–temporal correlations, multi-stream inputs, and cross-modal alignment. We further analyze architectural choices, loss formulations, and evaluation protocols, identifying recurring design patterns as well as systematic shortcomings. Based on this analysis, we highlight open challenges and research opportunities for transformer-based learning in highly structured multimodal domains, with implications extending beyond sign language to other low-resource spatio-temporal learning problems.
The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific knowledge. A core obstacle is the lack of established scientific terminology in these languages. We introduce AfriScience-MT, a parallel corpus covering six African languages (Amharic, Hausa, Luganda, Northern Sotho, Yorùbá, and isiZulu) across 11 scientific domains. Professional translators, working with expert science communicators, translated plain-language summaries of scientific papers into each target language and created new terms where none existed. We benchmark machine translation systems and large language models in zero-shot, few-shot, and fine-tuned settings. Our results show that closed-source models outperform all open-source models at both the sentence and document levels: GPT-5.4 and Gemini-3.1-Flash-Lite lead with average sentence-level COMET scores of 68.3 and 68.0, respectively, and tie at an average document-level COMET of 48.3. Among open systems, fine-tuned NLLB-1.3B reaches 67.3 at the sentence level, and TranslateGemma-12B reaches 44.0 at the document level with 1-shot in-context learning. We release AfriScience-MT to support benchmarking and document-level scientific MT for African languages.
Limited linguistic inclusivity in public health communication leaves many South African communities underserved, particularly regarding critical information on zoonotic diseases such as rabies. This pilot study addresses this gap by developing and evaluating AI-driven methods for delivering reliable rabies information to Sepedi speakers, a low-resource language group. The study presents a novel, curated Sepedi dataset of 60 question–answer pairs, created through a systematic pipeline: thematic analysis of authoritative English sources guided the synthetic generation of QA pairs, which were then translated and manually verified by a native-speaking expert. This dataset was used to compare two large language models, GPT-4o and Gemini-1.5 Flash, under both base and fine-tuned conditions. Evaluation used a human-centred rubric assessing fluency, accuracy, and cultural appropriateness. The findings reveal a key nuance in applying LLMs to low-resource domains. The base GPT-4o model, with strong foundational multilingual capabilities, outperformed all other configurations, including its own fine-tuned variant.In contrast, fine-tuning provided a marked improvement for the less capable base Gemini model. This result indicates that fine-tuning can enhance weaker models; its benefits are not universal and may be outweighed by the strong zero-shot performance of state-of-the-art architectures when training data is scarce. The curated Sepedi rabies QA dataset will be released under an open licence to support future work in low-resource public health communication.
The field of Natural Language Processing (NLP) has achieved significant success in various areas, such as developing large-scale datasets, algorithmic complexity, optimized computing capabilities, refined individual and community expertise, and more, particularly in languages such as English, French, and Spanish. However, such global north unilateral strides have inadvertently created a substantial representation bias towards many languages categorized as low-resourced languages, with the majority being African languages. As a result, rudimentary resources such as stopwords, lemmatizers, stemmers, and word embeddings, as well as advanced multilingual transformer-based models remain under-developed for these languages. Compounding these circumstances is the lack of insights surrounding the development of these resources in the low-resourced context (e.g., how to develop embeddings for morphologically rich languages). Looking back, research priorities aiming to create these resources, largely motivated by the high cost attached to remedying these issues shifted, leading to the rise of alternative methods such as cross-lingual transfer learning (CLTL). CLTL involves transferring domain knowledge gained from supervised training to a domain with limited supervision signals. This study conducts a systematic literature review of CLTL techniques, in the context of cross-lingual models and embeddings, looking at their mathematical foundations, application domains, evaluation metrics, languages covered, and the latest developments. The findings of this study offer valuable insights into the present scenario of CLTL techniques, identifying areas for future research and development to advance cross-lingual natural language processing applications specifically in low-resourced settings.
Food security is a vital aspect of the United Nations’ Sustainable Development Goals (SDGs) which aims to promote sustainable farming in the world. Farming-driven economies such as Ghana are faced with challenges due to plant diseases. Cacao, a vital crop in Ghana is severely impacted with diseases which affect its yield and decrease exports revenue through reduced exports. Leveraging deep learning techniques offers an effective solution for early detection of diseases in cacao plants. This study adopts a comprehensive approach, starting with an Exploratory Data Analysis (EDA) of the dataset containing images of both healthy and diseased cacao plants from Ghanaian farms. Using exploratory data analysis (EDA), we can identify patterns and understand the characteristics of the dataset, laying a solid foundation for developing robust machine learning models tailored to the specific challenges faced by Ghanaian cacao farmers. Our approach involves developing and evaluating deep learning models to detect and classify cacao plant diseases. These models are designed with the Predictability, Compatibility, and Stability (PCS) framework in mind, ensuring reliability and effectiveness in disease detection. The custom convolution neural network (CNN) model outperformed other models considered in experimental analysis. This study aims to revolutionize cacao farming through precise, stable, and ethical deep learning solutions, ultimately enhancing crop resilience, productivity, and the livelihood of Ghanaian farmers.
Disease attacks on crops like maize pose a significant threat to the global food supply chain in Africa, particularly in West Africa. Maize is a staple food source and the economic backbone of the population and farmers in West Africa. In recent years, maize yields have declined due to diseases. Systematic solutions, such as visual inspection through laboratory experiments for disease diagnosis, have not led to a significant improvement in production or ensured sustainable food security in Africa. In response to this challenge, we introduce a lightweight deep-learning ensemble model for early disease detection in maize plants. The study focuses on developing and validating a model specifically designed to identify and classify diseases in maize plants. We use computer vision technology to capture intricate patterns in maize leaf images. The model is trained to recognise six classes, with five representing different diseases and one representing a healthy state. In this paper, we explore the amalgamation of Residual Network (ResNet9) and Efficient-Net-b4 (ENetb4) as a pre-training model built on convolutional neural networks (CNN) to improve accuracy and robustness in maize disease detection and prediction. The results of the study indicate significant opportunities for improvement in agricultural technology in West Africa. For instance, the ResNet9 model accurately identified diseased images of maize crops with a performance accuracy of 98.2%, while the (ENetb4) model achieved a performance accuracy of 94.3% in the same task.
Social media—particularly X (formerly Twitter)—has become a critical platform for political discourse. It shapes public opinion, influences voter behaviour, and provides real-time insight into contentious issues. Xenophobia, defined as the hostility, or hatred towards foreigners, is a polarising topic in South Africa, especially during election seasons. This paper analyses South African Twitter data from the 2016 and 2021 municipal elections, as well as the 2019 and 2024 national elections, with a focus on Xenophobia-related discourse. We develop a novel machine learning model to identify xenophobic tweets despite the removal of explicit hate speech by platform moderation. Using a labelled dataset of xenophobic tweets, we fine-tuned a transformer-based classifier that achieves over 95
The critical lack of structured terminological data for South Africa's official languages hampers progress in multilingual NLP, despite the existence of numerous government and academic terminology lists. These valuable assets remain fragmented and locked in non-machine-readable formats, rendering them unusable for computational research and development. Mafoko addresses this challenge by systematically aggregating, cleaning, and standardising these scattered resources into open, interoperable datasets. We introduce the foundational Mafoko dataset, released under the equitable, Africa-centered NOODL framework. To demonstrate its immediate utility, we integrate the terminology into a Retrieval-Augmented Generation (RAG) pipeline. Experiments show substantial improvements in the accuracy and domain-specific consistency of English-to-Tshivenda machine translation for large language models. Mafoko provides a scalable foundation for developing robust and equitable NLP technologies, ensuring South Africa's rich linguistic diversity is represented in the digital age.
Social media platforms play a significant role in analyzing customer perceptions of financial products and services in today’s culture. These platforms facilitate the immediate and in-depth sharing of thoughts and experiences, offering valuable insights into consumer behaviour. Any customer looking for such a service would surf the internet for reviews and ratings before making a decision, which usually influences their ultimate pick. Feedback and suggestions from friends, family, and coworkers improve customer experiences. Customer reviews play a crucial role in shaping the reputation and profitability of businesses and products offered by financial institutions, often serving as the final assessment of quality and satisfaction during decision-making. Therefore, it is paramount for decision-makers to carefully evaluate customer feedback and understand the sentiment expressed in a given piece of text, which could lead to equity trading, and credit market assessment, and offer invaluable insights that boost the financial performance of the institution. Previous research has used human-annotated text, such as lexicon-based methods, to train machine learning models for sentiment analysis, but the approach did not capture the full range of structure and semantic relationships in natural language. Therefore, our research aims to develop a more comprehensive and accurate sentiment analysis model using advanced natural language processing techniques that could answer questions on various subjects and tasks. To do this, we first crawled customer reviews on Hellopeter, a popular review site, and financial data on the top five financial institutions listed on the Johannesburg Stock Exchange (JSE) in South Africa. After that, we used OpenAI’s ChatGPT as a zero-short learning model to generate human-like annotation tools for different sentiment tasks. The OpenAI ChatGPT feature vector was subsequently fed into BERT, BiLSTM, and a SoftMax function to detect and identify the sentiment of a given sentence. Lastly, we use feature vectors with oversampling methods to address the imbalanced data dilemma and visualise the contribution features of the given piece of text for the customer reviewers. The experiments demonstrated that the method performed as well as or better than the latest and most effective methods on the tested datasets, yielding comparable results. When OpenAI’s ChatGPT was combined with pre-trained BERT and BiLSTM models, it did better overall, with an average score of 98.9%, an F1-measure of 97.7%, and an AUC of 91.90% when oversampling was used. The traditional lexicon-based model got an 86.68% score using SVM and logistic regression and an AUC of 91.90%. The study shows the exceptional performance of OpenAI ChatGPT in detecting the emotional tone or polarity of a given sentence in a customer review, which helps with annotation and understanding the sentiment analysis of an event and how it influences decisions and outcomes. In conclusion, these results underscore the significant advantages of incorporating customer sentiment analysis into financial analysis and decision-making processes as a valuable tool for understanding and prioritizing customer needs and preferences.
The rapid spread of misinformation on platforms like Twitter, and Facebook, and in news headlines highlights the urgent need for effective ways to detect it. Currently, researchers are increasingly using machine learning (ML) and deep learning (DL) techniques to tackle misinformation detection (MID) because of their proven success. However, this task is still challenging due to the complexity of deceptive language, digital editing tools, and the lack of reliable linguistic resources for non-English languages. This paper provides a comprehensive analysis of relevant research, providing insights into advanced techniques for MID. It covers dataset assessments, the importance of using multiple forms of data (multimodality), and different language representations. By applying the Preferred Reporting Items for Systematic Review and Meta-Analysis (PRISMA) methodology, the study identified and analyzed literature from 2019 to 2024 across five databases: Google Scholar, Springer, Elsevier, ACM, and IEEE Xplore. The study selected thirty-one papers and examined the effectiveness of various ML and DL approaches with a focal point on performance metrics, datasets, and false or misleading information detection challenges. The findings indicate that most current MID models are heavily dependent on DL techniques, with approximately 81% of studies preferring these over traditional ML methods. In addition, most studies are text-based, with much less attention given to audio, speech, images, and videos. The most effective models are mainly designed for high-resource languages, with English datasets being the most used (67%), followed by Arabic (14%), Chinese (11%), and others. Less than 10% of the studies focus on low-resource languages (LRLs). Therefore, the study highlighted the need for robust datasets and interpretable, scalable MID models for LRLs. It emphasizes the critical need to prioritize and advance MID research for LRLs across all data types, including text, audio, speech, images, videos, and multimodal approaches. This study aims to support ongoing efforts to combat misinformation and promote a more informed understanding of under-resourced African languages.
Sentiment analysis is a well-known task that has been used to analyse customer feedback reviews and media headlines to detect the sentimental personality or polarisation of a given text. With the growth of social media and other online platforms, like Twitter (now branded as X), Facebook, blogs, and others, it has been used in the investment community to monitor customer feedback, reviews, and news headlines about financial institutions’ products and services to ensure business success and prioritise aspects of customer relationship management. Supervised learning algorithms have been popularly employed for this task, but the performance of these models has been compromised due to the brevity of the content and the presence of idiomatic expressions, sound imitations, and abbreviations. Additionally, the pre-training of a larger language model (PTLM) struggles to capture bidirectional contextual knowledge learnt through word dependency because the sentence-level representation fails to take broad features into account. We develop a novel structure called language feature extraction and adaptation for reviews (LFEAR), an advanced natural language model that amalgamates retrieval-augmented generation (RAG) with a conversation format for an auto-regressive fine-tuning model (ARFT). This helps to overcome the limitations of lexicon-based tools and the reliance on pre-defined sentiment lexicons, which may not fully capture the range of sentiments in natural language and address questions on various topics and tasks. LFEAR is fine-tuned on Hellopeter reviews that incorporate industry-specific contextual information retrieval to show resilience and flexibility for various tasks, including analysing sentiments in reviews of restaurants, movies, politics, and financial products. The proposed model achieved an average precision score of 98.45%, answer correctness of 93.85%, and context precision of 97.69% based on Retrieval-Augmented Generation Assessment (RAGAS) metrics. The LFEAR model is effective in conducting sentiment analysis across various domains due to its adaptability and scalable inference mechanism. It considers unique language characteristics and patterns in specific domains to ensure accurate sentiment annotation. This is particularly beneficial for individuals in the financial sector, such as investors and institutions, including those listed on the Johannesburg Stock Exchange (JSE), which is the primary stock exchange in South Africa and plays a significant role in the country’s financial market. Future initiatives will focus on incorporating a wider range of data sources and improving the system’s ability to express nuanced sentiments effectively, enhancing its usefulness in diverse real-world scenarios.
The world is witnessing a growing epidemic of misinformation. Misinformation can have severe impacts on society across multiple domains: including health, politics, security, the environment, the economy and education. With the constant spread of misinformation on social media networks, a need has arisen to continuously assess the veracity of digital content. This need has inspired numerous research efforts on the development of misinformation detection (MD) models. However, many models do not use all information available to them and existing research contains a lack of relevant datasets to train the models, specifically within the South African social media environment. The aim of this paper is to investigate the transferability of knowledge of a MD model between different contextual environments. This research contributes a multimodal MD model capable of functioning in the South African social media environment, as well as introduces a South African misinformation dataset. The model makes use of multiple sources of information for misinformation detection, namely: textual and visual elements. It uses bidirectional encoder representations from transformers (BERT) as the textual encoder and a residual network (ResNet) as the visual encoder. The model is trained and evaluated on the Fakeddit dataset and a South African misinformation dataset. Results show that using South African samples in the training of the model increases model performance, in a South African contextual environment, and that a multimodal model retains significantly more knowledge than both the textual and visual unimodal models. Our study suggests that the performance of a misinformation detection model is influenced by the cultural nuances of its operating environment and multimodal models assist in the transferability of knowledge between different contextual environments. Therefore, local data should be incorporated into the training process of a misinformation detection model in order to optimize model performance.
In the 1990s, the tools of natural language processing (NLP) underwent a big change, moving from rule-based to statistical-based methods to make computers understand language better (Magueresse et al., 2020). Today’s NLP research mainly focuses on 20 of the 7,000 languages spoken worldwide, which account for more than 95% of the world’s population, leaving the vast majority of African languages unstudied. These languages are often called low-resource languages (LRLs), even though this term is not always clear. To improve language translation in LRLs, NLP research has shifted focus to neural machine translation (NMT) to handle errors such as phoneme substitutions, grammatical structure, and sentence boundaries, all of which pose challenges to NMT robustness (Li et al., 2021). NMT systems have been designed to learn from the language data available in LRLs, creating models that can generalize and handle errors efficiently.
Learning morphologically supplemented embedding spaces using cross-lingual models has become an active area of research and facilitated many research breakthroughs in various applications such as machine translation, named entity recognition, document classification, and natural language inference. However, the field has not become customary for Southern African low-resourced languages. In this paper, we present, evaluate and benchmark a cohort of cross-lingual embeddings for the English-Southern African languages on two classification tasks: News Headlines Classification (NHC) and Named Entity Recognition (NER). Our methodology considers four agglutinative languages from the eleven official South African languages: Isixhosa, Sepedi, Sesotho, and Setswana. Canonical correlation analyses and VecMap are the two cross-lingual alignment strategies adopted for this study. Monolingual embeddings used in this work are Glove (source), and FastText (source and target) embeddings. Our results indicate that with enough comparable corpora, we can develop strong inter-joined representations between English and the considered Southern African languages. More specifically, the best zero-shot transfer results on the available Setswana NHC dataset were achieved using canonically correlated embeddings with Multi-layered perceptron as the training model (54.5% accuracy). Furthermore, our NER best performance was achieved using canonically correlated cross-lingual embeddings with Conditional Random Fields as the training model (96.4% F1 score). Collectively, this study’s results were competitive with the benchmarks of the explored NHC and NER datasets, on both zero-short NHC and NER tasks with our advantage being the use of very minimal resources.
Low-resource languages pose a particularly difficult challenge to neu-ral machine translation (NMT), and there appears to be insufficient machine translation (MT) systems to support African language accessibility. Masakhane Web, an NMT system for African languages, is proposed in this paper. Our approach is an open-source platform that is free, flexible, and produces reasonably accurate translations for African languages. The platform makes use of Masakhane community-trained MT models. It enables users to generate new data by providing feedback on translations, which is then used to retrain the models to improve them. Ultimately, our goal is to create a platform that can provide accurate translations for African languages and make the process of creating MT models easier for those who lack the technical expertise. Furthermore, we include strategies for domain experts to evaluate the system and explain how the platform can be used as a data collection source to improve MT for African languages.
The problem of unveiling the author of a given text document from multiple candidate authors is called authorship attribution. Manifold word-based stylistic markers have been successfully used in deep learning methods to deal with the intrinsic problem of authorship attribution. Unfortunately, the performance of word-based authorship attribution systems is limited by the vocabulary of the training corpus. Literature has recommended character-based stylistic markers as an alternative to overcome the hidden word problem. However, character-based methods often fail to capture the sequential relationship of words in texts which is a chasm for further improvement. The question addressed in this paper is whether it is possible to address the ambiguity of hidden words in text documents while preserving the sequential context of words. Consequently, a method based on bidirectional long short-term memory (BLSTM) with a 2-dimensional convolutional neural network (CNN) is proposed to capture sequential writing styles for authorship attribution. The BLSTM was used to obtain the sequential relationship among characteristics using subword information. The 2-dimensional CNN was applied to understand the local syntactical position of the style from unlabeled input text. The proposed method was experimentally evaluated against numerous state-of-the-art methods across the public corporal of CCAT50, IMDb62, Blog50, and Twitter50. Experimental results indicate accuracy improvement of 1.07%, and 0.96%, on CCAT50 and Twitter, respectively, and produce comparable results on the remaining datasets.
Financial inclusion promises to improve worldwide economies by eradicating poverty. In this study, we use high-dimensional household-level survey data from Eswatini, Namibia, Rwanda, and Madagascar to ascertain the main factors that prevent individuals from accessing financial services. This study uses deep learning techniques to perform dimensionality reduction for feature extraction to overcome the common challenge of obscured interpretation in high-dimensional data analysis. To test the effectiveness of reduced features in measuring financial inclusion within and across countries, different algorithms were evaluated on seven performance metrics. From the results, we observed that for Eswatini, Rwanda, and Madagascar, the catboost algorithm was the best predictor of financial inclusion. In each of these three countries, the top three predictors were quality of financial services, spending and income, and remittances; money management and remittances; financial capacity and e-payments and mobile money; and banking and other non-microfinance institutions, remittances, house information, and well-being through farming. For Namibia, the top three features were bank penetration, general and money management, risk, and mitigation with the random forest as the algorithm for financial inclusion prediction. Lastly, it was established that the general cross-region features were banks and non-banks, income/expenditure and money management, and demographics, with the Gaussian Naiv̈e Bayes as the best algorithm for a generalised prediction of financial inclusion.
Post-authorship attribution is a scientific process of using stylometric features to identify the genuine writer of an online text snippet such as an email, blog, forum post, or chat log. It has useful applications in manifold domains, for instance, in a verification process to proactively detect misogynistic, misandrist, xenophobic, and abusive posts on the internet or social networks. The process assumes that texts can be characterized by sequences of words that agglutinate the functional and content lyrics of a writer. However, defining an appropriate characterization of text to capture the unique writing style of an author is a complex endeavor in the discipline of computational linguistics. Moreover, posts are typically short texts with obfuscating vocabularies that might impact the accuracy of authorship attribution. The vocabularies include idioms, onomatopoeias, homophones, phonemes, synonyms, acronyms, anaphora, and polysemy. The method of the regularized deep neural network (RDNN) is introduced in this paper to circumvent the intrinsic challenges of post-authorship attribution. It is based on a convolutional neural network, bidirectional long short-term memory encoder, and distributed highway network. The neural network was used to extract lexical stylometric features that are fed into the bidirectional encoder to extract a syntactic feature-vector representation. The feature vector was then supplied as input to the distributed high networks for regularization to minimize the network-generalization error. The regularized feature vector was ultimately passed to the bidirectional decoder to learn the writing style of an author. The feature-classification layer consists of a fully connected network and a SoftMax function to make the prediction. The RDNN method was tested against thirteen state-of-the-art methods using four benchmark experimental datasets to validate its performance. Experimental results have demonstrated the effectiveness of the method when compared to the existing state-of-the-art methods on three datasets while producing comparable results on one dataset.