
Artificial intelligence (AI) offers numerous benefits, including increased automation levels, but can also harm individuals and communities, which raises concerns. The field of human-centered AI (HCAI) was created to address human involvement and the consideration of human rights and values involved in the development of AI systems. Document analysis of governance frameworks was used as a research approach to identify components that contribute to human-centered AI solutions. Four authoritative bodies were selected as the principal sources of AI-related principles, standards, and guidelines, based on their recognized authority, global relevance, and comprehensive regulatory scope. Thirty-seven human-centered components were extracted and classified into five classification categories: human-centered values and ethics, user experience and human interaction, data and model governance, technical robustness and system performance, and AI system capabilities and design considerations. The identified components can be used to develop AI solutions that are human-centered and uphold legal integrity, fundamental freedoms, and the principles of democratic governance. Including these components in a formal development methodology can assist in developing AI solutions that are human-centered, free from bias, beneficial to humans, supportive of human capabilities, and aligned with ethical, transparent, reliable, trustworthy, and explainable principles.
As Large Language Models (LLMs) increasingly support and automate business-critical workflows, the need for robust evaluation frameworks becomes paramount. This paper proposes a system-level testing approach designed to assess the performance and reliability of LLM-based applications integrated into enterprise processes. Moving beyond model-centric benchmarks, the framework adopts principles from software engineering, including black-box and end-to-end testing, to evaluate real-world outcomes in retrieval-augmented generation (RAG) systems. It features modular performance indicators such as context precision, hallucination detection, business tonality alignment, and answer correctness, many harnessing the LLM-as-a-Judge methodology. Answer correctness is used as a case study for the design of interpretable performance indicators grounded in concepts from information retrieval while considering the objectives of business and technical stakeholders. Empirical evaluations in four use cases demonstrate how this approach enables organisations to validate not only the accuracy of their systems but also business relevance.
Algorithms often reinforce societal biases and stereotypes. This is especially concerning for minorities, who are disproportionately impacted by it, thereby threatening their further marginalization. Data fundamentalists frame this issue of algorithmic bias as stemming from data bias, indicated by the underrepresentation of some groups (minorities) in the datasets. Consequently, measures adopted to address algorithmic bias have been data-focused. A relatively recent data-focused measure adopted to address this issue is the deployment of what I term artificially generated minorities (AGMs)—synthetic data used to increase the representation of underrepresented groups (minorities) in algorithms’ training datasets. Data fundamentalists make two central claims about AGMs, which I term the representation claim, which holds that AGMs are representative of minorities, and the normative intervention claim, which holds that the deployment of AGMs addresses algorithmic bias. In this paper, I argue that AGMs do not meet these claims, particularly in the context of algorithmic recruitment. First, I demonstrate that AGMs do not capture the experience of historic and systemic oppression, which defines minority status. Hence, I contend that they do not meaningfully represent minorities. Second, I demonstrate that while AGMs facilitate the realization of the futuristic component of an adequate normative intervention, they undermine the reparative component. Thus, I contend that AGMs do not adequately address algorithmic bias. Finally, I briefly highlight that the failure of AGMs to meet these claims indicates that a data-focused framing of algorithmic bias is overly simplistic and does not account for all the complexities involved in the issue of algorithmic bias and its correction, particularly in the context of algorithmic recruitment.
Driver distraction remains a significant contributor to road accidents; however, existing deep-learning detectors either sacrifice accuracy for speed on resource-constrained hardware or lose generality when confronted with unseen data. This work presents an end-to-end single-stage model that, in one forward pass, jointly identifies the driver, links the driver’s body parts, recognises nearby in-cabin objects, and determines whether the driver is distracted. By embedding spatial and semantic relationships directly into its output vector, the model avoids slow post-processing and significantly reduces false alarms caused by objects that merely appear near the driver. Evaluated on five public datasets and an additional real-world collection that were not used for training, the proposed detector boosts mean F1-score by 0.11 ( ≈ 20
Pretraining language models for low-resource languages poses significant challenges due to scarce and poor-quality data, a lack of comprehensive evaluation benchmarks, and often limited computational resources. Research on compute-optimal language modeling typically focuses on scaling up decoder language models efficiently for high-resource languages. While some studies have investigated the down-scaling of encoder language models for low-resource languages, they often prioritize optimizing for computational constraints rather than pretraining text volume constraints. We address this research gap by analyzing the scaling behaviors of encoder language models which use the Replace Token Detection (RTD) and Masked Language Modeling (MLM) objectives under limited pretraining text volumes. By downsampling three different high-resource languages (English, French, Korean) and two low-resource languages (Xhosa and Swahili), we simulate varying degrees of data scarcity and evaluate downstream performance using established benchmarks such as the GLUE benchmark for English, FLUE for French, KLUE for Korean, and MasakhaNEWS for Xhosa and Swahili. Our findings demonstrate that optimal MLM accuracy scales logarithmically with increasing pretraining text volume across these diverse languages. Additionally, our results show that RTD models consistently outperform MLM models in low-resource scenarios, achieving superior downstream performance with pretraining text volumes smaller than 1000MB for downsampled high-resource languages. However, we find that RTD performs worse than MLM for Xhosa and Swahili. We also find that dynamic masking significantly improves MLM accuracy in these settings. Furthermore, our results show that smaller models are more effective for smaller pretraining text volumes, highlighting the importance of adjusting model size according to data availability in order to maximize performance and efficiency.
Deep learning models for instance segmentation have achieved remarkable success, yet their deployment in specialised domains like ecological monitoring is constrained by the prohibitive cost of acquiring high-quality polygonal annotations. This annotation dependency creates a fundamental bottleneck, limiting both scalability and adaptability of these models in real-world conservation scenarios. This paper introduces a data-centric workflow that addresses this challenge through an iterative, multi-stage active learning strategy enhanced with foundation models. The methodology integrates CLIP diversity sampling as an acquisition function with a semi-automated annotation pipeline that combines YOLOv8x6 detection proposals, human-in-the-loop verification, and SAM2-prompted segmentation refinement. A progressive training strategy using YOLOv11s-seg with quality-controlled pseudo-labelling iteratively expands the training dataset while maintaining annotation quality standards. Validated on African Penguin monitoring using Open Images V7 data and independent SANParks field data, experimental results demonstrate that CLIP diversity sampling achieves mAP _50 of 0.82 with 400 training samples (without SAM2 refinement), compared to 0.69 for random sampling; with SAM2-refined annotations, performance reaches 0.88 mAP _50 using the same 400 samples. Cross-domain generalisation on independent SANParks field data achieves 0.81–0.84 mAP _50 . The framework reduces annotation requirements while providing a practical solution for deploying instance segmentation in data-scarce domains.
This paper discusses ELSA (Ethical, Legal, and Social Aspects of technology) as an emerging methodology for transdisciplinary AI research, characterized by anticipatory technology assessment through close collaboration with diverse (societal) stakeholders. We offer a methodological reflection based on a 1,5 year-long case study on public safety and AI in Lombardijen, a neighbourhood in Rotterdam, the Netherlands, where we engaged residents as citizen stakeholders. Lombardijen is paradoxically under-resourced, meaning historically neglected and stigmatized as a ‘problem district’, yet over-researched, i.e. scrutinized by countless researchers who engage in what has been called ‘drive-by’ research – driving by, extracting data, and disappearing, often without benefits for the community. The community’s ensuing alienation from governmental and academic institutions means that citizens’ valuable contextual knowledge is often overlooked in public deliberation on AI. This raises our research question: How can citizens in low-trust neighbourhoods be meaningfully and reciprocally engaged in transdisciplinary AI research, and what does an ELSA approach offer in this regard? The paper details our experiences in Lombardijen respectively from ethical, legal, social, and technological perspectives. We candidly discuss our learnings, (modest) successes and limitations, ultimately emphasizing the importance of situated responsibility as a precondition for transdisciplinary AI research.
This study examines the factors that influence the use of autonomous vehicles (AVs) on traditional road infrastructure in the Western Cape, with a focus on AV compatibility with existing infrastructure. The Diffusion of Innovations Theory serves as a conceptual framework for assessing the potential and challenges of AV implementation. A qualitative approach was employed through semi-structured interviews with experts from the Departments of Mobility, Infrastructure, and Environmental Affairs in the Western Cape. A thematic analysis of the interviews indicates that the Western Cape’s paved roads are mostly suitable for AV implementation. However, some adaptation would be necessary in rural areas with gravel roads and limited connectivity. Findings indicate that environmental and economic factors, such as funding limitations and public purchasing preferences, negatively influence AV adoption. Political advantage, however, may positively influence the diffusion process. Surprisingly, the study suggests that AVs may need to adapt to existing road infrastructure rather than vice versa, which contrasts with the established literature on AV implementation in developed regions.
Target re-identification (re-ID) systems face critical deployment challenges balancing accuracy with computational efficiency in resource-constrained environments. This paper presents a novel framework that integrates Mixture-of-Experts (MoE) with Knowledge Distillation (KD) to effectively leverage pretrained foundation models. The framework employs dynamic expert selection to combine CLIP and ALIGN models, then distils their collective knowledge into a compact student architecture. The experimental evaluation on VeRi-776 and Market-1501 demonstrates 75. 2
This exploratory study synthesizes insights from 14 in-depth expert interviews to examine the state of artificial intelligence (AI) in Africa across six thematic areas: infrastructure development, governance frameworks, cultural preservation, linguistic equity, the startup ecosystem, and youth empowerment. Through systematic qualitative analysis, we identify preliminary challenges, opportunities, and context-specific strategies for responsible AI growth on the continent. As an exploratory investigation, the sample size aligns with established guidelines for thematic saturation in qualitative research. Verbatim transcripts were analyzed using Braun and Clarke’s thematic analysis framework, with sentiment quantification through AI-assisted coding and manual validation. The initial analysis revealed significant AI readiness gaps, with less than 1
Interpreting the complex and multifactorial risk factors driving Cholera outbreaks remains a critical challenge for public health, particularly across diverse environmental and socio-economic contexts. This paper presents an integrated agentic framework that combines explainable machine learning (ML), statistical analysis, and a language model-powered question-answering system to support Cholera risk interpretation and public health decision-making. Using a multi-country dataset spanning 2000–2025, the framework applies three interpretable ML models, Explainable Boosting Machines (EBM), Natural Gradient Boosting (NGBoost), and TabNet, to predict Cholera incidence based on environmental, socio-economic, and infrastructural variables. In parallel, statistical methods including Pearson and Spearman correlation, and multivariate linear regression are used to validate and quantify associations between predictors and disease outcomes. A LangChain-powered agent, implemented with LangGraph, is integrated into the system to interpret model outputs, analyse tabular results, and generate expert-like responses to natural language queries. The agent draws evidence from multiple CSV-based analyses, including feature importance scores, correlation matrices, regression coefficients, and model performance comparisons to provide grounded, interpretable answers and policy recommendations. A Streamlit interface enables interactive exploration of Cholera risk factors by researchers, health professionals, and policy stakeholders. Results show strong agreement among models on key predictors, such as rainfall frequency, stagnant water presence, and open defecation, with statistically significant relationships confirmed through regression analysis. The EBM model achieved the lowest RMSE (0.421), indicating superior predictive performance. This work demonstrates how explainable AI and LLM agents can be combined into a transparent, interpretable, and actionable framework for public health analytics, offering valuable insights into data-driven disease prevention strategies.
Generative artificial intelligence (AI) applications have enhanced democratization of information to the extent that industry professionals can automate routine tasks, gain insights from complex data and execute tasks more efficiently through generation of text, image, and audio content. Although these applications augment human capabilities, there are concerns about veracity of AI prompting, which results in hallucinations that could have dire consequences on clinical workflow of healthcare professionals. The impact of prompting patterns on optimization of clinical workflows at points-of-care remains nascent with limited evidence especially in healthcare sectors of sub-Saharan Africa. This review explores existing literature on how generative AI prompt engineering optimizes clinical workflow of healthcare professionals by adopting Arksey and O’Malley five-stage scoping review framework to analyze peer-reviewed publications. A comprehensive search strategy was conducted in scholarly databases, including PubMed, IEEE Xplore, and Google Scholar between 2019 and 2025. The study highlights AI prompt engineering strategies, how prompting affects clinical and administrative activities, and how limitations of generative AI prompting could be addressed. Evidence of generative AI prompt engineering are limited in SSA while the Global North and China are the most dominant regions in the discourse. Consultations, clinical decision support, record summaries and documentation, research and prescription recommendations are leading activities in which AI prompting is perceived as most significant. To conclude, this study provides insights for health managers, healthcare professionals, data scientists, ethicists, health IT experts, human-computer interaction practitioners, and researchers on standardizing integration of generative AI use at points-of-care.
The environmental impacts of large language models (LLMs) often remain invisible in business adoption. This paper presents an awareness framework to support the sustainable selection of LLMs, developed using a design science research approach within the marketing department of a major European engineering and technology company. Addressing the lack of transparency and emissions data from LLM providers, the artefact calculates electricity use, carbon emissions, and material impacts of inference tasks and visualises them in an interactive dashboard. Evaluation workshops with stakeholders from marketing, sustainability, and AI strategy confirmed the framework’s potential to foster awareness, support sustainable decision-making, and align AI use with corporate environmental goals and the UN Sustainable Development Goals (SDG). The framework is transferable to other business contexts.
The emergence of generative artificial intelligence (GenAI) and its applicability within higher education institutions (HEIs) has gained momentum worldwide. GenAI tools have caused a paradigm shift in education including students’ research activities. Despite various studies being conducted on GenAI tools in education, most research remains concentrated on developed countries, with limited attention to how these technologies are perceived in developing nations. Therefore, this study explores the usage and perceptions of GenAI tools among postgraduate students enrolled for postgraduate diplomas, honours and master’s degrees at a private HEI in South Africa. Using a mixed-methods approach, the study surveyed 75 students to understand their usage and perceptions of GenAI tools for supporting research activities. The findings reveal that almost three-quarters of the students use GenAI tools, particularly ChatGPT, and have a positive attitude towards the use of GenAI tools to support their research activities. The high usage of GenAI tools is attributed to their capability to generate research ideas, summarise articles, and simplify difficult concepts. Over a quarter of the surveyed students do not use GenAI tools due to concerns about plagiarism, bias, privacy and the potential to impair cognitive development. 78
Visual inspection remains a common approach for assessing composite insulators, with unmanned aerial vehicles (UAVs) increasingly preferred due to their efficiency and reduced error rates. Recent developments have integrated artificial intelligence (AI) algorithms directly into UAV hardware to enable faster processing; however, such systems require optimized models owing to limited onboard computing resources. The recently introduced YOLOv10-N model, which offers greater efficiency compared to its predecessors, demonstrates potential for detecting insulator defects on resource-constrained UAV platforms. This study evaluates the effectiveness of YOLOv10-N for this application.
The integration of AI into cybersecurity is essential for addressing complex and evolving threats. However, much of the existing AI implementation research emphasises either technical or social dimensions, neglecting their socio-technical interdependence. This study addresses this gap by identifying the critical success factors (CSFs) for AI-driven cybersecurity implementation through a socio-technical lens. Using an interpretive case study of a South African state-owned entity, the research draws on thematic analysis of in-depth interviews with technical staff and end-users. Findings reveal that successful implementation depends on the interplay between technical and social elements. Key technical CSFs include data quality, scalability, automation, and efficiency, while social CSFs encompass change acceptance, top management support, user awareness, ethical considerations, human oversight, and usability. Crucially, the study confirms that neither technical nor social factors alone are sufficient and that effective implementation depends on their interdependence. By applying a socio-technical perspective, the research offers a more balanced understanding of AI-driven cybersecurity and presents a framework to support practitioners in implementing socially integrated, technically robust solutions. Future research should further examine how human-AI collaboration can be socio-technically integrated to enhance trust, ensure ethical compliance, and improve the operational reliability of AI-enabled cybersecurity systems within organisational settings.
A third of the food produced globally and in South Africa is lost or wasted annually. Fresh fruits and vegetables (FFVs) contribute 44
Cross-lingual language models enable the transfer of linguistic knowledge across languages, however, they often perform worse for low-resource or typologically distant languages. Prior work has explored alignment and adapter methods, but the use of code-switching remains limited and typically confined to fine-tuning with static word substitutions. In this work, we propose an approach that integrates code-switching directly into masked language model pretraining. Instead of applying word substitutions after pretraining, we introduce a multiview probabilistic translation strategy that samples candidate translations based on alignment likelihoods, applying substitutions only to unmasked tokens. This exposes the model to cross-lingual ambiguity and encourages more robust cross-lingual representations. Our results on a diverse set of eight language pairs show that this approach improves zero-shot cross-lingual natural language understanding performance across all languages relative to bilingual baselines. We further observe gains on downstream named entity recognition tasks in most languages when incorporating our code-switched pretraining approach.
While fine-tuning transformer-based pre-trained speech models improves speech recognition for low resource languages, the approach increases the risk of speaker attribute bias in the resulting target language automatic speech recognition (ASR) systems. This work investigates gender bias in two state-of-the-art pre-trained speech models, MMS and Whisper, fine-tuned for ASR on three African languages: Bemba, Nyanja, and Swahili. We fine-tune models on gender-specific as well as gender-balanced datasets, and estimate and compare gender bias across different settings. Our results show varying degrees of gender bias in the fine-tuned models, even with gender-balanced fine-tuning, suggesting influence from pre-trained models. Inconsistencies in gender-specific fine-tuning further confirm the transfer of bias from pre-trained models. Additionally, an ablation study shows no relationship between training data size and gender bias.
Federated Learning (FL) is a distributed learning paradigm which entails the training of Machine Learning (ML) models across multiple computing devices, while keeping the training data local to the devices. One of the key challenges of FL is heterogeneity of both the computing devices and data. This challenge might ultimately lead to the FL model instability, slow convergence, and performance degradation. This work introduces Adaptive FedProx, a new FedProx algorithm extension that dynamically modifies its proximal regularisation term in response to real-time heterogeneity detection. In order to direct adaptive regularisation, we present the Heterogeneity-Aware Performance Index (HAPI), a metric that measures the difference between local and global models. We uncover an important trade-off through extensive experiments on CIFAR-10 across Independent and Identically Distributed (IID), mild non-IID, and strong non-IID scenarios: Adaptive FedProx exhibits superior robustness to data heterogeneity, despite a 1.6 p<0.001 ). When moving from IID to strong non-IID data, Adaptive FedProx shows 3.5