Accurate predictions of crop yields at the farm level are essential for improving agricultural productivity, enhancing food security, and supporting informed decision-making among smallholder farmers. However, conventional field assessments and simple statistical models are often time-consuming, limited in scope, and unable to capture complex interactions among climatic and soil factors. To address these challenges, this paper proposes a machine learning-based model for predicting the productivity of multiple crops, including maize, rice, and beans, using multi-source farm-level data from Tanzania. The dataset integrates climate variables such as temperature and rainfall, soil type, farm size, and crop type. Four ensemble learning models, namely Random Forest, Gradient Boosting, Extreme Gradient Boosting, and Extra Trees, were evaluated using an 80/20 train–test split on 9,897 farm-level records acquired from the Mbeya, Ruvuma, and Songwe regions between 2022 and 2024. Hyperparameter tuning with a fivefold cross-validation was applied to improve model generalization and reduce overfitting. Among the evaluated models, the Extra Trees ensemble achieved the highest performance, with a pooled multi-crop R² of 95%, while crop-specific R² values ranged from 79% to 81% for maize, rice, and beans. These findings demonstrate the potential of the proposed approach to support farm-level cultivation planning and climate adaptation decisions for smallholder farmers.
Nicotiana tabacum is a kind of plant cultivated for its leaves used for manufacturing medicine and cigarettes. With the common name, the Tobacco plant is grown in many countries including China, Indonesia, Malawi and Tanzania just to mention a few. Literatures suggest a technical gap in the proper identification of grade labels for various parts of the plant. In addition, manual grading has resulted in various gaps and biases. To mitigate this, a data-driven grading solution is necessary. However, relevant datasets to train grade classifiers from various countries become of the essence. This article presents images concentrated on tobacco leaf plant position namely Leaf position which normally carries 23 grade labels. Due to high rainfall which swiped away the applied fertilizer on the tobacco plants in the farms, we failed to get images of one grade. Therefore, this research could capture and label 22 grade labels. Images of tobacco leaves based on the tobacco plant position were collected in Tanzania through participatory community research. Canon 5D mark III cameras with 100 mm micro lens were used to take pictures of tobacco leaves based on the tobacco plant position. Domain experts were used for image labelling and cleaning according to tobacco grade labels identified in Tanzania. The dataset carries 49,779 images, which can be used to develop machine learning models for tobacco leaf grade label identification. The collected dataset can be used to train models and enhance the performance of pre-trained models in any country of interest.
Traditional clustering algorithms have often been used to categorize farmers but tend to overlook the underlying reasons for these groupings. Typically, clusters are formed based on common metrics such as dispersal and centrality, which provide limited insights into the relationships among key attributes. This study introduces an innovative approach using pattern and association rules analysis to better understand the characteristics of dairy production clusters. Focusing on Tanzanian smallholder farmers, the research moves beyond identifying clusters to uncovering the hidden relationships within them. Through pattern analysis, the study logically examines the behavioral mechanisms that define these clusters, highlighting service gaps that, if addressed, could enhance smallholder dairy farmers' productivity. Frequent patterns with support ranging from 57% to 93% and confidence levels between 85% and 100% were identified, revealing critical challenges faced by these farmers. For instance, farmers using Artificial Insemination—typically younger or new entrants—face constraints related to farm size, land holdings, fodder production, lack of farmer groups, and insufficient formal training in dairy care. Meanwhile, seasoned farmers deal more with institutional barriers such as limited access to marketplaces, extension services, and distant water sources. The study highlights the diverse challenges faced by different farmer groups and provides strategic recommendations for improving dairy productivity. Enhancing access to formal training, improving fodder production, supporting the formation of farmer groups, and addressing institutional barriers are key actions that could help Tanzanian smallholder dairy farmers increase milk yield and overall productivity.
Uncertainty quantification and sensitivity analysis are essential for improving the modeling and estimation of greenhouse gas emissions in livestock farming to evaluate and reduce the impact of uncertainty in input parameters to model output. The present study is a comprehensive review of the sources of uncertainty and techniques used in uncertainty analysis, quantification, and sensitivity analysis. The search process involved rigorous selection criteria and articles retrieved from the Science Direct, Google Scholar, and Scopus databases and exported to RAYYAN for further screening. This review found that identifying the sources of uncertainty, implementing quantifying uncertainty, and analyzing sensitivity are of utmost importance in accurately estimating greenhouse gas emissions. This study proposes the development of an EcoPrecision framework for enhanced precision livestock farming, and estimation of emissions, to address the uncertainties in greenhouse gas emissions and climate change mitigation.
Deaths are caused by breathing oxygen-deficient air all around the world. Nitrogen gas displaces oxygen in the air, bringing the percentage of oxygen down below 21
Objectives: To identify the hidden patterns in the K-means clustered dataset for the Pangani Basin using the Apriori algorithm through frequent patterns and association rules to enrich cluster characteristics. Methods: Frequent patterns and association rule mining were used to discover the hidden attributes in the K-means clustered dataset. Measures of minimum support ranging from 0.5% to 5% and minimum confidence ranging from 50% to 100% were used to generate a manageable number of rules which were then filtered for redundancy. Lift value >1.0 was used to determine the rule's interestingness while Arules and ArulesViz in R were used to visualize generated rules. Findings: Clusters one to four generated 25, 31, 47, and 49 rules respectively at a minimum confidence of 50% and minimum support of 2% in the first two clusters and 1% in other clusters. Furthermore, water users in cluster one were observed to abstract more water than the three clusters, while their water use fee also reflected on the amount they abstracted. In clusters two and three, water users identified the same amount of water source capacity but differed in the amount requested and water use fee. Water users in cluster four were identified with less water source capacity and fewer amounts abstracted than other clusters. However, their water use fee identified was higher than those in cluster three, with high water source capacity and high amount requested. Such a difference is attributed to the type of water use for cluster three users being domestically supplied through community water supply entities to help villagers access water. In contrast, the water use for users in cluster four is domestic and commercial. Novelty: When aggregated with the clustering observations, the identified association rules mining results provide a broad understanding of water users' characteristics for better water allocation and rationing. Keywords: Association rule, Frequent Patterns, Apriori, Characterization, Pangani Basin
Technology is involved in different sectors to improve service delivery. Habari Node PLC (Public Limited Company), located in Arusha, Tanzania, offers Internet services and various additional ICT-based business solutions. The company has a website that is used to provide information related to the services they provide with their cost. However, the current website is not mobile user- friendly and is not integrated with an electronic payment to pay for those services because the fees are currently paid manually. This study aimed to develop a mobile-based application for e-services and e-payment which will allow the user to access all information related to the services provided by this company and be able to perform e-payment to the subscribes services. The payment will be made through mobile money or credit card, depending on the customer’s choice.
Through a literature review, it has been observed that water scarcity results from increased demand due to population growth, economic progress, and climate change, leading to disparities between required and available water resources. Addressing this challenge requires segmenting water users into homogeneous groups and thoroughly examining their characteristics regarding water utilization to develop efficient and effective water governance strategies. This study employed data-driven multi-model validation techniques to characterize water users in Pangani Basin in Tanzania. The Kmeans, Agglomerative Hierarchical, and Fuzzy C-means clustering algorithms were used to ascertain the efficacy of the characterization. Cluster validation showed that K-means outperformed Agglomerative hierarchy by owning a high Calinski–Harabasz Index and low Davies–Bouldin Index of 692.3 and 1.8, respectively, compared to Agglomerative hierarchy with values of 578.2 and 1.9, respectively. The clustered dataset was tested for prediction accuracy by fitting the logistic regression. K-means showed a prediction accuracy of 98.2% over 97.5% of the Agglomerative Hierarchical method. The four clusters identified were large-scale irrigation water users, moderate irrigation water users, community water supply entities, and domestic water users. We argue that understanding users’ characteristics could efficiently and effectively add value to water governance along the basins.
Objectives: This work aims to contribute towards Tanzanian Central Bank Digital Currency (CBDC) users’ privacy preservation. It proposes the design of a privacy preserving CBDC which might be issued by Tanzania's Central Bank (CB), the Bank of Tanzania (BoT), which is currently in CBDC research phase. The work also aims to contribute to literature, the CBDC research being done by BoT, other CBs and CBDC stakeholders around the world. Methods: By using the Design Science Research (DSR) methodology, a privacy preserving CBDC design suitable for Tanzania was proposed, demonstrated and evaluated. This is the result of existing literature showing that different countries have different CBDC designs due to their differences in contexts and purposes for CBDC issuance. This consequently emphasized the fact that a CBDC design should not be treated as a one-size fits all solution. Findings: As opposed to the existing general and other country specific CBDC designs, we proposed a privacy preserving CBDC design suitable for Tanzania by consulting literature and taking into consideration the Tanzanian context. The design appears to be promising Tanzanian CBDC users’ privacy preservation though further work needs to be done. The work should not only be on practical evaluation of the proposed design but also on other factors impacting the success of CBDC projects. This will consequently further increase the success probability of CBDC projects, hence the potential for practical realization of CBDC project benefits. Novelty: Existing literature has shown that, considering the countries’ differences in context and CBDC issuance purposes, CBDC design should not be treated as a generic solution thereby obliging the need for country-specific CBDC designs. Consequently, the privacy preserving CBDC design suitable specifically for Tanzania consists of and provides an outline of privacy preserving interactions among the identified key Tanzanian CBDC participants or actors. The actors are the BoT, the intermediaries (i.e., other banks and payment service providers), Tanzania’s National Identification Authority (NIDA), financial transactions violation detection engine, and the expected CBDC users. Keywords: Digital currency, database privacy, central bank digital currency, privacy
Data scarcity is a significant challenge in the field of Machine Learning (ML), as data collection can be expensive, time-consuming, and difficult, particularly in developing countries. This challenge is exaggerated on the need to use dataset for livestock disease predictions for early intervention and surveillance. To address this challenge, this paper presents a data synthesis method that has been used to accurately generate new data samples from few real-world data. With much data available to train the ML models, overfitting is eliminated. We present the use of Generative Adversarial Networks mainly the Conditional Tabular Generative Adversarial Network to synthesize categorical data for training machine learning models for prediction of the Pestes des Petits Ruminants (PPR) disease. The results showed that training score became 0.89 and the cross-validation score was 0.87 after synthesized data was used with Random Forest algorithm. The resulting dataset can be used to support the prediction and surveillance of the Pestes des Petits Ruminants (PPR) disease. The proposed method can also be applied to any domain with categorical data, and has the potential to improve the performance of machine learning models with increased data availability.
Applying deep learning models requires design and optimization when solving multifaceted artificial intelligence tasks. Optimization relies on human expertise and is achieved only with great exertion. The current literature concentrates on automating design; optimization needs more attention. Similarly, most existing optimization libraries focus on other machine learning tasks rather than image classification. For this reason, an automated optimization scheme of deep learning models for image classification tasks is proposed in this paper. A sequential-model-based optimization algorithm was used to implement the proposed method. Four deep learning models, a transformer-based model, and standard datasets for image classification challenges were employed in the experiments. Through empirical evaluations, this paper demonstrates that the proposed scheme improves the performance of deep learning models. Specifically, for a Virtual Geometry Group (VGG-16), accuracy was heightened from 0.937 to 0.983, signifying a 73% relative error rate drop within an hour of automated optimization. Similarly, training-related parameter values are proposed to improve the performance of deep learning models. The scheme can be extended to automate the optimization of transformer-based models. The insights from this study may assist efforts to provide full access to the building and optimization of DL models, even for amateurs.
RAHA Beverages Company (RABEC) is one of the banana wine production companies that utilizes fuel in steam production in Arusha-Tanzania, where fuel data conditions such as temperature, pressure, discharge, fuel level, and gas leakage with humidity were a challenge to monitor, which provoked boiler malfunction and plant breakdown. Today, RABEC manually uses a dropping stick into the fuel tank to monitor fuel data conditions which is time-consuming and gives inaccurate readings, inefficiency, fuel economy discrepancy, and accidents. This study aimed to design and develop an IoT-based fuel monitoring system. The flow meter, ultrasonic level, thermistor fuel temperature, humidity, and pressure sensors were used to gather fuel information where the GSM module was employed to send fuel data messages to the operator’s phone. An AT mega 328 microcontroller was used to process and analyse the fuel data and send them to the Thing Speak IoT platform using Wi-Fi connectivity. The results showed that when the fuel level was less than the threshold value, an operator was alerted by a refilling message via GSM technology. A pressure of 0.1psi, a fuel temperature of 120 ℃, and 80% of humidity, the system notifies the operator by an alert message to check the injector pressure and if the fuel–air mixture is perfect. In addition, these data were observed on the LCD and ThingSpeak webpage. To conclude, the developed system proved the best performance with a 99.98% of success rate with high accuracy, security, and efficiency rate compared to the current monitoring system.
Increasing the milk production of small dairy producers is necessary to cover the increase in milk demand in Tanzania. Currently, the population of people in both Tanzania and the world has increased and is predicted to increase more in the year 2050. The use of multilevel association rule mining methods to mine strong patterns among smallholder dairy farmers could help in identifying the best dairy farming practices and increase their milk production by adopting them. This study employed multi-level association rule mining to discover strong rules in three clusters, resulting in three levels of rules in each cluster. These three clusters were high, medium, and low milk producers. Rules were obtained for feeding practices, milk production, and breeding and health practices. These rules represent strong patterns among smallholder dairy farmers that could help them improve their dairy farming practices and have a gradual increase in milk production, from low to medium and from medium to higher milk production. Smallholder dairy producers would be provided with recommendations on their dairy farming practices, using rules based on the cluster to which they belong that could help them achieve higher milk production.
Peste des petits ruminants (PPR) is a viral disease that affects small ruminants and is prevalent in many developing countries, particularly in Africa and Asia. It can spread through direct contact, air, and contaminated feed and water. PPR can result in significant economic losses and has a detrimental impact on small ruminant production and trade. Clinical signs include fever, respiratory distress, and diarrhoea, and prevention is primarily through vaccination with a live attenuated vaccine. In this study, 24 samples were selected, pre-processed and synthesized using the Conditional Tabular Generative Adversarial Networks (CTGAN) model. Feature extraction was performed, revealing difficult_breathing as the most important feature in predicting PPR in ruminants. The study used Random Forest Classifier which was fine-tuned using Bayesian Optimization to attain an accuracy of 91%.
Healthcare services are dependent on health information systems (HIS), which enable health data collection and storage with improved management of healthcare service provision. However, several data security weaknesses in HIS have been identified by various studies in Tanzania. The majority are on data and information exchange leading to breaches of system security. In order to reduce security threats and accomplish security objectives (i.e., confidentiality, integrity, availability), disruptive technologies like blockchain must be used to address security breaches, attacks, and vulnerabilities of HISs. In this chapter, a problem-solving technique with digital innovative solutions for addressing real-world problems was used to arrive at the problem's solution. This chapter then shows how the data security weakness of HISs can be improved using the hyper-ledger blockchain. The system was virtually integrated with an existing HIS. Data storage security was attained through private data collection to meet security goals.
Multi-agent-based modelling and simulation provides an adequate environment to study the real world. This paper presents the use of a multi-agent research and simulation (MARS) framework and model design based on the overview, design concepts, design (ODD) protocol to model and simulate small-scale management strategies that are important for increased milk yield per cow. In reality, strategies for farm management at a small-scale level are purely based on heuristics that cost farmers and lead to inadequate milk yields. A differential assessment of the farming strategies was conducted to yield a data-driven approach for selection of the best strategies, which in turn will optimize investments and increase milk yield. The agent-based modelling and simulation revealed that, the studied strategies based on income, farm, and farmer-based characteristics influenced an increase of up to 7.72 L of milk above the average (12.7 ± 4.89). Generally, there was an increase in milk yield based on the identified evolvement strategies; from a baseline data average milk yield of 12.7 ± 4.89 to simulated milk yield average of 17.57 ± 0.72. Evaluating the agent-based models in real-world scenarios will strengthen the assurance that the identified strategies can move small-scale dairy farmers from low to higher milk producers.
The Arusha Urban Water Supply and Sanitation Authority (AUWSA) is a public utility responsible for ensuring the reliable delivery of clean water to the people of Arusha, Tanzania. However, issues with water bill payments have emerged, which have impacted the AUWSA’s operations and ability to provide a sustainable and affordable water supply service. Late payment of water bills and revenue loss due to the outdated postpayment water metering technology are the major challenges faced by the AUWSA. To address these issues, a study has designed and developed an intelligent GSM-based smart prepaid water meter with improved prepayment methods. The prepaid water meter includes a custom user interface, water flow sensor, automatic token generator, and Unstructured Supplementary Service Data (USSD) application. Customers can easily register; load their water credits using Airtel Money, TigoPesa, or Mpesa; and check their usage. An LCD displays water usage information, and customers receive short message service (SMS) notifications via GSM technology when their tokens are about to expire or have been successfully loaded. Water flows, and the valve stops when credits are exhausted. The system’s administrator (admin) communicates with the meter remotely via a web interface developed using PHP, HTML, JavaScript, and SQL. This interface enables the AUWSA to monitor and manage the meter, which should help reduce revenue losses and improve overall efficiency. Future research should focus on predicting water leakage and cybersecurity using artificial intelligence (AI)/machine learning (ML) technologies.
African countries need to strengthen surveillance and control of arboviral diseases such as dengue due to increased outbreaks and spread of arboviruses. Climatic, socio-environment, and ecological variables influence the spread of dengue fever in Sub-Saharan Africa. This paper presents an Agent-Based conceptual and design model for dengue fever developed using the Multi-Agent Research and Simulation (MARS) framework. The study analyzes dengue fever's spatial distribution and identifies the causal relationship between the disease and its climatic and environmental variables. Agent-based modeling (ABM) was used to comprehend the spatial patterns of variation to determine the ecological association between the observed spatio-temporal variations in dengue fever. The domain and design model of an ABM for the surveillance of dengue fever is presented based on the Overview, Design Concepts, and Details (ODD) protocol. Model input parameters and input data for the study area are also presented. The dengue ABM can be adopted and reused for modeling other diseases and other complex problems from different domains while ensuring that their unique characteristics and appropriate modifications are considered to ensure the model's validity and relevance to the new context.
Tanzania's small-scale dairy industry faces similar challenges to those of other developing nations whereby insufficient infrastructure, outdated technology, and low productivity are serious problems for higher milk yield. Tanzania urgently needs to adopt cutting-edge solutions in order to boost dairy performance. With 3500 households' secondary data and 202 households' primary data from 8 villages throughout the Kilimanjaro and Arusha regions, this chapter demonstrates the use of machine learning (ML) techniques to derive homogeneous production clusters and recommendations for more milk yield among dairy farmers. The likelihood for higher milk yield is demonstrated for various clusters with the use of support, confidence, and lift values of association rules analysis. Finally, the production clusters and recommendations are deployed through a mobile application. Recommendations for future improvement are suggested especially on further deployment of learning recommendations and development of a platform-independent mobile solution.