Livestock population data is crucial in researching and modeling animal disease (Sibhat et al., 2017), economic health burden (Muraguri et al., 1998), climate change (Fordyce et al., 2023), antimicrobial resistance calculations (Mulchandani et al., 2023), and other topics. The issue is that there are conflicting population data from the World Organization for Animal Health (WOAH), the Food and Agriculture Organization of the United Nations (FAOSTAT), and others for identical populations over the same period. This can pose a challenge in identifying the most accurate source for research. This paper outlines methodologies for comparing FAOSTAT and WOAH data against each other and against other influences to grade and outline data so that researchers can make informed decisions on which data to use and when. The advantages and disadvantages of multiple methods that compare FAOSTAT and WOAH data with other external sources are discussed. External influences such as economic recessions, and government policies are researched and discussed. These practices were utilized for comparing cattle, pig, chicken, and sheep population data across fifteen countries to gauge which sources appeared more probable over which times. A software tool was created to identify outliers in data trends which can be a starting point for researchers to do their investigations in the future. After the points were acquired from the software, research was performed into historical documentation for influences on livestock. It was found during the investigation of the magnitudes of FAOSTAT and WOAH data that government surveys and research paper populations had the most comparable data. Comparing historical records of natural disasters, market pressures, and government policies explained most trends in the data.
Livestock provide nutritional and socio-economic security for marginalized populations in low and middle-income countries. Poorly-informed decisions impact livestock husbandry outcomes, leading to poverty from livestock disease, with repercussions on human health and well-being. The Global Burden of Animal Diseases (GBADs) programme is working to understand the impacts of livestock disease upon human livelihoods and livestock health and welfare. This information can then be used by policy makers operating regionally, nationally and making global decisions. The burden of animal disease crosses many scales and estimating it is a complex task, with extensive requirements for data and subsequent data synthesis. Some of the information that livestock decision-makers require is represented by quantitative estimates derived from field data and models. Model outputs contain uncertainty, arising from many sources such as data quality and availability, or the user’s understanding of models and production systems. Uncertainty in estimates needs to be recognized, accommodated, and accurately reported. This enables robust understanding of synthesized estimates, and associated uncertainty, providing rigor around values that will inform livestock management decision-making. Approaches to handling uncertainty in models and their outputs receive scant attention in animal health economics literature; indeed, uncertainty is sometimes perceived as an analytical weakness. However, knowledge of uncertainty is as important as generating point estimates. Motivated by the context of GBADs, this paper describes an analytical framework for handling uncertainty, emphasizing uncertainty management, and reporting to stakeholders and policy makers. This framework describes a hierarchy of evidence, guiding movement from worst to best-case sources of information, and suggests a stepwise approach to handling uncertainty in estimating the global burden of animal disease. The framework describes the following pillars: background preparation; models as simple as possible but no simpler; assumptions documented; data source quality ranked; commitment to moving up the evidence hierarchy; documentation and justification of modelling approaches, data, data flows and sources of modelling uncertainty; uncertainty and sensitivity analysis on model outputs; documentation and justification of approaches to handling uncertainty; an iterative, up-to-date process of modelling; accounting for accuracy of model inputs; communication of confidence in model outputs; and peer-review.
The heterogeneity that exists across the global spectrum of livestock production means that livestock productivity, efficiency, health expenditure and health outcomes vary across production systems. To ensure that burden of disease estimates are specific to the represented livestock population and people reliant upon them, livestock populations need to be systematically classified into different types of production system, reflective of the heterogeneity across production systems. This paper explores the data currently available of livestock production system classifications and animal health through a scoping review as a foundation for the development of a framework that facilitates more specific estimates of livestock disease burdens. A top-down framework to classification is outlined based on a systematic review of existing classification methods and provides a basis for simple grouping of livestock at global scale. The proposed top-down classification framework, which is dominated by commodity focus of production along with intensity of resource use, may have less relevance at the sub-national level in some jurisdictions and will need to be informed and adapted with information on how countries themselves categorize livestock and their production systems. The findings in this study provide a foundation for analysing animal health burdens across a broad level of production systems. The developed framework will fill a major gap in how livestock production and health are currently approached and analysed.
With the growth of AI and data modelling, the old saying by George Fuechsel regarding data quality ‘Garbage in, garbage out’ holds more truth than ever. Data Scientists are learning the quality of their models depends on the quality of data. Data used by the Global Burden of Animal Diseases (GBADs) is available to modellers around the world, and the quality of the data provided is important as it is used in modelling disease, greenhouse gas emissions, and more. These are important topics, so the data given to the modellers must be investigated and checked for internal and external inconsistencies. The goal of this paper is to investigate data provided by GBADs to find inconsistencies in the data. Data quality was analysed using a five-year trailing average comparison, the interquartile range for the yearly rates of change, and observing outliers on a normal distribution for the yearly rates of change for livestock populations over time. The normal distribution and interquartile range analysis is an internal data analysis that can find outliers that indicate possible data inconsistencies. The five-year trailing average helps identify external data inconsistencies between sources. Using purpose-built data analysis tools and performing analysis on the data shows there are inconsistencies in the data. The consequences of these findings show that researchers need to be cognisant of the data they are using and need to perform their own analysis before they use it in their models as the data can show incorrect results.
The estimation of the global burden of animal diseases requires the integration of multidisciplinary models: economic, statistical, mathematical and conceptual. The output of one model often serves as input for another; therefore, consistency of the model components is critical. The Global Burden of Animal Diseases (GBADs) Informatics team aims to strengthen the scientific foundations of modelling by creating tools that address challenges related to reproducibility, as well as model, data and metadata interoperability. Aligning with these aims, several tools are under development: a) GBADs'Trusted Animal Information Portal (TAIL) is a data acquisition platform that enhances the discoverability of data and literature and improves the user experience of acquiring data. TAIL leverages advanced semantic enrichment techniques (natural language processing and ontologies) and graph databases to provide users with a comprehensive repository of livestock data and literature resources. b) The interoperability of GBADs'models is being improved through the development of an R-based modelling package and standardisation of parameter formats. This initiative aims to foster reproducibility, facilitate data sharing and enable seamless collaboration among stakeholders. c) The GBADs Knowledge Engine is being built to foster an inclusive and dynamic user community by offering data in multiple formats and providing user-friendly mechanisms to garner feedback from the community. These initiatives are critical in addressing complex challenges in animal health and underscore the importance of combining scientific rigour with user-friendly interfaces to empower global efforts in safeguarding animal populations and public health.
Understanding the global economic importance of farmed animals to society is essential as a baseline for decision making about future food systems. We estimated the annual global economic (market) value of live animals and primary production outputs, e.g., meat, eggs, milk, from terrestrial and aquatic farmed animal systems. The results suggest that the total global market value of farmed animals ranges between 1.61 and 3.3 trillion USD (2018) and is expected to be similar in absolute terms to the market value of crop outputs (2.57 trillion USD). The cattle sector dominates the market value of farmed animals. The study highlights the need to consider other values of farmed animals to society, e.g., finance/insurance value and cultural value, in decisions about the sector’s future.
The Global Burden of Animal Diseases (GBADs) programme will provide data-driven evidence that policy-makers can use to evaluate options, inform decisions, and measure the success of animal health and welfare interventions. The GBADs' Informatics team is developing a transparent process for identifying, analysing, visualising and sharing data to calculate livestock disease burdens and drive models and dashboards. These data can be combined with data on other global burdens (human health, crop loss, foodborne diseases) to provide a comprehensive range of information on One Health, required to address such issues as antimicrobial resistance and climate change. The programme began by gathering open data from international organisations (which are undergoing their own digital transformations). Efforts to achieve an accurate estimate of livestock numbers revealed problems in finding, accessing and reconciling data from different sources over time. Ontologies and graph databases are being developed to bridge data silos and improve the findability and interoperability of data. Dashboards, data stories, a documentation website and a Data Governance Handbook explain GBADs data, now available through an application programming interface. Sharing data quality assessments builds trust in such data, encouraging their application to livestock and One Health issues. Animal welfare data present a particular challenge, as much of this information is held privately and discussions continue regarding which data are the most relevant. Accurate livestock numbers are an essential input for calculating biomass, which subsequently feeds into calculations of antimicrobial use and climate change. The GBADs data are also essential to at least eight of the United Nations Sustainable Development Goals.
The Linked Infrastructure for Networked Cultural Scholarship (LINCS) infrastructure project is converting cultural datasets in a range of forms to Linked Open Data with a view to improving findability, interoperability, and reusability (LINCS (2021)). In converting or extracting granular linked data from researcher source data content alongside object metadata, it differs from large-scale cultural heritage and humanities projects focused on linked open metadata for large collections or aggregations of research objects. The linked data produced by LINCS is granular but uneven in its level of detail, given its provenance in individual scholarly projects and the gaps in historical records. The project thus faces the challenge of creating data robust enough to support not only querying and browsing but also reasoning across graphs generated from different interrelated but uneven datasets. We aim to test the potential of the Semantic Web Rule Language (SWRL) rules and Shapes Constraint Language (SHACL) to generate new relationships from existing ones. Medical researchers are employing SWRL and SHACL to support reasoning with complex patterns in health data (Lezcano et al. (2011); Pareti et al. (2019); Somodevilla et al. (2021)). Similar methods can be applied to cultural data, for instance to move beyond direct links between individual people to generate speculative ”webs of textuality” or ”webs of influence” (second-degree or friend-of-a-friend relationships, for instance) that can then be evaluated by constraints in SHACL to see which parts of these webs are most likely or possible. Because cultural datasets are far from comprehensive, this strategy, may help to ”fill in the gaps” in the historical record. This speculative mode of building on existing data aims to prompt new inquiry into underrepresented activities and figures and suggest ways in which silences in historical data might be tackled computationally. In addition, because most LINCS data is being generated from datasets that were not created with RDF representation in mind, SHACL and SWRL present the potential for improving data quality by weeding out logically impossible statements that are being generated via scripts, including from the XML-encoded prose of the literary historical textbase Orlando: Women’s Writing in the British Isles from the Beginnings to the Present (Brown et al. (2019a, 2021)).
Solving complex global problems involving data and data analysis can require data from both the public and private sectors. The sharing of data has traditionally been restricted to open data. To facilitate the use of both open and private data, a new data-sharing framework has been constructed as an extension to the popular Findable-Accessible-Interoperable-Reusable (FAIR) framework. The "Secure by Design" approach has been taken to define the FAIRS data-sharing framework where S stands for Secure. A Cloud infrastructure architecture is proposed that would allow data brokers to implement FAIRS. This architecture is being constructed for the Global Burden of Animal Diseases (GBADs) to facilitate the sharing of livestock data.
Investments in animal health and Veterinary Services can have a measurable impact on the health of people and the environment. These investments require a baseline metric that describes the burden of animal health and welfare in order to justify and prioritise resource allocation and from which to measure the impact of interventions. This paper is part of a process of scientific enquiry in which problems are identified and solutions sought in an inclusive way. It poses the broad question: what should a system to measure the animal disease burden on society look like and what value would it add? Moreover, it aims to do this in such a way as to be accessible by a wide audience, who are encouraged to engage in this debate. Given that farmed animals, including those raised by poor smallholders, are an economic entity, this system should be based on economic principles. These poor farmers are negatively impacted by disparities in animal health technology, which can be addressed through a mixture of supply-led and demand-driven interventions, reinforcing the relevance of targeted financial support from government and non-governmental organisations. The Global Burden of Animal Diseases (GBADs) Programme will glean existing data to measure animal health losses within carefully characterised production systems. Consistent and transparent attribution of animal health losses will enable meaningful comparisons of the animal disease burden to be made between diseases, production systems and countries, and will show how it is apportioned by people's socio-economic status and gender. The GBADs Programme will produce a cloud-based knowledge engine and data portal, through which users will access burden metrics and associated visualisations, support for decisionmaking in the form of future animal health scenarios, and the outputs of wider economic modelling. The vision of GBADs, strengthening the food system for the benefit of society and the environment, is an example of One Health thinking in action.
This session explores how Complex Adaptive Systems provide a framework for analyzing important social, biological, and environmental systems in One Health. Anthropogenic disturbances, many of which are technological, pose a threat to key ecological and sociological processes. They lead us to consider questions such as: Is artificial intelligence a saviour or a demon? What are the political, ethical, and scientific implications for One Health? How might the Global Burden of Disease (human), the Global Burden of Animal Diseases (GBADs) and other Global Burdens constitute a broader "One Health Burdens of Disease" and provide an evidence-base for One Health decisions? It will be necessary to address different data challenges in the developed and developing worlds, many of which are ethical and political, not just technical. Panelists will discuss the GBADs approach to data sharing, including how FAIR-principled metadata can be used to create trustworthy data systems and how the Data Governance Handbook provides important guidance for communicating data sharing principles to data contributors and users. Each panelist will provide a 5-10 minute "primer" talk which will introduce and link the key themes. This will be followed by a moderated panel discussion with opportunities for the audience to pose questions.
A consistent and comparable description of animal diseases, the risk factors associated with them, and the effectiveness of intervention strategies to mitigate these diseases are important for decision making and planning. The economic impact of a pathogen or animal disease is a function of disease frequency, infection intensity, the effect of the disease on mortality and productivity in animals and its effects on human health, and efforts to respond to the disease.1 All of these factors can vary over time between species and the contexts in which people and animals live, and need to be measured to understand the patterns of impact at local, national, and global levels.
This work reports on an ontological organization (framework) that separates domain knowledge from knowledge of specific views and formalizes conceptual relationships by linking to the meta-ontology structure. We use parameters of animal disease spread simulation models as an example, although all concepts presented could apply to human disease spread simulation as well. A meta-ontology is created to document parameter concepts in different comparable simulation models. It formalizes relationships between parameter concepts. This offers several advantages such as allowing explicit domain knowledge representation and provenance, allowing for the assessment of parameters with respect to domain knowledge, and assisting in usage and evaluation of the models. The meta-ontology allows views about parameter concepts to be captured. This is important because it establishes a neutral view point which allows the assessment of parameter semantics in respect to documented domain knowledge. While this work uses the domain of animal disease spread, the principles of ontological representation of model parameters is applicable to a wide range of domains.
OBJECTIVE To evaluate mean body weight (BW) over the lifespan of domestic cats stratified by breed and sex (including reproductive status [neutered vs sexually intact]). ANIMALS 19,015,888 cats. PROCEDURES Electronic medical records from veterinary clinics in the United States and Canada from 1981 to 2016 were collected through links to practice management software programs and anonymized. Age, breed, sex and reproductive status, and BW measurements and measurement dates were recorded. Data were cleaned, and descriptive statistics were determined. Linear regression models were created with data for 8-year-old domestic shorthair, medium hair, and longhair (SML) cats to explore changes in BW over 3 decades (represented by the years 1995, 2005, and 2015). RESULTS 9,886,899 of 19,015,888 (52%) cats had only 1 BW on record. Mean BW for cats of the 4 most common recognized breeds (Siamese, Persian, Himalayan, and Maine Coon Cat) peaked between 6 and 10 years of age and then declined. Mean BW of SML cats peaked at 8 years and was subjectively higher for neutered than for sexually intact cats. Mean BW of neutered 8-year-old SML cats increased between 1995 and 2005 but was steady between 2005 and 2015. CONCLUSIONS AND CLINICAL RELEVANCE The large dataset for this study yielded useful information on mean BW over the lifespan of domestic cats. This could be a basis for BW management discussions during veterinary visits. A low frequency of repeated BW measurements suggested a low frequency of repeated veterinary visits, especially after 1 year of age, making engagement of cat owners in the health of their animals particularly relevant.
AbstractResearch in big data, informatics, and bioinformatics has grown dramatically (Andreu-Perez J, et al., 2015, IEEE Journal of Biomedical and Health Informatics 19, 1193–1208). Advances in gene sequencing technologies, surveillance systems, and electronic medical records have increased the amount of health data available. Unconventional data sources such as social media, wearable sensors, and internet search engine activity have also contributed to the influx of health data. The purpose of this study was to describe how ‘big data’, ‘informatics’, and ‘bioinformatics’ have been used in the animal health and veterinary medical literature and to map and chart publications using these terms through time. A scoping review methodology was used. A literature search of the terms ‘big data’, ‘informatics’, and ‘bioinformatics’ was conducted in the context of animal health and veterinary medicine. Relevance screening on abstract and full-text was conducted sequentially. In order for articles to be relevant, they must have used the words ‘big data’, ‘informatics’, or ‘bioinformatics’ in the title or abstract and full-text and have dealt with one of the major animal species encountered in veterinary medicine. Data items collected for all relevant articles included species, geographic region, first author affiliation, and journal of publication. The study level, study type, and data sources were collected for primary studies. After relevance screening, 1093 were classified. While there was a steady increase in ‘bioinformatics’ articles between 1995 and the end of the study period, ‘informatics’ articles reached their peak in 2012, then declined. The first ‘big data’ publication in animal health and veterinary medicine was in 2012. While few articles used the term ‘big data’ (n = 14), recent growth in ‘big data’ articles was observed. All geographic regions produced publications in ‘informatics’ and ‘bioinformatics’ while only North America, Europe, Asia, and Australia/Oceania produced publications about ‘big data’. ‘Bioinformatics’ primary studies tended to use genetic data and tended to be conducted at the genetic level. In contrast, ‘informatics’ primary studies tended to use non-genetic data sources and conducted at an organismal level. The rapidly evolving definition of ‘big data’ may lead to avoidance of the term.
Large amounts of animal health data are available to researchers, but are often stored in different formats and information silos. Analysis of this existing information can provide new insights into the health and welfare of animals and possibly reduce the need to collect additional data. The objective of this study was to develop a method of managing and analyzing large amounts of data on a personal computer that can be run within 24 h to limit the time and resources spent deploying models on larger servers. This paper describes an overall approach that makes use of existing methods for data acquisition and modeling, but adapts and combines them in a way that allows manipulation and analysis of large volumes of data on a PC. This included a total of five steps: removing errors; removing data points outside the scope of a specific hypothesis; creating descriptive statistics; developing explanatory and/or predictive models; and assessing the fit or accuracy of the models created. The approach was developed using electronic medical records for 19,416,753 feline patients from 3972 anonymized veterinary clinics in the United States and Canada, recorded between January 1981 and June 2016. Data regarding patient signalment (age, sex, breed, reproductive status) and body weight were extracted from the records and used to create linear regression models to describe body weight in cats of different ages, breeds, genders and reproductive status. Ordinary least squares linear regression and stochastic gradient descent linear regression were compared to determine their effectiveness and suitability for creating predictive models with large datasets, using 10 fold cross validation. This approach could be used to build workflows to create models to determine exploratory and predictive properties of health parameters for animals and people. The ability to work with large datasets on a PC or equivalent technology was demonstrated. Significant interactions were present among sex, reproductive status and age. A peak in weight occurred between 6 and 9 years depending on the sex, reproductive status and breed. The predictive ability of the two models was similar, with both producing a root mean square error of 1.45 and a mean absolute error of 1.09, and mean error that was approximately zero on the validation dataset.
Objective: Our objective was to assess the suitability of the data collected by the Animal Poison Control Center, run by the American Society for the Prevention of Cruelty to Animals, for the surveillance of toxicological exposures in companion animals in the United States.Introduction: There have been a number of non-infectious intoxication outbreaks reported in North American companion animal populations over the last decade1. The most devastating outbreak to date was the 2007 melamine pet food contamination incident which affected thousands of pet dogs and cats across North America1. Despite these events, there have been limited efforts to conduct real-time surveillance of toxicological exposures in companion animals nationally, and there is no central registry for the reporting of toxicological events in companion animals in the United States. However, there are a number of poison control centers in the US that collect extensive data on toxicological exposures in companion animals, one of which is the Animal Poison Control Center (APCC) operated by the American Society for the Prevention of Cruelty to Animals (ASPCA). Each year the APCC receives thousands of reports of suspected animal poisonings and collects extensive information from each case, including location of caller, exposure history, diagnostic findings, and outcome. The records from each case are subsequently entered and stored in the AnTox database, an electronic medical record database maintained by the APCC. Therefore, the AnTox database represents a novel source of data for real-time surveillance of toxicological events in companion animals, and may be used for surveillance of pet food and environmental contamination events that may negatively impact both veterinary and human health.Methods: Recorded data from calls to the APPC were collected from the AnTox database from January 1, 2005 to December 31, 2014, inclusive. Sociodemographic data were extracted from the American 2010 decennial census and the American Community Surveys. Choropleth maps were used for preliminary analyses to examine the distribution of reporting to the hotline at the county-level and identify any “holes” in surveillance. To further identify if gaps in reporting were randomly distributed or tended to occur in clusters, as well as to look for any predictable spatial clusters of high rates of reporting, spatial scan statistics, based on a Poisson model, were employed. We fitted multilevel logistic regression models, to account for clustering within county and state, to identify factors (e.g., season, human demographic factors) that are related to predictable changes in call volume or reporting, which may bias the results of quantitative methods for aberration/outbreak detection.Results: Throughout the study period, over 40% of counties reported at least one call to the hotline each year, with the majority of calls coming from the Northeast. Conversely, there was a large “hole” in coverage in Midwestern and southeastern states. The location of the most likely high and low call rate clusters were relatively stable throughout the study period and were associated with socioeconomic status (SES), as the most likely high risk clusters were identified in areas of high SES. Similar results were identified using multivariable analysis as indicators of high SES were found to be positively associated with rates of calls to the hotline at the county-level.Conclusions: Socioeconomic status is a major factor impacting the reporting of toxicological events to the APCC, and needs to be accounted for when applying cluster detection methods to identify outbreaks of mass poisoning events. Large spatial gaps in the network of potential callers to the center also need to be recognized when interpreting the spatiotemporal results of analyses involving these data, particularly when statistical methods that are highly influenced by edge effects are used.
Conventional image processing techniques have been applied to the field of agricultural machine vision for the purposes of identifying crops for quality control, weed detection, automated spraying and harvesting. With the recent advancements in computational hardware Region-based Convolutional Networks have met with varying levels of success in the area of object detection and classification. In this study we found that a Region-based Convolutional Neural Network was able to achieve a 92% accuracy rating while a Region-based Fully Convolutional Network was able to achieve an 87% accuracy rating in the area of object detection operating on a newly create agricultural mushroom dataset.