Society is undergoing rapid transformation through the integration of big data and artificial intelligence (AI). In this context, the dairy sector must address pressing challenges in efficiency, sustainability and data management by adopting intelligent, scalable and privacypreserving technologies. With global dairy demand rising alongside population growth, the adoption of big data and AI has become essential to ensure operational efficiency, informed decisionmaking and the capacity to meet future needs. This paper presents an integrated, multimodal AI framework designed to support dataintensive dairy farm operations by applying big data principles and advancing them through AI technologies. It extends existing literature by outlining collaborative pathways toward a more resilient, efficient and sustainable dairy future, while introducing two additional dimensions – vulnerability and vigilance – to the classic “5Vs” of big data, emphasizing the importance of security and integrity. The proposed architecture combines edge computing, autonomous AI agents, and federated learning to deliver realtime, privacypreserving analytics at the farm level, while enabling knowledge sharing and refinement through research farms and cloudbased collaboration. This architecture supports data integrity, scalability and realtime personalization, while fostering partnerships between farms, research institutions and regulatory bodies to advance secure, crosssector innovation.
Animal behavior analysis is central to understanding welfare, health, and productivity in livestock, yet manual observation is time-consuming, subjective, and difficult to scale. We present a modular pipeline that integrates open-source, state-of-the-art computer vision models to automate individual-level behavior analysis in group-housing environments. The pipeline does not introduce a new learning algorithm or training paradigm; instead, it combines existing zero-shot object detection, motion-aware segmentation and tracking, and vision-transformer feature extraction into a reproducible end-to-end workflow that isolates each animal from its background before behavior is recognized. This individual-level, background-independent design directly targets the poor cross-environment generalization that limits group-level approaches, and it addresses challenges such as occlusion and crowding in indoor pig monitoring. We validated the system on the Edinburgh Pig Behavior Video Dataset across detection, tracking, and behavior-classification tasks. A temporal model achieved 94.2% overall accuracy on nine behaviors, a 21.2 percentage-point improvement over the previous benchmark, while tracking reached 93.3% identity preservation (IDF1) and detection reached 89.3% average precision. The same pipeline, unchanged in structure, transfers across species and tasks: companion studies report 97.6% accuracy for calf play behavior and 98.3% for dairy-cow posture. By releasing an open-source, end-to-end implementation, this work provides a reproducible and scalable tool for automated, objective, and continuous behavior monitoring in precision livestock farming and welfare assessment.
Despite the global adoption of the FAIR guiding principles, significant barriers remain in their practical implementation due to the lack of domain-specific standards for “rich metadata”. While generic metadata schemes provide basic information, they often fail to capture the complexity of data quality and the nuances of specialized scientific fields. This gap is particularly evident in veterinary epidemiology, a field characterized by complex, multi-scale data spanning various domains where detailed description of data analysis and modelling tend to be prioritized over raw data description, leading to inconsistent reporting and low data reusability. This paper introduces the first community-developed rich metadata guidelines specifically designed for veterinary epidemiology datasets, accompanied by practical templates and real-world examples. By integrating existing standards with specific guidance on domain-specific attributes and data quality, these guidelines provide a comprehensive, single-source framework accessible to researchers across academic, public, and private sectors. They aim to support researchers and stakeholders in enhancing the reusability of animal health data.
Animal behavior analysis is central to understanding welfare, health, and productivity in livestock, yet manual observation is time-consuming, subjective, and difficult to scale. We present a modular pipeline that integrates open-source, state-of-the-art computer vision models to automate individual-level behavior analysis in group-housing environments. The pipeline combines zero-shot object detection, motion-aware segmentation and tracking, and vision-transformer feature extraction for robust behavior recognition, addressing challenges such as occlusion and crowding in indoor pig monitoring. We validated the system on the Edinburgh Pig Behavior Video Dataset across detection, tracking, and behavior-classification tasks. A temporal model achieved 94.2% overall accuracy on nine behaviors, a 21.2 percentage-point improvement over the previous benchmark, while tracking reached 93.3% identity preservation (IDF1) and detection reached 89.3% average precision. The modular design lets individual components be replaced as methods advance and suggests potential for adaptation to other species and settings, although cross-species use would require further validation. By releasing an open-source, end-to-end implementation, this work provides a reproducible and scalable tool for automated, objective, and continuous behavior monitoring in precision livestock farming and welfare assessment.
Bovine respiratory disease (BRD) remains one of the most impactful health challenges in cattle, driving antimicrobial use, economic losses, and welfare concerns. Effective control requires insight into circulating pathogens. However, differences in sampling strategies, diagnostic methods, and data formats hinder cross-country interpretation. During a European stakeholder consultation, farmers called for simple tools that clearly communicate trends in pathogen occurence, while veterinarians stressed the need for data-driven dashboards that turn complex data into practical insights to support evidence-based advice and farmer–veterinarian collaboration. To address these priorities, we aimed to develop and evaluate the Cattle Barometer: an interactive, cross-country dashboard integrating laboratory-confirmed BRD pathogen data from multiple European diagnostic laboratories. Diagnostic laboratories in Europe provided anonymised test results from cattle with respiratory disease. Data included identified pathogens, sample type, diagnostic method, date, and location proxies. Heterogeneous datasets were harmonised using scripted pipelines, including cleaning, recoding, anonymisation, and transformation into standardised formats. Farm-level monthly pathogen positivity was calculated and incorporated into a dashboard. Usability of the first prototype was evaluated through a survey (n = 25), including a System Usability Scale (SUS) assessment, open-ended questions, and task-based testing among veterinarians, researchers, laboratory scientists, advisors, and farmers. Fourteen datasets from six laboratories, representing over 8000 farms sampled between 2016 and 2024, were integrated into one dataset. The dashboard enables interactive analysis of trends in proportion of positive samples by time, location, production type, and diagnostic methodology. The survey showed the tool to be promising, but still requiring refinement to achieve optimal user-friendliness (SUS score 60, 60
Livestock farming faces persistent challenges in animal health management, particularly in the surveillance and management of infectious diseases in terrestrial and aquatic species. These diseases affect productivity, economic sustainability, and food security. While smart agriculture and precision livestock farming (PLF) generate large volumes of animal health data, issues such as data fragmentation, poor interoperability, security concerns, and low farmer adoption limit their use. Ontologies as explicit representations of domain knowledge offer a promising way to standardize and integrate heterogeneous data. However, existing literature lacks a comprehensive analysis of their applicability and limitations in livestock disease surveillance. This examines data integration challenges, the role of ontologies, and their limitations in covering livestock diseases. A systematic literature review was conducted following PRISMA 2020 guidelines and supported by a machine learning–based screening tool (ASReview) to ensure transparency, reproducibility, and efficiency in identifying relevant literature. Ontology-based and non-ontology-based approaches were reviewed, with ontologies categorized as active or inactive and assessed for scope, availability, species coverage, and terminological depth. A total of 286 records were screened, of which 100 were included in the final review. Among 32 identified ontologies, 15 remain active while 17 are inactive or no longer publicly accessible, reducing practical use. Active ones often lack full disease coverage across species. Common challenges include system complexity, maintenance, low adoption, and limited domain representation. The review also discusses initiatives such as DECIDE, which illustrate how ontology-driven surveillance can be strengthened through open access, training, and collaborative tools. These findings highlight the urgent need to improve interoperability and develop ontology-driven surveillance systems for livestock.
Following a significant increase in herd and farm sizes after the removal of milk quotas in Europe, the past 10 years have seen a slight yet steady decline in the population of cattle. This includes a reduction of approximately 5% in dairy and beef cattle. This trend is driven by various factors, such as changing market demands, economic shifts, and sustainability challenges in the livestock sector. Despite this, technological advancements in reproductive management have continued to enhance efficiency and sustainability, particularly in dairy production. The main areas of rapid development, which will continue to grow for improving fertility and management, include: i) genetic selection (including improved phenotypes for use in breeding programs), ii) nutritional management (including transition cow management), iii) control of infectious disease, iv) rapid diagnostics of reproductive health, v) development of more efficient ovulation/estrous synchronization protocols, vi) assisted reproductive management (and automated systems to improve reproductive management), vii) increased implementation of sexed semen and embryo transfer, viii) more efficient handling of substantial volumes of data, ix) routine implementation of artificial intelligence technology for rapid decision-making at the farm level, x) climate change and sustainable cattle production awareness, xi) new (reproductive) strategies to improve cattle welfare, and xii) improved management and technology implementation for male fertility. This review addresses the current status and future outlook of key factors that influence cattle herd health and reproductive performance, with a special focus on dairy cattle. These insights are expected to contribute to improved performance, health, and fertility of ruminants in the next 20 years.
Footpad lesions (FPL) are a prevalent welfare concern in broilers, influenced by various factors such as farm management practices and season. In the Netherlands, FPL scores are monitored at slaughter and linked to corrective measures. Early prediction of FPL scores could enable timely interventions. This study investigated the potential of routinely collected data to predict FPL scores at slaughter. Data from 592 broiler houses, each with at least 30 consecutive flocks, across 190 farms were included. The ability of various models to predict FPL scores above or below the threshold of 80 was compared. These models included univariate dynamic linear models (DLMs); multivariate DLMs using weather data of the first week of the production cycle; and random forest models using previous flock scores or DLM output, first-week weather variables, and current and previous flock and farm characteristics. Incorporation of DLM output in the random forest model provided the numerically highest performance, although this was not significantly better than the random forest model with raw previous flock scores. This model achieved an ROC AUC of 0.70, with the best threshold yielding a sensitivity of 74.4% and specificity of 60.2%. Previous flock FPL was the most important predictor, followed by the fraction of birds thinned, flock size difference between previous and current flock, and outside humidity. These findings highlight the value of weather variables in predicting FPL scores. Future research should explore additional factors which could explain within-house variation, such as indoor climate and feed changes, to improve predictive accuracy.
First-week mortality (FWM) is considered an important indicator of chick quality in broiler production, but its association with later performance is understudied. Data from 1,142 production cycles across 175 farms belonging to an Italian integrator were analyzed to identify links between management factors at the broiler farm and FWM, and explore the association between FWM and mortality after week 1, daily gain, and FCR. Median FWM was 0.95 % and median total mortality was 3.42 %. Factors associated with FWM were year and sex. For each 1 % increase in FWM, mortality after the first week increased by a factor 1.08 (95 % CI: 1.03-1.13, p = 0.002), adjusted for year and sex. However, FWM was not associated with daily gain or FCR. These results suggest that early mortality reflects vulnerabilities that persist throughout the production cycle, increasing later mortality without compromising growth performance. Variance in FWM was largely attributable to within-farm variance. Future research could explore how day-old chick quality contributes to both early and late mortality in broilers.
The dairy sector should overcome challenges in productivity, sustainability, and data management by adopting intelligent, scalable, and privacy-preserving technological solutions. Adopting data and artificial intelligence (AI) technologies is essential to ensure efficient operations and informed decision making and to keep a competitive market advantage. This paper proposes an integrated, multimodal AI framework to support data-intensive dairy farm operations by leveraging big data principles and advancing them through AI technologies. The proposed architecture incorporates edge computing, autonomous AI agents, and federated learning to enable real-time, privacy-preserving analytics at the farm level and promote knowledge sharing and refinement through research farms and cloud collaboration. Farms collect heterogeneous data, which can be transformed into embeddings for both local inference and cloud analysis. These embeddings form the input of AI agents that support health monitoring, risk prediction, operational optimization, and decision making. Privacy is preserved by sharing only model weights or anonymized data externally. The edge layer handles time-sensitive tasks and communicates with a centralized enterprise cloud hosting global models and distributing updates. A research and development cloud linked to research farms ensures model testing and validation. The entire system is orchestrated by autonomous AI agents that manage data, choose models, and interact with stakeholders, and human oversight ensures safe decisions, as illustrated in the practical use case of mastitis management. This architecture could support data integrity, scalability, and real-time personalization, along with opening up space for partnerships between farms, research institutions, and regulatory bodies to promote secure, cross-sector innovation.
Farmers, veterinarians and other animal health managers in the livestock sector are currently missing sufficient information on prevalence and burden of contagious endemic animal diseases. They need adequate tools for risk assessment and prioritization of control measures for these diseases. The DECIDE project develops data-driven decision-support tools, which present (i) robust and early signals of disease emergence and options for diagnostic confirmation; and (ii) options for controlling the disease along with their implications in terms of disease spread, economic burden and animal welfare. DECIDE focuses on respiratory and gastro-intestinal syndromes in the three most important terrestrial livestock species (pigs, poultry, cattle) and on reduced growth and mortality in two of the most important aquaculture species (salmon and trout). For each of these, we (i) identify the stakeholder needs; (ii) determine the burden of disease and costs of control measures; (iii) develop data sharing frameworks based on federated data access and meta-information sharing; (iv) build multivariate and multi-level models for creating early warning systems; and (v) rank interventions based on multiple criteria. Together, all of this forms decision-support tools to be integrated in existing farm management systems wherever possible and to be evaluated in several pilot implementations in farms across Europe. The results of DECIDE lead to improved use of surveillance data and evidence-based decisions on disease control. Improved disease control is essential for a sustainable food chain in Europe with increased animal health and welfare and that protects human health.
Lameness in dairy cows, linked to claw disorders and pain, is a major welfare concern. Studies worldwide use various scoring methods, resulting in differing prevalences. To address this, we developed a 3-level comparative locomotion scale (Welfare Quality equivalent [WQE]), to compare studies using different lameness scoring methods and provide insights into the distribution of lameness prevalence. This scale defines lameness and severe lameness according to stride length, weightbearing and back posture. To account for different locomotion scales used, we created 2 versions, with the main difference being the threshold for lameness. According to a "conservative" version of the scale, no irregularities are allowed in a normal gait, lameness is defined as any irregularities in the gait and severe lameness as: shortened strides, (strong) reluctance to bear weight ≥1 limb and an arched back standing and walking. According to a "lenient" version, some slight irregularities are allowed in a normal gait; lameness is defined as: shortened strides, unequal weightbearing, and arched back walking/standing; and severe lameness as: shortened strides, (strong) reluctance to bear weight ≥1 limb and arched back standing and walking. We also aimed to describe the prevalence of lameness by using meta-analysis, and to inventory risk factors associated with lameness. We identified 53 studies on lameness prevalence from SCOPUS and Web of Science, focusing on Northwest Europe over the last 25 years. Other eligibility criteria included scoring of cows by trained assessors and criteria relating to housing and methodology. We found that cow-level prevalence of lameness reported by studies ranged from 2.6% to 63.7% (with a median of 29.5%). A random effects model was used to analyze lameness distribution based on the WQE scale and to estimate weighted average lameness prevalences. The estimated mean prevalences were 28% (95% CI 23%-33%) for the lenient definition of lameness, 23% (95% CI 18%-29%) for the conservative definition of severe lameness, and 9% (95% CI 5%-12%) for the lenient definitions of severe lameness. The estimated prevalence of the conservative definition of lameness was extremely variable. Subgroup analyses were performed to explain the high heterogeneity between studies. Although some unexplained heterogeneity remained, lameness prevalence estimates differed between country and were lower in pasture access subgroups. Finally, we reviewed risk factors for lameness in Northwest Europe, based on 38 studies. Key risk factors for lameness included lack of pasture access, concrete floors, lying surfaces other than deep bedding, poor body condition and higher parity. This meta-analysis revealed significant variability in lameness prevalence across studies using different lameness scoring methods. Sparse herd information and vague lameness scoring descriptions complicated interpretation and WQE scale development. Our findings highlight the need for more randomized controlled trials investigating risk factors, and standardized, large-scale observational lameness prevalence research with detailed reporting. Furthermore, our results suggest that reduction of lameness prevalence is possible through providing pasture access, optimizing housing and other management interventions. We recommend using the WQE scale or an equivalent standard scale to obtain comparable scores and prevalences.
The increased uptake of sensor technologies and precision farming tools for the dairy cattle sector is enabling real-time monitoring of animal health, welfare, and productivity. These digital advancements provide high-frequency, objective, and large-scale phenotypic data for breeding purposes. This review explores the potential of sensor-derived data to improve genetic and genomic evaluations in dairy cattle and outlines key challenges, opportunities, and approaches associated with their implementation. While these data streams have great potential for genetic evaluations, their integration into national and international breeding programs remains limited due to fragmentation across sensor brands, lack of standardization, and challenges related to data accessibility, data access and portability rights, business interests, and governance. A crucial aspect of leveraging digital technologies in dairy cattle breeding is data harmonization and integration. We highlight the importance of establishing standardized data collection and data sharing protocols, implementing robust quality control and data cleaning methodologies, as well as defining novel sensor-based traits and estimating their genetic background. In this context, we compiled heritability estimates for novel traits derived from data recorded by sensors and other technologies in dairy cattle populations. The development of phenomics in breeding programs, which involves integrating multisource data-including sensor-based, genomic, and management information-will be key to accelerating genetic progress, especially for traits related to animal welfare, health, resilience, and efficiency. This review presents a roadmap for the effective use of sensor-derived data in genetic evaluations, advocating for centralized data infrastructures, transparent data-sharing agreements, and the role of different stakeholders from academia and industry, including organizations such as the International Committee on Animal Recording (ICAR) in establishing global standards and guidelines. By addressing these challenges, dairy breeding programs can fully harness precision dairy farming technologies to enhance production and environmental efficiency, improve animal health and welfare, and drive sustainable genetic advancements in the dairy cattle sector.
Farmers, veterinarians and other animal health managers in the livestock sector are currently missing sufficient information on the prevalence and burden of contagious endemic animal diseases. They need adequate tools for risk assessment and prioritization of control measures for these diseases. The DECIDE project develops data-driven decision-support tools, which present (i) robust and early signals of disease emergence and options for diagnostic confirmation; and (ii) options for controlling the disease along with their implications in terms of disease spread, economic burden and animal welfare. DECIDE focuses on respiratory and gastro-intestinal syndromes in the three most important terrestrial livestock species (pigs, poultry, cattle) and on reduced growth and mortality in two of the most important aquaculture species (salmon and trout). For each of these, we (i) identify the stakeholder needs; (ii) determine the burden of disease and costs of control measures; (iii) develop data sharing frameworks based on federated data access and meta-information sharing; (iv) build multivariate and multi-level models for creating early warning systems; and (v) rank interventions based on multiple criteria. Together, all of this forms decision-support tools to be integrated in existing farm management systems wherever possible and to be evaluated in several pilot implementations in farms across Europe. The results of DECIDE lead to improved use of surveillance data and evidence-based decisions on disease control. Improved disease control is essential for a sustainable food chain in Europe with increased animal health and welfare and that protects human health.
Determining the optimal insemination moment for individual cows is complex, particularly when considering the effects of pregnancy on milk production. The effect of pregnancy on the absolute milk yield has already been reported in several studies. Currently, there is limited quantitative knowledge about the association between days post-conception (DPC) and lactation persistency, based on a lactation curve model, and, specifically, how persistency changes during pregnancy and relates to the days in milk at conception (DIMc). Understanding this association might provide valuable insights to determine the optimal insemination moment. This study, therefore, aimed to investigate the association between DPC and lactation persistency, with an additional focus on the influence of DIMc. Available milk production data from 2005 to 2022 were available for 23,908 cows from 87 herds located throughout the Netherlands and Belgium. Persistency was measured by a lactation curve characteristic decay, representing the time taken to halve milk production after peak yield. Decay was calculated for 8 DPC (0, 30, 60, 90, 120, 150, 180, and 210 d after DIMc) and served as the dependent variable. Independent variables included DPC, DIMc (<= 60, 61-90, 91-120, 121-150, 151-180, 181-210, >210), parity group, DPC x parity group, DPC x DIMc, and variables from 30 d before DIMc as covariates. The results showed an increase in decay, which is to say, a decrease in persistency, during pregnancy for both parity groups, albeit in different ways. Specifically, from DPC 150 to DPC 210, multiparous cows showed a greater decline in persistency compared with primiparous cows. Furthermore, a later DIMc (cows conceiving later) was associated with higher persistency. Except for the early DIMc groups (DIMc <90), DIMc does not affect the change in persistency by gestation. The findings from this study contribute to a better derstanding of how DPC and DIMc during lactation influence lactation persistency, enabling more informed decision-making by farmers who wish to take persistency into account in their reproduction management.
The transition period is one of the most challenging periods in the lactation cycle of high-yielding dairy cows. It is commonly known to be associated with diminished animal welfare and economic performance of dairy farms. The development of data-driven health monitoring tools based on on-farm available milk yield development has shown potential in identifying health-perturbing events. As proof of principle, we explored the association of these milk yield residuals with the metabolic status of cows during the transition period. Over 2 yr, 117 transition periods from 99 multiparous Holstein-Friesian cows were monitored intensively. Pre- and postpartum dry matter intake was measured and blood samples were taken at regular intervals to determine β-hydroxybutyrate, nonesterified fatty acids (NEFA), insulin, glucose, fructosamine, and IGF1 concentrations. The expected milk yield in the current transition period was predicted with 2 previously developed models (nextMILK and SLMYP) using low-frequency test-day (TD) data and high-frequency milk meter (MM) data from the animal's previous lactation, respectively. The expected milk yield was subtracted from the actual production to calculate the milk yield residuals in the transition period (MRT) for both TD and MM data, yielding MRTTD and MRTMM. When the MRT is negative, the realized milk yield is lower than the predicted milk yield, in contrast, when positive, the realized milk yield exceeded the predicted milk yield. First, blood plasma analytes, dry matter intake, and MRT were compared between clinically diseased and nonclinically diseased transitions. MRTTD and MRTMM, postpartum dry matter intake and IGF1 were significantly lower for clinically diseased versus nonclinically diseased transitions, whereas β-hydroxybutyrate and NEFA concentrations were significantly higher. Next, linear models were used to link the MRTTD and MRTMM of the nonclinically diseased cows with the dry matter intake measurements and blood plasma analytes. After variable selection, a final model was constructed for MRTTD and MRTMM, resulting in an adjusted R2 of 0.47 and 0.73, respectively. While both final models were not identical the retained variables were similar and yielded comparable importance and direction. In summary, the most informative variables in these linear models were the dry matter intake postpartum and the lactation number. Moreover, in both models, lower and thus also more negative MRT were linked with lower dry matter intake and increasing lactation number. In the case of an increasing dry matter intake, MRTTD was positively associated with NEFA concentrations. Furthermore, IGF1, glucose, and insulin explained a significant part of the MRT. Results of the present study suggest that milk yield residuals at the start of a new lactation are indicative of the health and metabolic status of transitioning dairy cows in support of the development of a health monitoring tool. Future field studies including a higher number of cows from multiple herds are needed to validate these findings.
Despite decades of research, little is known regarding physiologic temporal limits for initiation of lactation in pregnant non-lactating cattle the aim of this study was to compare the lactation performances in primiparous Holstein cows after a short gestation length (GL) or abortion to those after a normal GL. The data were collected using an automated data collection system. The 94 herds evaluated were located in Belgium, France, Italy, the Netherlands and Germany. Data from a wide range of physiological cow-life events including birth and calving events, reproduction events (insemination, pregnancy checks, and abortions), and milking events were collected. The GL was defined as the interval between the last insemination and the subsequent calving (or abortion) within a range of 150–297 days. Animals were categorized into one of five categories based on GL quantiles (C-I to C-V). Lactation curve parameters including the scale, ramp, and decay were estimated using the Milkbot model. Then, the derived 305-day milk yield (M305-d), peak yield, and time to peak were compared between different GL categories. Of 13,732 lactations, 15 (0.11%) were found with a GL shorter than 210 days (ranging from 158 to 208 days). The 305-day milk yield was significantly lower in the C-I (7,566 ± 186) and C-II groups (7,802 ± 136 kg), compared to the C-III (8,254 ± 116 kg), C-IV (8,148 ± 119 kg), and C-V (8,255 ± 117 kg) groups. The same trends were found for the scale and peak yield of the lactation; the lowest scale were found for the C-I (31.5 ± 0.73) and C-II (32.8 ± 0.53) groups, and the highest were found for the C-III (34.5 ± 0.46), C-IV (34.9 ± 0.45), and C-V (35.0 ± 0.45) groups. Peak yield increased significantly from C-I (27.8 ± 0.66 kg) and C-II group (28.8 ± 0.48 kg) to the C-III (30.2 ± 0.42 kg) and further to the C-IV (30.6 ± 0.40 kg) and C-V (30.6 ± 0.41 kg) groups. Moreover, primiparous cows in the C-II GL category showed a higher milk yield persistency (decay of 1.30E−4 ± 3.55E−5) compared to those belonging to the C-IV (decay of 1.38E−4 ± 2.51E-5) and C-V (decay of 1.38E−4 ± 2.58E-5) group. In conclusion, results showed that primiparous cows with a shorter GL produced significantly less 305-day milk and peak yields, had a higher lactation persistency, and showed a lower upward slope of the lactation curve compared to those with a normal GL.
In the last years, the livestock sector is moving towards a more sustainable animal-based industry, mitigating the environmental impact of livestock while meeting the demand for high-quality food. To achieve these goals, farms are using a more technological approach, adopting algorithms to manipulate the vast amount of data from sensors and routine operations. The results will be useful for making more objective decisions. In this context, machine learning - a branch of Artificial Intelligence applied to the study of prediction, inference, and clustering algorithms - can be successfully employed. Nowadays, machine learning algorithms are successfully used to solve many issues in the livestock sector, such as early disease detection, and they are expected to be employed in the future for welfare monitoring. This brief review gives an overview of the current state of the art of the most popular applications for dairy science and the most widely used and best-performing algorithms, highlighting the challenges and obstacles for broad acceptance of these techniques in the dairy sector.
(Sub)clinical hypocalcaemia occurs frequently in the dairy industry, and is one of the earliest symptoms of an impaired transition period. Calcium deficiency is accompanied by changes in cows’ daily behavioural variables, which can be measured by sensors. The goal of this study was to construct a predictive model to identify cows at risk of hypocalcaemia in dairy cows using behavioural sensor data. For this study 133 primiparous and 476 multiparous cows from 8 commercial Dutch dairy farms were equipped with neck and leg sensors measuring daily behavioural parameters, including eating, ruminating, standing, lying, and walking behaviour of the 21 days before calving. From each cow, a blood sample was taken within 48 h after calving to measure their blood calcium concentration. Cows with a blood calcium concentration ≤2.0 mmol/L were defined as hypocalcemic. In order to create a more context based cut-off, a second way of dividing the calcium concentrations into two categories was proposed, using a linear mixed-effects model with a k-Means clustering. Three possible binary predictive models were tested; a logistic regression model, a XgBoost model and a LSTM deep learning model. The models were expanded by adding the following static features as input variables; parity (1, 2 or 3+), calving season (summer, autumn, winter, spring), day of calcium sampling relative to calving (0, 1 or 2), body condition score and locomotion score. Of the three models, the deep learning model performed best with an area under the receiver operating characteristic curve (AUC) of 0.71 and an average precision of 0.47. This final model was constructed with the addition of the static features, since they improved the model’s tuning AUC with 0.11. The calcium label based on the cut-off categorization method proved to be easier to predict for the models compared to the categorization method with the k-means clustering. This study provides a novel approach for the prediction of hypocalcaemia, and an ameliorated version of the deep learning model proposed in this study could serve as a tool to help monitor herd calcium status and to identify animals at risk for associated transition diseases.