Selecting and raising dairy animals that are more likely to reach their potential is a strategy to increase milk production efficiency and overall profitability. However, indicators are necessary for the early identification of animals that are less likely to perform well, allowing for their early culling and ensuring that resources are allocated to those with the highest potential. The objective of this study was to analyze the association between early-life animal health and performance with longevity, production, and profitability. After data cleaning, the following early-life measures (i.e., predictors) were available for 363 female calves born between June 2014 and November 2015 in eight dairy herds from New Brunswick, Canada (average: 45 calves/farm; SD: 26.1 calves/farm; median: 42 calves/farm; range: 15-95 calves/farm): birth weight, weaning weight, weaning age, weaning average daily gain (weaning ADG), immunoglobulin G (IgG) serum concentration, the occurrence of navel infection, diarrhea, and pneumonia, and if animals received antibiotic treatment between birth and weaning. Their subsequent length of life (LL), length of productive life (LPL), lifetime cumulative energy-corrected milk (ECM), and lifetime cumulative milk value (i.e., response variables) were provided by the Canadian dairy herd improvement agency. Bayesian Additive Regression Tree models were trained for each response variable using 5-fold cross-validation. Models were evaluated using the RMSE and R-2. The three most important predictors were identified using permutation, and the relationship between response variables and important predictors was assessed using accumulated local effect plots. The RMSE for LL, LPL, ECM, and milk value were 1.43 years, 1.37 years, 16 314.94 kg, and $CAD 11 525.68, respectively, whereas the R-2 values were 0.30, 0.25, 0.29, and 0.29, respectively, indicating a moderate relationship between predictors and response variables. Non-linear relationships were found between the response variables and important predictors. Animals born with low or high birth weights were associated with decreased LL, LPL, ECM, and milk value. The highest LL, LPL, and milk value was observed for calves weaned between 1.9 and 2.0 months old, followed by a decline for calves weaned at older ages. The lowest LL and ECM were associated with weaning ADG of 0.786 kg/day, while 0.787 kg/day was associated with the lowest LPL. Lastly, both ECM and milk value were highest when serum IgG values were 1 659 mg/dL. These findings provide valuable insights for optimizing early culling decisions and enhancing the productivity and profitability of dairy farms.
Three weeks prior to calving to three weeks after calving, the transition period poses challenges for dairy cattle and farmers. Vast changes in housing, feeding, and reproduction might result in milk drop, metabolic and reproductive diseases. Moreover, most of the metabolic processes are intricately linked as many conditions can coexist. This challenge means that dairy producers and their advisors have difficulty drawing concise conclusions because of all aspects and relationships in transition cow management. Herein, machine-learning techniques and knowledge-graph theory were explored with a view to creating a decision-support system that could provide producers and their advisors with knowledge from domain literature. Specifically, knowledge is modelled as entities and relationships in knowledge graph theory, and natural language models were developed to extract information as knowledge graphs. A dataset comprising 1152 sentences from 20 papers was created and split into 922 sentences for training and 230 sentences for testing. Sequentially, two deep learning models were trained to extract entities and relationships respectively. For training results, a Bi-directional Long-Short-Term Memory model was applied for the entity extraction task and obtained an F1 score of 80 %. As for relationship extraction, a Transformer-based model was deployed but yielded a low F1 of 23 %, thus another pre-trained Transformer model with 89 % accuracy was deployed into the system. After feeding the domain literature into the deep-learning models, a knowledge graph of 1,576 nodes and 3,456 edges was constructed and stored in the graph database Neo4j. Afterward, a semantic parsing method was used to allow users to conduct question answering through the knowledge graph in natural language. In addition, to determine the quality of answers that the knowledge built from the papers, answers were sampled and evaluated based on human judgment. On average, answers scored 7.5 out of 10 and proved informative with respect to the original literature. Although the final interactive results demonstrated a high degree of visualization and scalability, this study primarily sought to demonstrate its feasibility. For tailored commercial applications, further improvements could be implemented in knowledge graph expansion and reasoning.
Nowadays, in many practical situations, analytical tasks need to be performed on complex heterogeneous data, often described by a domain ontology (DO). Such cases abound in life science fields such as agro-informatics, where observations and measures on animals/plants are logged for subsequent mining. The data is naturally structured as graph(s), unlabelled and missing some values, hence it fits well pattern mining. In our own precision farming project aimed at decision support for dairy cow management, we mine for knowledge in milk production data. In one task, we aim at contrast patterns explaining the relative impact of independent production factors. To that end, ontologically-generalized graph patterns (OGPs), a variety of generalized graph patterns, where vertices and edges are labelled by DO classes and properties, respectively, were defined. A mining methodology was also designed that reconciles OWL DOs, abstraction from RDF graphs and literals in data. To address the well-known cost-related limitations of graph mining -exacerbated here by class/property specializations and data properties- we split the mining task into (1) mining of generic object property topology patterns and (2) label refinement. Those focus on two sorts of OGPs, called topologies and class stars, respectively, which, after being mined separately, get (3) assembled into fully-fledged OGPs.
Increasing the productive lifespan of dairy cows is important to achieve a sustainable dairy industry, but making strategic culling decisions based on cow profitability is challenging for farmers. The objective of this study was to carry out a lifetime cost-benefit analysis based on production and health records and to explore different culling decisions among farmers. The cost-benefit analysis was conducted for 22 747 dairy cows across 114 herds in Quebec, Canada for which feed costs and the occurrence of diseases were reported. Costs and revenues related to productive lifespan were compared among cohorts of cows that left their respective herd at the end of their last completed lactation or stayed for a complete additional lactation. Hierarchical clustering analysis was carried out based on costs and revenues to explore different culling decisions among farmers. Our results showed that the knowledge of lifetime cumulative costs and revenues was of great importance to identify low-profitable cows at an earlier lactation, while only focusing on current lactation costs and revenues can lead to an erroneous assessment of profitability. While culling decisions were mostly based on current lactation costs and revenues and disregarded the occurrence of costly events on previous lactations, there was variation among farmers as we identified three different culling decision clusters. Monitoring cumulative costs and revenues would help farmers to identify low-profitable cows at an earlier lactation and make the decision to increase herd productive lifespan and farm profitability by keeping the most profitable cows.
Nowadays, linked data (LD) are ubiquitous and mining them for knowledge, e.g. frequent patterns, needs not be argued for.A domain ontology (DO) on top of a LD dataset enables the discovery of abstract patterns, a.k.a. generalized, capturing –rather than identical sub-structures–conceptual regularities in data. Yet with the resulting ontologically-generalized graph patterns (OGP), a miner faces the combined challenges of graph topology and a label hierarchy, which amplifies well-known difficulties with graphs such as support counting or non redundant pattern listing. As OGP mining is yet to be addressed in its generality, we propose a formalization and study two workaround methods that avoid tackling it head-on, i.e. deal with each aspect separately. Both perform pure graph mining with adapted label sets: gSpan-OF merely strips labels of hierarchical structure while Tax-ON first mines frequent graph topologies with only root classes as labels, then successively refines labels on each topology.
Studies of dairy cow longevity usually focus on the animal life after first calving, with few studies considering early life conditions and their effects on longevity. The objective was to evaluate the effect of birth conditions routinely collected by Dairy Herd Improvement agencies on offspring longevity measured as length of life and length of productive life. Lactanet provided 712,890 records on offspring born in 5,425 Quebec dairy herds between January 1999 and November 2015 for length of life, and 506,066 records on offspring born in 5,089 Quebec dairy herds between January 1999 and December 2013 for length of productive life. Offspring birth conditions used in this study were calving ease (unassisted, pull, surgery, or malpresentation), calf size (small, medium, or large), and twinning (yes or no). Observations were considered censored if the culling reason was "exported," "sold for dairy production," or "rented out" as well as if the animals were not yet culled at the time of data extraction. If offspring were not yet culled when the data were extracted, the last test-day date was considered the censoring date. Conditional inference survival trees were used in this study to analyze the effect of offspring birth conditions on offspring longevity. The hazard ratio of culling between the groups of offspring identified by the survival trees was estimated using a Cox proportional hazard model with herd-year-season as a frailty term. Five offspring groups were identified with different length of life based on their birth condition. Offspring with the highest length of life [median = 3.61 year; median absolute deviation (MAD) = 1.86] were those classified as large or medium birth size and were also the result of an unassisted calving. Small offspring as a result of a twin birth had the lowest length of life (median = 2.20 year; MAD = 1.69) and were 1.52 times more likely to be culled early in life. Six groups were identified with different length of productive life. Offspring that resulted from an unassisted or surgery calving and classified as large or medium when they were born were in the group with the highest length of productive life (median = 2.03 year; MAD = 1.63). Offspring resulting from a malpresentation or pull in a twin birth were in the group with the lowest length of productive life (median = 1.15 year; MAD = 1.11) and were 1.70 times more likely to be culled early in life. In conclusion, birth conditions of calving ease, calf size, and twinning greatly affected offspring longevity, and such information could be used for early selection of replacement candidates.
A domain ontology (DO) is a machine-readable knowledge repository which, whenever properly exploited, can help to discover meaningful and intelligible patterns from compatible datasets. Yet since such data is naturally graph-shaped, the corresponding task amounts to mining what we call ontologically-generalized graph patterns. We study the underlying problem within a dairy production context where a dedicated DO has been designed beforehand. Two alternative mining approaches have been designed, both representing adaptations of methods from the literature. We evaluated them on an excerpt from our dairy production dataset and report here their respective limitations. We also sketch a way to approach the design of ontology-powered graph miner.
The ability of dairy farmers to keep their cows for longer could positively enhance the economic performance of the farms, reduce the environmental footprint of the milk industry, and overall help in justifying a sustainable use of animals for food production. However, there is little published on the current status of cow longevity and we hypothesized that a reason may be a lack of standardization and an over narrow focus of the longevity measure itself. The objectives of this critical literature review were: (1) to review metrics used to measure dairy cow longevity; (2) to describe the status of longevity in high milk-producing countries. Current metrics are limited to either the length of time the animal remains in the herd or if it is alive at a given time. To overcome such a limitation, dairy cow longevity should be defined as an animal having an early age at first calving and a long productive life spent in profitable milk production. Combining age at first calving, length of productive life, and margin over all costs would provide a more comprehensive evaluation of longevity by covering both early life conditions and the length of time the animal remains in the herd once it starts to contribute to the farm revenues, as well as the overall animal health and quality of life. This review confirms that dairy cow longevity has decreased in most high milk-producing countries over time and its relationship with milk yield is not straight forward. Increasing cow longevity by reducing involuntary culling would cut health costs, increase cow lifetime profitability, improve animal welfare, and could contribute towards a more sustainable dairy industry while optimizing dairy farmers’ efficiency in the overall use of resources available.
Precision farming is about improving farming processes through in-depth analysis of the generated data. Dairy farming, in particular, is being intensively computerized and hence a fertile soil for such applications. In our own project, we investigate the benefit of data analytics in optimizing dairy production. To that end, the Valacta centre of expertise shares a dataset recording the performances of dairy cows and farms in Eastern Canada. Here, we tackle the design of a domain ontology (DONT) on top of it. The dairy cattle performance ontology (DCPO) reconciles the complex structure to the heterogeneous nature of dairy data within a unified framework that ensures extensibility to external data. It also provides a common vocabulary for both stakeholders and automated knowledge management tools, and, in the longer term, should support explainability for predictive neural models. We present here the bottom-up process of DCPO design and summarize its current content. We also illustrate its present and future usages.
A domain ontology (DO) is a machine-readable knowledge repository compatible with the popular knowledge graph (KG) format. An intriguing question is how to leverage a DO plus a KG in a neural learning process. We propose to use ontology-rooted graph patterns mined from a DO-compatible graph translation of the raw data as a vector for injecting some domain knowledge into the neural network. Such patterns represent a frequently occurring regularities in the data yet they are expressed in terms of the ontological entities (classes, properties, etc.) and reflect additional knowledge from the KG. Using them as an additional input to the learning process seems a promising way to guide it towards improved explainability, accuracy and convergence, as well as, in a more general vein, increase the generalization power of the neural models.
Our first objective was to estimate the prevalence of foot lesions by type of milking system in dairy cows examined during regular hoof-trimming sessions between 2015 and 2018 in Québec dairy herds. A secondary objective was to describe the effect of day-to-day variation, cow, and herd characteristics on the prevalence of foot lesions. Data included 52,427 observations (on a cow during a specific trimming session) performed on 28,470 cows (≥2 yr old) from 355 herds. Only observations from trimming sessions in which ≥90% of the lactating herd was trimmed were considered. Lesions were recorded at the hoof level by 17 trained hoof trimmers between March 23, 2015, and July 10, 2018, using a computerized recording system. Hoof-level information was then matched with cow information and centralized at the Eastern Canada Dairy Herd Improvement. Foot lesions were classified into 6 categories: infectious, white line disease, heel erosion, ulcers, hemorrhages, and any type of foot lesions. Prevalence of each outcome was quantified using the marginal predicted mean probability estimated from a null generalized linear mixed model with a logit link, and accounted for clustering of observations by cow and by herd. Variance was partitioned to assess the variation in the probability of the outcomes attributable to each level of the data structure (day of exam, cow, and herd). Prevalence of a given foot lesion as function of milking system and of various explanatory variables (mean herd size, herd average daily production, breed of the cow, age of the cow at trimming, and year of the visit) was then estimated using a generalized linear mixed model. At least 1 foot lesion was observed in 29% of cows examined during regular trimming sessions in Québec from 2015 to 2018. Prevalence for any type of lesion was 27% for pipeline, 38% for robotic milking, and 41% for milking parlors. The highest prevalence of infectious lesions (mainly digital dermatitis) was observed in milking parlors and robotic systems, while the most prevalent lesions in pipeline were hemorrhages. Herd-level factors explained most of the disease probability for infectious diseases, heel erosion, and hemorrhages. Therefore, control of these diseases should be based on applying best herd-management practices. On the other hand, probabilities of white line disease and sole ulcers were mainly determined by cow-level characteristics.
Lameness is a common welfare issue in dairy farms and can negatively influence productivity and farm profitability. On-farm measurements are time consuming and continuous monitoring of lameness can be thus challenging. A predictive model that is suitable for routine field applications can be thus an efficient strategy to improve lameness status in dairy herds. We explored the use of a machine learning approach based on decision tree induction to detect and monitor the risk level of dairy herds for lameness based on 20 routinely pre-collected farm-based records related to herd management and housing, milk production performance, reproduction performance, longevity, and genetic merit. The risk of lameness was determined for 229 dairy herds based on herd prevalence. The results based on 10-fold repeated cross-validation suggest that there was little difference between a simple decision tree and more advanced machine learning algorithms such as random forests or boosting algorithms with an area under the receiver operating characteristic curve (AUC) of 0.73–0.75, but machine learning performed better than classical multivariate logistic regression (AUC of 0.28). Model sensitivity based on machine learning was highest at 0.58, whereas model specificity was highest at 0.89 among all models tested. Stacking the individual machine learning algorithms to a higher-order model slightly lifted model performance (AUC of 0.76) at a model sensitivity of 0.54 and a model specificity of 0.94 but at a loss in model interpretability. Additional information more closely related to lameness are required to further improve model performance. Nevertheless, decision tree induction was found to be useful in analysing small data sets that frequently occur in various livestock farming disciplines. Overall, the machine learning approach described in our study may be an appropriate and powerful decision support and monitoring tool, and can be implemented in a computerized information system to detect herds with potential deficiencies in lameness.
Continuous assessment of the herd status is important in order to monitor and adjust to changes in the welfare and health status but can be time consuming and expensive. In this study, herd status indicators from routinely collected dairy herd improvement (DHI) records were used to develop a remote herd assessment tool with the aim to help producers and advisors benchmark the herd status and identify herd management issues affecting welfare and health. Thirteen DHI indicators were selected from an initial set of 72 potential indicators collected on 4324 dairy herds in Eastern Canada. Data were normalized to percentile ranks and aggregated to a composite herd status index (HSI) with equal weights among indicators. Robustness analyses indicated little fluctuation for herds with a small HSI (low status) or large HSI (high status), suggesting that herds in need of support could be prioritized and effectively monitored over time, limiting the need for time-consuming farm visits. This tool allows evaluating herds relative to their peers through the composite index and highlighting specific areas with opportunities for improvements through the individual indicators. This procedure could be applied to similar multidimensional livestock farming issues, such as environmental and socio-economic studies.
Life-time profitability is a leading factor in the decision to keep a cow in a herd, or sell it, that a dairy farmers face regularly. A cow’s profit is a function of the quantity and quality of its milk production, health and herd management costs, which in turn may depend on factors as diverse as animal genetics and weather. Improving the decision making process, e.g. by providing guidance and recommendation to farmers, would therefore require predictive models capable of estimating profitability. However, existing statistical models cover only partially the set of relevant variables while merely targeting milk yield. We propose a methodology for the design of extensive predictive models reflecting a wider range of factors, whose core is a Long Short-Term Memory neural network. Our models use the time series of individual features corresponding to earlier stages of cow’s life to estimate target values at following stages. The training data for our current model was drawn from a dataset captured and preprocessed for about a million cows from more than 6000 different herds. At validation time, the model predicted monthly profit values for the fifth year of each cow (from data about the first four years) with a root mean squared error of 8.36 $/cow/month, thus outperforming the ARIMA statistical model by 68% (14.04 $/cow/month). Our methodology allows for extending the models with attention and initializing mechanisms exploiting precise information about cows, e.g. genomics, global herd influence, and meteorological effects on farm location.
Petko Valtchev合作论文数Departement d'informatique, University of Quebec at Montreal6