The integration of Earth Observation (EO) into human health research has expanded significantly, particularly since 2009, highlighting its potential for disease modelling, environmental exposure assessment, and public health decision-making. This review explores the evolving role of EO in health applications through a bibliometric analysis of 1751 research documents retrieved from the Web of Science (WoS) database. These documents were selected using targeted keywords and after excluding non-primary literature such as reviews, editorials, and meeting abstracts. Findings revealed a substantial increase in EO-health research outputs, growing from 2 publications in 1991 to 266 in 2024, with a notable surge beginning in 2009. More than 65 % of the selected studies contributed to Sustainable Development Goal (SDG) 13 on Climate Action, followed by SDG 3 on Good Health and Wellbeing (n = 994) and SDG 11 on Sustainable Cities and Communities (n = 980), illustrating EO's cross-cutting relevance. Despite this growth, the field remains fragmented due to inconsistent data formats, limited accessibility, and weak interdisciplinary collaboration. A key challenge is the persistent divide between EO data producers and health practitioners, which hampers the effective translation of EO insights into practice. This review highlights the importance of co-production approaches that bring together researchers, policymakers, and communities to address these barriers. By promoting standardisation, enhancing data interoperability, and fostering interdisciplinary collaboration, EO can be more effectively leveraged to support disease surveillance, environmental health monitoring, and evidence-based policy interventions aligned with global health and sustainability goals.
The One Health approach unites efforts across human-animal-environment interfaces against shared threats like zoonotic diseases. T-Racing is a Shiny web application, that supports epidemiological investigations and helps contain livestock-related disease spread, aligning with multidisciplinary principles to safeguard public health. The application uses Temporal Network Analysis techniques to address the dynamic nature of animal trade, facilitating backward and forward tracing strategies. T-Racing leverages web services to retrieve data from multiple sources simultaneously and in near real-time through the plumber package and is distributed using Shinyproxy. T-Racing manages and analyze extensive and diverse datasets within the same environment, including animal movement data, disease outbreak data, and genomic data, all obtained from Italian National databases. In this work, we show T-Racing's capabilities by simulating epidemiological investigations of brucellosis and tuberculosis outbreaks that occurred in non-endemic areas of Italy. To further highlight its capabilities, an interactive demo of T-Racing is available, showcasing its potential and features. This tool supports epidemiological investigations by adopting a data-driven approach, guiding users through the analysis via an iterative process while leveraging their expertise. Therefore, it enables faster data analysis, improves understanding of disease transmission patterns, and facilitates prompt and targeted interventions.
The analysis of networks from cattle movements is an important approach to investigate areas and premises where disease outbreaks can occur and be contained. Vulnerability analysis allows a more profound understanding of the network by combining network measures with the strategic removal of nodes, helping to identify more vulnerable areas and the best metrics to support disease control planning. Therewith, the aim of this study was to analyze the network vulnerability of cattle movements from 2013 to 2022, in Minas Gerais, Brazil and to identify the spatial spreaders into the network to improve infectious disease control programs by targeted risk-based surveillance and intervention. The vulnerability was calculated considering the graphs diameter and the spatial spreaders with a threshold distance of 300 km, for incoming (IN) and outgoing (OUT) movements. Additionally, a risk-based analysis was performed in the more vulnerable region. The results showed Triângulo Mineiro/ Alto Paranaíba with higher vulnerability and many IN spatial spreaders, as well as Vale do Mucurí region with many OUT spatial spreaders. The risk-based analysis revealed betweenness and out degree as the most effective measures to be considered for intervention. Therefore, the vulnerability analysis and the spatial spreader were observed as great tools for risk-based interventions and surveillance. Furthermore, Triângulo Mineiro/ Alto Paranaíba and Vale do Mucuri regions were important regions, considering restriction of animal infectious disease spread in Minas Gerais, Brazil.
The aim of this study was to analyze the network vulnerability of cattle movements from 2013 to 2022, in Minas Gerais, Brazil and to identify the spatial spreaders into the network to improve infectious disease control programs by targeted risk-based surveillance and intervention. The vulnerability was calculated considering the graphs diameter and the spatial spreaders with a threshold distance of 300 km, for incoming (IN) and outgoing (OUT) movements. Additionally, a risk-based analysis was performed in the more vulnerable region. The results showed Triângulo Mineiro / Alto Paranaíba with higher vulnerability and many IN spatial spreaders, as well as Vale do Mucurí region with many OUT spatial spreaders. The risk-based analysis revealed betweenness and out degree as the most effective measures to be considered for intervention. Therefore, the vulnerability analysis and the spatial spreader were observed as great tools for risk-based interventions and surveillance. Furthermore, Triângulo Mineiro / Alto Paranaíba and Vale do Mucuri regions were important regions, considering restriction of animal infectious disease spread in Minas Gerais, Brazil. ### Competing Interest Statement The authors have declared no competing interest.
Animal movements are a key factor in the spread of pathogens. Consequently, network analysis of animal movements is a well-developed and well-studied field. The relationships between animals facilitate the diffusion of infectious agents and, in particular, shared environments and close interactions can facilitate cross-species transmission. Cattle are often the focus of these studies since they are among the most widely distributed and traded species globally. This remains true for Italy as well, but with an important additional consideration. Indeed, another important productive reality in the peninsula is buffalo farming. These farms have an interesting characteristic: approximately two-thirds of them also rear cattle. This coexistence between cattle and buffalo could have an impact on the diffusion of pathogens. Given that buffalo farms are often overlooked in the literature, the primary goal of this work is to investigate the potential consequences of omitting buffalo from cattle network analyses. To investigate this impact, we will focus on Q fever, a disease that can infect both species and is present on the Italian territory and for which the impact of the buffalo population has not been thoroughly studied, and simulate its spread to the farms of both species through compartmental models. Our analysis reveals that despite the significant difference in network sizes, the unique characteristic of Italian buffalo farms makes the buffalo network essential for a comprehensive understanding of bovine disease dynamics in Italy.
Brucellosis is one of the world's major zoonotic pathogens and is responsible for enormous economic losses as well as considerable human morbidity in endemic areas. Definitive control of human brucellosis requires control of brucellosis in livestock through practical solutions that can be easily applied to the field. In Italy, brucellosis remains endemic in several southern provinces, particularly in Sicily Region. The purpose of this paper is to describe the developed brucellosis model and its applications, trying to reproduce as faithfully as possible the complex transmission process of brucellosis accounting for the mixing of grazing animals. The model focuses on the contaminated environment rather than on the infected animal, uses real data from the main grazing areas of the Sicily Region, and aims to identify the best control options for minimizing the spread (and the prevalence) and to reach the eradication within the concerned areas. Simulation results confirmed the efficacy of an earlier application of the controls, showed the control should take place 30 days after going to pasture, and the culling time being negligible. Moreover, results highlighted the importance of the timing of both births and grazing pastures (and their interaction) more than other factors. As these factors are region‑specific, the study encourages the adoption of different and new eradication tools, tuned on the grazing and commercial behavior of each region. This study will be further extended to improve the model's adaptability to the real world, with the purpose of making the model an operational tool able to help decision makers in accelerating brucellosis eradication in Italy.
West Nile virus (WNV) is a mosquito-borne virus potentially causing serious illness in humans and other animals. Since 2004, several studies have highlighted the progressive spread of WNV Lineage 2 (L2) in Europe, with Italy being one of the countries with the highest number of cases of West Nile disease reported. In this paper, we give an overview of the epidemiological and genetic features characterising the spread and evolution of WNV L2 in Italy, leveraging data obtained from national surveillance activities between 2011 and 2021, including 46 newly assembled genomes that were analysed under both phylogeographic and phylodynamic frameworks. In addition, to better understand the seasonal patterns of the virus, we used a machine learning model predicting areas at high-risk of WNV spread. Our results show a progressive increase in WNV L2 in Italy, clarifying the dynamics of interregional circulation, with no significant introductions from other countries in recent years. Moreover, the predicting model identified the presence of suitable conditions for the 2022 earlier and wider spread of WNV in Italy, underlining the importance of using quantitative models for early warning detection of WNV outbreaks. Taken together, these findings can be used as a reference to develop new strategies to mitigate the impact of the pathogen on human and other animal health in endemic areas and new regions.
Background: Tunisia has experienced several West Nile virus (WNV) outbreaks since 1997. Yet, there is limited information on the spatial distribution of the main WNV mosquito vector Culex pipiens suitability at the national level. Objectives: In the present study, our aim was to predict and evaluate the potential and current distribution of Cx. pipiens in Tunisia. Methods: To this end, two species distribution models were used, i.e. MaxEnt and Random Forest. Occurrence records for Cx. pipiens were obtained from adult and larvae sampled in Tunisia from 2014 to 2017. Climatic and human factors were used as predictors to model the Cx. pipiens geographical distribution. Mean decrease accuracy and mean decrease Gin i indices were calculated to evaluate the importance of the impact of different environmental and human variables on the probability distribution of Cx. pipiens. Results: Suitable habitats were mainly distributed next to oases, in the north and eastern part of the country. The most important predictor was the population density in both models. The study found out that the governorates of Monastir, Nabeul, Manouba, Ariana, Bizerte, Gabes, Medenine and Kairouan are at highest epidemic risk. Conclusions: The potential distribution of Cx. pipiens coincides geographically with the observed distribution of the disease in humans in Tunisia. Our study has the potential for driving control effort in the fight against West Nile vector in Tunisia.
From 24 December 2020 to 8 February 2021, 163 cases of SARS-CoV-2 Alpha variant of concern (VOC) were identified in Chieti province, Abruzzo region. Epidemiological data allowed the identification of 14 epi-clusters. With one exception, all the epi-clusters were linked to the town of Guardiagrele: 149 contacts formed the network, two-thirds of which were referred to the family/friends context. Real data were then used to estimate transmission parameters. According to our method, the calculated Re(t) was higher than 2 before the 12 December 2020. Similar values were obtained from other studies considering Alpha VOC. Italian sequence data were combined with a random subset of sequences obtained from the GISAID database. Genomic analysis showed close identity between the sequences from Guardiagrele, forming one distinct clade. This would suggest one or limited unspecified viral introductions from outside to Abruzzo region in early December 2020, which led to the diffusion of Alpha VOC in Guardiagrele and in neighbouring municipalities, with very limited inter-regional mixing.
In a temperature-increasing scenario, due to global warming, the individual thermic resilience of the male assumes a crucial role in the reproductive efficiency of a male since the thermic stress, such as the inability of the male to reduce body or regional temperature on a physiological level, impairs testicular function. In this study, the effect of the environmental conditions on the fresh semen quality, in terms of volume, concentration, total sperm in the ejaculate, total motility, normal morphology, membrane integrity, and discarding rate, were compared longitudinally in Belgian Blue (BB) and Brown Swiss (BS) bulls. The environmental conditions, summarized in the mean temperature-humidity index (THI), were calculated on the day of collection, as well as 7 days (epididymal maturation), 35 days (late spermatogenesis), and 70 days (early spermatogenesis) before the collection, to reflect spermatogenesis time. Our findings showed that limited seasonal effects were present in the semen quality of BS bulls. On the other hand, in BB bulls lower semen quality was found between July and November, with a different timing depending on the seminal parameter. This effect of the season on BB semen parameters appears to be related to the THI. The data presented in this study shows that the temperature and humidity, summarized in THI, could affect the semen quality of the bull on breed basis, given that volume, concentration, total sperm in the ejaculate, total motility, membrane integrity, and sperm normal morphology were significantly reduced by an increasing THI in the Belgian Blue bulls, but not in Brown Swiss bulls.
In February 2020, Italy became the epicentre for COVID-19 in Europe and at the beginning of March, in response to the growing epidemic, the Italian Government put in place emergency measures to restrict the movement of the population. Human mobility represents a crucial element to be considered in modelling human infectious diseases. In this paper, we examined the mechanisms underlying COVID-19 propagation using a Susceptible-Infected stochastic model (SI) driven mainly by commuting network in Italy. We modelled a municipality-specific contact rate to capture the disease permeability of each municipality, considering the population at different times of the day and describing the characteristic of the municipalities as attractors of commuters or places that make their workforce available elsewhere. The purpose of our analysis is to provide a better understanding of the epidemiological context of COVID-19 in Italy and to characterize the territory in terms of vulnerability at local or national level. The use of data at such a high spatial resolution allows highlighting particular situations on which the health authorities can promptly intervene to control the disease spread. Our approach provides decision-makers with useful geographically detailed metrics to evaluate those areas at major risk for infection spreading and for which restrictions of human mobility would give the greatest benefits, not only at the beginning of the epidemic but also in the last phase, when the risks deriving from the gradual lockdown exit strategies must be carefully evaluated.
Chemical and organoleptic properties of dairy products largely depend on the action of microorganisms that tend to be selected in cheese during ripening in response to the availability of specific substrates. The aim of this work was to evaluate the effects of a diet enriched with hemp seeds on the microbiota composition of fresh and ripened cheese produced from milk of lactating ewes. Thirty-two half-bred ewes were involved in the study, in which half (control group) received a standard diet, and the other half (experimental group) took a diet enriched with 5% hemp seeds (on a DM basis) for 35 d. The dietary supplementation significantly increased the lactose in milk, but no variations in total fat, proteins, caseins, and urea were observed. Likewise, no changes in total fat, proteins, or ash were detected in the derived cheeses. The metagenomic approach was used to characterize the microbiota of raw milk and cheese. The phyla Proteobacteria and Firmicutes were in equally high abundance in both control and experimental raw milk samples, whereas Bacteroidetes was less abundant. The scenario changed when considering the dairy products. In all cheese samples, Firmicutes was clearly predominant, with Streptococcaceae being the most abundant family in the experimental group. The reduction of taxa observed during ripening was in accordance with the increment (relative abundance) of the starter culture Lactococcus lactis and Streptococcus thermophilus, which together dominate the microbial community. The analysis of the volatile profile in ripened cheeses led to the identification of 3 major classes of compounds: free fatty acids, ketones, and aldehydes, which indicate a prevalence of lipolysis compared with the other biochemical mechanisms that characterize the cheese ripening.
West Nile Disease (WND) is one of the most spread zoonosis in Italy and Europe caused by a vector-borne virus. Its transmission cycle is well understood, with birds acting as the primary hosts and mosquito vectors transmitting the virus to other birds, while humans and horses are occasional dead-end hosts. Identifying suitable environmental conditions across large areas containing multiple species of potential hosts and vectors can be difficult. The recent and massive availability of Earth Observation data and the continuous development of innovative Machine Learning methods can contribute to automatically identify patterns in big datasets and to make highly accurate identification of areas at risk. In this paper, we investigated the West Nile Virus (WNV) circulation in relation to Land Surface Temperature, Normalized Difference Vegetation Index and Surface Soil Moisture collected during the 160 days before the infection took place, with the aim of evaluating the predictive capacity of lagged remotely sensed variables in the identification of areas at risk for WNV circulation. WNV detection in mosquitoes, birds and horses in 2017, 2018 and 2019, has been collected from the National Information System for Animal Disease Notification. An Extreme Gradient Boosting model was trained with data from 2017 and 2018 and tested for the 2019 epidemic, predicting the spatio-temporal WNV circulation two weeks in advance with an overall accuracy of 0.84. This work lays the basis for a future early warning system that could alert public authorities when climatic and environmental conditions become favourable to the onset and spread of WNV.
Pecorino di Farindola is a typical cheese produced in the area surrounding the village of Farindola, located in the Abruzzo Region (central Italy), unique among Italian cheese because only raw ewe milk and pig rennet are used for its production. In the literature it is well documented that raw milk is able to support the growth of pathogenic microorganisms such as Listeria monocytogenes. Predictive microbiology can be useful in order to predict growth-death kinetics of pathogenic bacteria, on the basis of known environmental conditions. Aim of this study was to compare predictions obtained from a model, originally designed to predict the kinetics of L. monocytogenes in the dynamic growth-death environment of drying fresh sausage, with the results of challenge tests performed during the ripening of Pecorino di Farindola produced from artificially contaminated raw ewe milk. A challenge test was carried out using ewe raw milk inoculated with L. monocytogenes, in order to produce Pecorino di Farindola cheese stored at 18 degrees C for 149 days of ripening. During the ripening period, pH and a(w) values decreased in all samples analysed; lactic acid bacteria become the prevailing microbial population, while for L. monocytogenes a period of stability (neither growth nor death) followed the initial situation. The growth inhibition and the following inactivation may mostly be due to competition with the autochthonous microbiota and to the reduction of water activity. Mathematical modelling was used in order to predict microbial kinetics in the dynamic ripening environment, joining growth and death patterns in a continuous way, and including the highly uncertain growth/no growth range separating the two regions. The effect of lactic acid bacteria on the growth of pathogens was also included. Predicted microbial kinetics were satisfactory, as confirmed by the absence of statistically significant difference between observed and predicted values (p > 0.05). The present study proved, via challenge tests, that a dynamic growth/death model, previously used for a meat product, can be fruitfully used in cheese characterized by active competitive microbiota and progressive drying during ripening.
Purpose: To assess the consistency of the observed patterns of West Nile virus (WNV) during the last six years and possible drivers for the observed increased incidence in 2018. Methods & Materials: The data on confirmed West Nile Neuro-invasive Disease (WNND) human cases notified in Italy since 2012 indicate an increase of incidence in 2018 (131 cases as of 30 August). An integrated surveillance system is in place in Italy since 2008, which includes RT-PCR testing of mosquito pools, birds belonging to three target species (magpie, hooded crow, jay) and wild birds of other species found sick/dead. Data from veterinary activities are recorded in a national database, which is integrated with the data on WNND human cases. The possible correlation between the monthly numbers of positive mosquito pools and the monthly numbers of WNND cases, of positive birds of target species, of positive wild birds since 2012 have been tested. Results: The monthly numbers of WNND human cases, positive birds of target species and positive wild birds were all significantly correlated to monthly numbers of positive mosquito pools: Kendall correlation coefficients (tau) equal to, respectively, 0.6466 (p = 5.6e-11), 0.6983 (p = 4.3e-13) and 0.5548 (p = 1.5e-8). Conclusion: The variations of WNV infection in mosquitoes led to similar variations in the incidence in both the vertebrate reservoirs of infection (birds) and in the accidental hosts (human beings), thus supporting the hypothesis that the increased number of human cases observed so far in Italy in 2018 is linked to a higher incidence of infection in the mosquito populations. The analysis of climatic patterns in 2018 is indicating that during the first six months of this year the temperatures were higher than usual (+1.1 °C), and also rainfalls were more frequent, especially during June (http://www.meteo.it/clima-italia-2018-pioggia-caldo-temperature/). These differences in temperatures and rainfall patterns in 2018 could explain the observed increased incidence of cases in vertebrate hosts. Further analyses are ongoing to verify this hypothesis using a more consolidate dataset, to be expected in October, after the peak of WNV infection.
Bovine viral diarrhea (BVD) is a viral disease that affects cattle and that is endemic to many European countries. It has a markedly negative impact on the economy, through reduced milk production, abortions, and a shorter lifespan of the infected animals. Cows becoming infected during gestation may give birth to Persistently Infected (PI) calves, which remain highly infective throughout their life, due to the lack of immune response to the virus. As a result, they are the key driver of the persistence of the disease both at herd scale, and at the national level. In the latter case, the trade-driven movements of PIs, or gestating cows carrying PIs, are responsible for the spatial dispersion of BVD. Past modeling approaches to BVD transmission have either focused on within-herd or between-herd transmission. A comprehensive portrayal, however, targeting both the generation of PIs within a herd, and their displacement throughout the country due to trade transactions, is still missing. We overcome this by designing a multiscale metapopulation model of the spatial transmission of BVD, accounting for both within-herd infection dynamics, and its spatial dispersion. We focus on Italy, a country where BVD is endemic and seroprevalence is very high. By integrating simple within-herd dynamics of PI generation, and the highly-resolved cattle movement dataset available, our model requires minimal arbitrary assumptions on its parameterization. We use our model to study the role of the different productive contexts of the Italian market, and test possible intervention strategies aimed at prevalence reduction. We find that dairy farms are the main drivers of BVD persistence in Italy, and any control strategy targeting these farms would lead to significantly higher prevalence reduction, with respect to targeting other production compartments. Our multiscale metapopulation model is a simple yet effective tool for studying BVD dispersion and persistence at country level, and is a good instrument for testing targeted strategies aimed at the containment or elimination of this disease. Furthermore, it can readily be applied to any national market for which cattle movement data is available.
The confined environment of the dog shelter, particularly over extensive time-periods can impact severely on welfare. Surveillance and assessment are therefore essential components of the welfare protocol. The aim of this study was to generate a descriptive analysis of a sample of Italian long-term shelters and identify potential hazards regarding the welfare of shelter dogs. This was achieved through application of the Shelter Quality Protocol (SQP) to link income/outcome variables and the inclusion of sixty-four long-term shelters in Italy. Descriptive and logistic regression analyses were conducted. Key findings showed feeding regime, type of diet and access to outdoor area to be significantly associated with inadequate body condition score (BCS). The probability of observing skin lesions was shown to be influenced by bedding inadequacy and bedding type. Limiting beds to one per dog and utilising clean bedding materials was significantly associated with a reduced probability of observing dirty/wet dogs. Protection from adverse weather conditions and inadequate bedding were significantly associated with the manifestation of polypnea. Non-existent dog training facilities, outdoor access or leash walking were all found to significantly increase the likelihood of fearful or aggressive attitudes to people. Outdoor access also, in conjunction with feeding regime, was associated with the presence of diarrhoea. The SQP proved useful in identifying welfare hazards, both as regards shelter environment and shelter management. Identification of these hazards creates the opportunity for interventions to be applied, minimising the risks and improving the welfare of long-term shelter dogs.
Emerging and re-emerging infectious diseases are a significant public and animal health threat. In some zoonosis, the early detection of virus spread in animals is a crucial early warning for humans. The analyses of animal surveillance data are therefore of paramount importance for public health authorities to identify the appropriate control measure and intervention strategies in case of epidemics. The interaction among host, vectors, pathogen and environment require the analysis of more complex and diverse data coming from different sources. There is a wide range of spatiotemporal methods that can be applied as a surveillance tool for cluster detection, identification of risk areas and risk factors and disease transmission pattern evaluation. However, despite the growing effort, most of the recent integrated applications still lack of managing simultaneously different datasets and at the same time making available an analytical tool for a complete epidemiological assessment. In this paper, we present EpiExploreR, a user-friendly, flexible, R-Shiny web application. EpiExploreR provides tools integrating common approaches to analyze spatiotemporal data on animal diseases in Italy, including notified outbreaks, surveillance of vectors, animal movements data and remotely sensed data. Data exploration and analysis results are displayed through an interactive map, tables and graphs. EpiExploreR is addressed to scientists and researchers, including public and animal health professionals wishing to test hypotheses and explore data on surveillance activities.
Event Abstract Back to Event C. imicola occurrence prediction in Italy using Machine-learning and satellite data Luca Candeloro1*, Romolo Salini1, M Goffredo1, M Quaglia1 and Annamaria Conte1 1 Experimental Zooprophylactic Institute of Abruzzo and Molise G. Caporale, Italy In recent years, the availability of satellite-derived products is increased either in quantity or in quality and consequently their use in species distribution modelling. At the same time, machine-learning approaches showing great performance in terms of prediction accuracy have been developed and used in several fields. Environmental factors such as temperature and water accessibility due to precipitation are the most important drivers for vector’s life cycle, like mosquitos and Culicoides. Culicoides imicola is the most important Bluetongue vector in the Mediterranean basin and its distribution (in Italy) has been successfully modelled using satellite data and classical approaches like spatial logistic regression (Conte et al. 2007, Thibaut et al. 2012). However, accurate prediction in both space and time is still a big challenge because the interaction of environmental factors, along with their evolution in time, significantly affect the specie’s presence through complex relationships. The present work aims to answer the questions: will ML algorithms be able to improve C. imicola’s occurrence prediction accuracy in space and time, using freely available and commonly used satellite data? What is the achievable accuracy, and can we trust it as reliable? Data about Culicoides imicola presence, along with time and geographical coordinates, where derived from the entomological surveillance plan in place since 2000 (Goffredo and Meiswinkel, 2004). The dataset includes about 4,500 catch sites repeatedly sampled in time (although not evenly) since 2000, resulting in around 150,000 catches (12,000 of which positive for C. imicola presence). Selected satellite data (MOD11A2 day and night land surface temperature - 1km spatial resolution and 8 days temporal resolution; MOD13Q1 enhanced and normalized vegetation index – 250 m spatial resolution and 16 days temporal resolution; TRMM precipitation data - 1° longlat spatial resolution and 1 day temporal resolution) were downloaded and processed to account for their availability and difference in space time resolution. To account for satellite data availability, we considered only catches performed after May 2001. Satellite images falling into a 80 days back window were selected for each catch, resulting in 10 images for LSTD, 10 for LSTN, 10 for TRMM, 5 for EVI, 5 for NDVI, for a total of 40 variables. Two further variables were included along with satellite values: longitude and latitude catch sites’ coordinates. The number of catches included in the final dataset, however, was further decreased (for a final dataset of 91,467 records) by missing data in satellite images (mostly due to cloud cover). A graphical description of the final C. imicola dataset used in the study, in terms of seasonality and spatial distribution is shown in Fig.1. The study has been structured into two steps. In the first one, we evaluate the performance of four machine-learning algorithms (random forest, xgBoost tree, k-nearest neighbor and multi-layer perceptron (MLP)). Data was randomly split into train and test data (about 80% and 20% respectively) accounting for prevalence and avoiding spatial overlapping between each other (as catch sites). Given data imbalance and to prevent from over-fitting, down-sampling negatives inside a ten-fold cross validation (ten repetition) was used for all algorithms except for MLP (in which case imbalance was managed during network training using a weighted loss function). Hyper-parameters tuning was implemented through grid search or random search and AUC metric was used for model selection (Andrew). Model’s performance were measured in terms of AUC, accuracy, sensitivity and specificity, on the test dataset. The best resulting model capability in predicting the occurrence in space (almost one occurrence in time) was also evaluated. Given the strong seasonality of C. imicola presence (Fig.1), we investigated the spatial distribution of the probability of presence in spring (critical transition period when the abundance start arising after winter) and in autumn (period when the abundance reaches its maximum). Using the best performing model we created two raster of the probability of C. imicola presence for the whole country (1 km spatial resolution) at the 2018/04/01 and 2018/10/15. In the second step, we use the best model previously identified, to evaluate the capability of the model in generalizing results, decreasing both spatial and temporal link between training and test. Given the availability of data, we applied a more restrictive criteria to the train-test splitting procedure. Firstly, we ensured an not overlapping selection of sites creating wider regions (in the first step points do not overlap, but might be very close each other). For this purpose, the Italian territory was divided into 19 cells (2.5 longlat degree spatial resolution) to be sampled independently. Secondly, to avoid temporal overlapping, data until 2015 were used for training whilst data later than 2016 for test. A brute-force search among all possible combinations of picking from six to nine cells into test dataset was performed using the two non-overlapping spatio and temporal criteria, resulting in about 140,000 possible splits. However, most of those were extremely unbalanced in terms of prevalence and test train dimension. Among all, we choose those having a difference in prevalence between test and train less than 2% and a test data dimension ranging from 22.5% and 27.5%. This process of data splitting and selection resulted in 277 feasible sample splits (Fig. 2 shows the results of this procedure). Finally, we ran the best model on the selected sample splits and evaluated the predictive capability on the test dataset summarizing distribution of Sensibility, Specificity and AUC. Because of the extremely exacerbated sampling design, we trust the model must be able in catching factors really driving the presence of C. imicola. Variable importance was ranked for each model and summarized as overall importance through the median ranking. All analyses were performed using R (R Core Team, 2019). MODIS and raster packages (Mattiuzzi 2018, Hijmans 2018) were used to download and process satellite data. dplyr package (Wickham 2018) was used to prepare data for ML algorithm. doSNOW package (Weston 2017) was used to parallelize tasks. All algorithms were performed using caret package (Kuhn 2018), whilst MLP was implemented using keras (Allaire 2018). In the first step, classification was well performed by all methods, as the AUC values ranged from 0.982 (knn) to 0.984 (Xgboost) (Fig. 3 panel a). Xgboost tree algorithm showed a great performance, being at the same time the less time consuming, making it the best model to be used. Spatial prediction (neglecting the timing of catches) using Xlgboost tree model is shown in Fig. 3, panel b for absence (on the left) and presence (on the right) sites in the test dataset. Red and green dots represent presence and absence prediction respectively. Class prediction in space well classified real positive sites except for a few points (sensitivity equals to 89%). Those false negative points, however, are located in high prevalence areas. As expected, there are several real negative sites predicted as positive (positive predictive value equals to 45%). Fig.4 shows the predicted probability of occurrence for the 1st of April and the 15th of October 2018. The prediction seems to catch well the of C. imicola seasonality trend. The characteristics of Xgboost tree algorithm made it really suitable for the second step study, where the training and testing process was repeated for all the selected sample splits (277 times). Distribution of AUC, Se and Sp are shown in Fig.5 panel a). Despite the exacerbated sampling design, the method reached a median value of AUC, Se and Sp of 0.9, 0.8 and 0.92 respectively. Variable importance ranking (Fig. 5 panel b) showed how latitude and longitude are the most important variables when predicting C. imicola, followed by the night land surface temperature with the highest lag considered (LSTN-10). Moreover looking at grouped variables (independently from the temporal lag), it is evident the importance of the temperature group in comparison with the TRMM. The study highlighted the advantages of using ML methods when predicting C. imicola presence in space and time in Italy. The adopted experimental design ensured reliable results and reinforced the capability of the trained model in extending performance to never seen data neither in space or in time. This is particularly useful for targeting surveillance programs, discovering the presence of vectors in unchecked regions and for predicting future occurrence. Future work will see the implementation of such trained model in a near real time web distributed GIS framework to allow policy maker targeted decisions related to BT surveillance in order to reduce the sampling rate and consequently the cost and resources. Such a method might be successfully implemented for other vectors (mosquitos, ticks, etc.) and other countries wherever data are abundantly available, even in selected regions. Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 References 1. Conte A., Goffredo M., Ippoliti C. and Meiswinkel R., 2007. Influence of biotic and abiotic factors on the distribution and abundance of Culicoides imicola and the Obsoletus Complex in Italy. Vet. Par. Vol 150/4 pp 333-344. 2. Goffredo M., Meiswinkel R., 2004. Entomological surveillance of bluetongue in Italy: methods of capture, catch analysis and identification of Culicoides biting midges. Vet. Ital. 40, 260–265. 3. Andrew P.Bradley, 1997. The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recognition Volume 30, Issue 7, July 1997, Pages 1145-1159 4. R Core Team (2019). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL https://www.R-project.org/. 5. Matteo Mattiuzzi and Florian Detsch (2018). MODIS: Acquisition and Processing of MODIS Products. R package version 1.1.4. 6. Robert J. Hijmans (2018). raster: Geographic Data Analysis and Modeling. R package version 2.8-4. 7. Hadley Wickham, Romain François, Lionel Henry and Kirill Müller (2018). dplyr: A Grammar of Data Manipulation. R package version 0.7.8. 8. Microsoft Corporation and Stephen Weston (2017). doSNOW: Foreach Parallel Adaptor for the 'snow' Package. R package version 1.0.16. 9. Max Kuhn. Contributions from Jed Wing, Steve Weston, Andre Williams, Chris Keefer, Allan Engelhardt, Tony Cooper, Zachary Mayer, Brenton Kenkel, the R Core Team, Michael Benesty, Reynald Lescarbeau, Andrew Ziem, Luca Scrucca, Yuan Tang, Can Candan and Tyler Hunt. (2018). caret: Classification and Regression Training. R package version 6.0-81. 10. JJ Allaire and François Chollet (2018). keras: R Interface to 'Keras'. R package version 2.2.4. Keywords: machine learning, satellite data, space - time prediction, accuracy, C. imicola Conference: GeoVet 2019. Novel spatio-temporal approaches in the era of Big Data, Davis, United States, 8 Oct - 10 Oct, 2019. Presentation Type: Regular oral presentation Topic: Spatio-temporal surveillance and modeling approaches Citation: Candeloro L, Salini R, Goffredo M, Quaglia M and Conte A (2019). C. imicola occurrence prediction in Italy using Machine-learning and satellite data. Front. Vet. Sci. Conference Abstract: GeoVet 2019. Novel spatio-temporal approaches in the era of Big Data. doi: 10.3389/conf.fvets.2019.05.00112 Copyright: The abstracts in this collection have not been subject to any Frontiers peer review or checks, and are not endorsed by Frontiers. They are made available through the Frontiers publishing platform as a service to conference organizers and presenters. The copyright in the individual abstracts is owned by the author of each abstract or his/her employer unless otherwise stated. Each abstract, as well as the collection of abstracts, are published under a Creative Commons CC-BY 4.0 (attribution) licence (https://creativecommons.org/licenses/by/4.0/) and may thus be reproduced, translated, adapted and be the subject of derivative works provided the authors and Frontiers are attributed. For Frontiers’ terms and conditions please see https://www.frontiersin.org/legal/terms-and-conditions. Received: 10 Jun 2019; Published Online: 27 Sep 2019. * Correspondence: Dr. Luca Candeloro, Experimental Zooprophylactic Institute of Abruzzo and Molise G. Caporale, Teramo, Abruzzo, Italy, l.candeloro@izs.it Login Required This action requires you to be registered with Frontiers and logged in. To register or login click here. Abstract Info Abstract Supplemental Data The Authors in Frontiers Luca Candeloro Romolo Salini M Goffredo M Quaglia Annamaria Conte Google Luca Candeloro Romolo Salini M Goffredo M Quaglia Annamaria Conte Google Scholar Luca Candeloro Romolo Salini M Goffredo M Quaglia Annamaria Conte PubMed Luca Candeloro Romolo Salini M Goffredo M Quaglia Annamaria Conte Related Article in Frontiers Google Scholar PubMed Abstract Close Back to top Javascript is disabled. Please enable Javascript in your browser settings in order to see all the content on this page.
Ecoregionalization is the process by which a territory is classified in similar areas according to specific environmental and climatic factors. The climate and the environment strongly influence the presence and distribution of vectors responsible for significant human and animal diseases worldwide. In this paper, we developed a map of the eco-climatic regions of Italy adopting a data-driven spatial clustering approach using recent and detailed spatial data on climatic and environmental factors. We selected seven variables, relevant for a broad set of human and animal vector-borne diseases (VBDs): standard deviation of altitude, mean daytime land surface temperature, mean amplitude and peak timing of the annual cycle of land surface temperature, mean and amplitude of the annual cycle of greenness value, and daily mean amount of rainfall. Principal Component Analysis followed by multivariate geographic clustering using the k-medoids technique were used to group the pixels with similar characteristics into different ecoregions, and at different spatial resolutions (250 m, 1 km and 2 km). We showed that the spatial structure of ecoregions is generally maintained at different spatial resolutions and we compared the resulting ecoregion maps with two datasets related to Bluetongue vectors and West Nile Disease (WND) outbreaks in Italy. The known characteristics of Culicoides imicola habitat were well captured by 2/22 specific ecoregions (at 250 m resolution). Culicoides obsoletus/scoticus occupy all sampled ecoregions, according to its known widespread distribution across the peninsula. WND outbreak locations strongly cluster in 4/22 ecoregions, dominated by human influenced landscape, with intense cultivations and complex irrigation network. This approach could be a supportive tool in case of VBDs, defining pixel-based areas that are conducive environment for VBD spread, indicating where surveillance and prevention measures could be prioritized in Italy. Also, ecoregions suitable to specific VBDs vectors could inform entomological surveillance strategies.