Continuous-discrete state space models (CDSSMs) enable modelling of irregular time series by learning the underlying continuous-time dynamics with noisy measurements obtained at discrete timestamps. Recent studies have shown remarkable performance involving CDSSMs with neural networks. However, challenges still remain in the application of general non-linear non-Gaussian CDSSMs to irregular time series. To address these challenges, we propose a new method, named continuous-discrete differentiable particle filters (CD-DPFs), to model probabilistic irregular time series. Representing the latent state probability density function by a Gaussian mixture model (GMM), an adaptive Gaussian sum particle filter is applied into CDSSMs, and the weights of the GMM are optimised adaptively through solving a convex optimisation problem by direct matching the Fokker-Planck-Kolmogorov equation. Performance is evaluated on a stochastic Lorenz 63 model, a highly non-linear chaotic system. Compared with the state-of-the-art, the proposed method demonstrates significant improvement for forecasting whilst maintaining competitive performance on imputation.
We are seeing a steady build up in momentum of two trends that will lead to a significant, and necessary, transformation of agriculture in the 21st Century. Firstly, the move to digital with IoT facilitating the use of sensor networks to support precise decision making. Secondly, the move to "ecological intensification"; working with natural processes to lower the carbon footprint of agricultural processes and increase the biodiversity of agricultural units without sacrificing yield. Clearly data mining and machine learning have an important role to play in supporting this transformation. However, given the range of biotic, abiotic and social contexts that need to inform the development of models for decision support in agriculture, we need data mining techniques that support the use of qualitative and unstructured data as well as hard numerical data. In this tutorial we show how Bayesian Networks can be built from a wide range of sources of expert knowledge and data. Whist we use digital agriculture as a key beneficiary of these techniques, this tutorial will also be of interest to all those with an interest in building IoT systems that require combinations of social, technical and environmental understanding.
Plant diseases are one of the main causes of crop loss in agriculture. Machine Learning, in particular statistical and neural nets (NNs) approaches, have been used to help farmers identify plant diseases. However, since new diseases continue to appear in agriculture due to climate change and other factors, we need more data-efficient approaches to identify and classify new diseases as early as possible. Even though statistical machine learning approaches and neural nets have demonstrated state-of-the-art results on many classification tasks, they usually require a large amount of training data. This may not be available for emergent plant diseases. So, data-efficient approaches are essential for an early and precise diagnosis of new plant diseases and necessary to prevent the disease’s spread. This study explores a data-efficient Inductive Logic Programming (ILP) approach for plant disease classification. We compare some ILP algorithms (including our new implementation, PyGol) with several statistical and neural-net based machine learning algorithms on the task of tomato plant disease classification with varying sizes of training data set (6, 10, 50 and 100 training images per disease class). The results suggest that ILP outperforms other learning algorithms and this is more evident when fewer training data are available.
Large Language Models (LLMs) can appear to generate expert advice on legal matters. However, at closer analysis, some of the advice provided has proven unsound or erroneous. We tested LLMs’ performance in the procedural and technical area of insolvency law in which our team has relevant expertise. This paper demonstrates that statistically more accurate results to evaluation questions come from a design which adds a curated knowledge base to produce quality responses when querying LLMs. We evaluated our bot head-to-head on an unseen test set of twelve questions about insolvency law against the unmodified versions of gpt-3.5-turbo and gpt-4 with a mark scheme similar to those used in examinations in law schools. On the “unseen test set”, the Insolvency Bot based on gpt-3.5-turbo outper-formed gpt-3.5-turbo (p = 1.8%), and our gpt-4 based bot outperformed unmodified gpt-4 (p = 0.05%). These promising results can be expanded to cross-jurisdictional queries and be further improved by matching on-point legal information to user queries. Overall, they demonstrate the importance of incorporating trusted knowledge sources into traditional LLMs in answering domain-specific queries.
Purpose To compare supervised transfer learning to semisupervised learning for their ability to learn in-depth knowledge with limited data in the optical coherence tomography (OCT) domain. Methods Transfer learning with EfficientNet-B4 and semisupervised learning with SimCLR are used in this work. The largest public OCT dataset, consisting of 108,312 images and four categories (choroidal neovascularization, diabetic macular edema, drusen, and normal) is used. In addition, two smaller datasets are constructed, containing 31,200 images for the limited version and 4000 for the mini version of the dataset. To illustrate the effectiveness of the developed models, local interpretable model-agnostic explanations and class activation maps are used as explainability techniques. Results The proposed transfer learning approach using the EfficientNet-B4 model trained on the limited dataset achieves an accuracy of 0.976 (95% confidence interval [CI], 0.963, 0.983), sensitivity of 0.973 and specificity of 0.991. The semisupervised based solution with SimCLR using 10% labeled data and the limited dataset performs with an accuracy of 0.946 (95% CI, 0.932, 0.960), sensitivity of 0.941, and specificity of 0.983. Conclusions Semisupervised learning has a huge potential for datasets that contain both labeled and unlabeled inputs, generally, with a significantly smaller number of labeled samples. The semisupervised based solution provided with merely 10% labeled data achieves very similar performance to the supervised transfer learning that uses 100% labeled samples. Translational Relevance Semisupervised learning enables building performant models while requiring less expertise effort and time by using to good advantage the abundant amount of available unlabeled data along with the labeled samples.
In this paper, we present a Business Analytics (BA) framework, which addresses the challenge of analysing primary care outcomes for both patients and clinicians from multiple data sources in an accurate manner. A review of the process monitoring literature has been conducted in the context of healthcare management and decision making and its findings have informed the formulation of a BA conceptual framework for process monitoring and decision support in primary care. Furthermore, a real case study is conducted to demonstrate the application of the BA framework to implement a BA dashboard tool within one of the largest primary care providers in England. Findings: The main contributions of the presented work are the development of a conceptual BA framework and a BA dashboard tool to support management and decision making in primary care. This was evaluated through a case study of the implementation of the BA dashboard tool in London’s largest primary care provider. This BA tool provides real-time information to enable simpler decision-making processes and to inform business transformation in a number of areas. The resulting increased efficiency has led to significant cost savings and improved delivery of patient care.
Recent advances in smart connected vehicles and Intelligent Transportation Systems (ITS) are based upon the capture and processing of large amounts of sensor data. Modern vehicles contain many internal sensors to monitor a wide range of mechanical and electrical systems and the move to semi-autonomous vehicles adds outward looking sensors such as cameras, lidar, and radar. ITS is starting to connect existing sensors such as road cameras, traffic density sensors, traffic speed sensors, emergency vehicle, and public transport transponders. This disparate range of data is then processed to produce a fused situation awareness of the road network and used to provide real-time management, with much of the decision making automated. Road networks have quiet periods followed by peak traffic periods and cloud computing can provide a good solution for dealing with peaks by providing offloading of processing and scaling-up as required, but in some situations latency to traditional cloud data centres is too high or bandwidth is too constrained. Cloud computing at the edge of the network, close to the vehicle and ITS sensor, can provide a solution for latency and bandwidth constraints but the high mobility of vehicles and heterogeneity of infrastructure still needs to be addressed. This paper surveys the literature for cloud computing use with ITS and connected vehicles and provides taxonomies for that plus their use cases. We finish by identifying where further research is needed in order to enable vehicles and ITS to use edge cloud computing in a fully managed and automated way. We surveyed 496 papers covering a seven-year timespan with the first paper appearing in 2013 and ending at the conclusion of 2019.
In this final chapter, we outline a vision for a technological revolution in agriculture that would work to regain a sense of balance between food production and natural ecosystems. We promote an ecological engineering approach to crop production that draws on experience from the organic and conservation agriculture movements. However, we expand on this by promoting in addition: 1. An Internet of Things enabled biomonitoring system that enables key (above and below ground) environmental indicators to be automatically monitored across an agricultural unit. 2. A combination of network and thermodynamic ecosystem modelling approaches to enable a deep understanding of the response of ecosystem service functioning to changes in biodiversity, and the abiotic context, in any given agroecosystem. We also call for an explicit recognition of the “ethnosphere” (the sphere of human social and cultural experience) as a fifth geosphere which emerged from the biosphere, and the other three geospheres, and whose continued existence is therefore contingent on the health and stability of the other four geospheres.
In this chapter, we study some research issues from IoT-based spectrum trading in Wireless Communication in a strategic setting. We consider the scenario in which there are multiple secondary users (such as non-governmental organizations (NGOs), institutional organizations, foundations, etc.) having available un-utilized spectrum and multiple tertiary users (such as small farms, agricultural enterprises or people residing in different localities). Tertiary users provide preferences over the subset of all the available secondary users (NGOs, hereafter). Based on their preference ordering, the tertiary users are allocated the best possible NGOs among the available ones and under the restrictions that each user is assigned to at most one NGO. However, it is to be noted that, in this model, the allocated spectrum might not be available through out a long period of time but rather for a short duration of time within a time period. Therefore, tertiary users have to be able to work off-line and access cached data at the Edges of Internet even if the Internet access is not available. For the purpose of storing and retrieving the cached data several algorithms are designed and their computational complexity is analyzed. In order to empirically measure the efficacy of the proposed mechanisms the simulations are carried out and are compared with the benchmark mechanism. The proposed allocation mechanisms are envisaged as especially useful tools for emerging scenarios of smart farming and precision agriculture, where in situ infrastructures are not available.
BACKGROUND:In diabetic retinopathy (DR) screening programmes feature-based grading guidelines are used by human graders. However, recent deep learning approaches have focused on end to end learning, based on labelled data at the whole image level. Most predictions from such software offer a direct grading output without information about the retinal features responsible for the grade. In this work, we demonstrate a feature based retinal image analysis system, which aims to support flexible grading and monitor progression.METHODS:The system was evaluated against images that had been graded according to two different grading systems; The International Clinical Diabetic Retinopathy and Diabetic Macular Oedema Severity Scale and the UK's National Screening Committee guidelines.RESULTS:External evaluation on large datasets collected from three nations (Kenya, Saudi Arabia and China) was carried out. On a DR referable level, sensitivity did not vary significantly between different DR grading schemes (91.2-94.2.0%) and there were excellent specificity values above 93% in all image sets. More importantly, no cases of severe non-proliferative DR, proliferative DR or DMO were missed.CONCLUSIONS:We demonstrate the potential of an AI feature-based DR grading system that is not constrained to any specific grading scheme.
Soil health is an environmental factor that impacts a range of important issues including food production, water retention and soil organic carbon storage. Agriculture relies on healthy soil for crop growth and animal grazing. Water retention reduces the risks of desertification and of flooding as the capacity to retain water reduces the rate of surface water flow. In addition, soil organic carbon represents the largest terrestrial carbon stock and is second only to the oceans. Yet, soil health is threatened by intensive farming practices and changes of land use such as deforestation. Thus, it is important to manage soil health to maintain food security, avoid desertification and maintain or ideally increase soil organic carbon storage. A useful tool to inform this management function would be a machine learning model that can predict soil health given land cover and parameters of the abiotic context, such as terrain elevation and historical weather data. The first step in developing such a model is to be able to identify the land cover for a chosen area. Satellites provide multi-spectral images that include the visual bands. Land cover databases provide the ground truth labels for a supervised learning approach to train an image semantic segmentation model. This chapter describes how Sentinel-2 satellite image data was combined with data from the UK Centre for Ecology and Hydrology Land Cover Map 2015, to train a convolutional neural network for land cover classification for the South of England.
In this introductory chapter we highlight some fundamental concepts, architectures and definitions related to IoT-based Computational Modeling for Next Generation Agro-ecosystems. We distinguish and discus the paradigms of Cloud-to-thing Continuum as a large digital ecosystem comprising IoT, Edge, Fog, and Cloud Computing, data cycles from data gathering, processing and analysis to knowledge generation and decision making. Machine learning and stream processing, optimization, simulation frameworks, symbiotic modeling and the digital twin as well as emerging research trends, ethics and health & safety issues are also introduced and discussed. Challenges arising from processing and analyzing large and heterogeneous data sets are pointed out with some examples of killer applications from Agriculture 4.0. These concepts, models, technologies, frameworks and benchmarks are covered in the chapters of the book and exemplified with real life use cases and applications. In all, Cloud-to-thing continuum and IoT as part of it are envisaged as game changers for Computational Modeling of Next Generation Agro-ecosystems.
In the Internet of Things (IoT) + Fog + Cloud architecture, with the unprecedented growth of IoT devices, one of the challenging issues that needs to be tackled is to allocate Fog service providers (FSPs) to IoT devices, especially in a game-theoretic environment. Here, the issue of allocation of FSPs to the IoT devices is sifted with game-theoretic idea so that utility maximizing agents may be benign. In this scenario, we have multiple IoT devices and multiple FSPs, and the IoT devices give preference ordering over the subset of FSPs. Given such a scenario, the goal is to allocate at most one FSP to each of the IoT devices. We propose mechanisms based on the theory of mechanism design without money to allocate FSPs to the IoT devices. The proposed mechanisms have been designed in a flexible manner to address the long and short duration access of the FSPs to the IoT devices. For analytical results, we have proved the economic robustness, and probabilistic analyses have been carried out for allocation of IoT devices to the FSPs. In simulation, mechanism efficiency is laid out under different scenarios with an implementation in Python.
Purpose The aim of this work is to demonstrate how a retinal image analysis system, DAPHNE, supports the optimization of diabetic retinopathy (DR) screening programs for grading color fundus photography. Method Retinal image sets, graded by trained and certified human graders, were acquired from Saudi Arabia, China, and Kenya. Each image was subsequently analyzed by the DAPHNE automated software. The sensitivity, specificity, and positive and negative predictive values for the detection of referable DR or diabetic macular edema were evaluated, taking human grading or clinical assessment outcomes to be the gold standard. The automated software's ability to identify co-pathology and to correctly label DR lesions was also assessed. Results In all three datasets the agreement between the automated software and human grading was between 0.84 to 0.88. Sensitivity did not vary significantly between populations (94.28%–97.1%) with specificity ranging between 90.33% to 92.12%. There were excellent negative predictive values above 93% in all image sets. The software was able to monitor DR progression between baseline and follow-up images with the changes visualized. No cases of proliferative DR or DME were missed in the referable recommendations. Conclusions The DAPHNE automated software demonstrated its ability not only to grade images but also to reliably monitor and visualize progression. Therefore it has the potential to assist timely image analysis in patients with diabetes in varied populations and also help to discover subtle signs of sight-threatening disease onset. Translational Relevance This article takes research on machine vision and evaluates its readiness for clinical use.
In this chapter, we study some research issues from IoT-based crowdsourcing in a strategic setting. We have considered the scenario in IoT-based crowdsourcing, where there are multiple task requesters and multiple IoT devices as task executors. Each task requester has multiple tasks, with the tasks having start and finish times. Based on the start and finish times, the tasks are to be distributed into different slots. On the other hand, in each slot, each IoT device requests for the set of tasks that it wants to execute along with the valuation that it will charge in exchange for its service. Both the requested set of tasks and the valuations are private informations. Given such scenario, the objective is to allocate the subset of IoT devices to the tasks in a non-conflicting manner with the objective of maximizing the social welfare. For the purpose of determining the unknown quality of the IoT devices we have utilized the concept of peer grading. Therefore, we have designed a truthful mechanism for the problem under investigation that also allows us to have the true information about the quality of the IoT devices.
The fast development of Internet of Things (IoT) computing and technologies has prompted a decentralization of Cloud-based systems. Indeed, sending all the information from IoT devices directly to the Cloud is not a feasible option for many applications with demanding requirements on real-time response, low latency, energy-aware processing and security. Such decentralization has led in a few years to the proliferation of new computing layers between Cloud and IoT, known as Edge computing layer, which comprises of small computing devices (e.g. Raspberry Pi) to larger computing nodes such as Gateways, Road Side Units, Mini Clouds, MEC Servers, Fog nodes, etc. In this paper, we study the challenges of processing an IoT data stream in an Edge computing layer. By using a real life data stream set arising from a car data stream as well as a real infrastructure using Raspberry Pi and Node-Red server, we highlight the complexities of achieving real time requirements of applications based on IoT stream processing.
This study evaluated the effectiveness and usability of a developed collaborative online tool (chit-chat) for children with Attention Deficit Hyperactivity Disorder (ADHD). We studied whether this tool influenced children’s Knowledge and experience exchange, motivation, behavioral abilities and social skills while using another learning tool, ACTIVATE. A total of seven Saudi children with ADHD aged from 6 to 8 years were assigned to the collaborative intervention using iPads. They were asked to play mini games that positively affect children with ADHD cognitively and behaviorally, then chat using our developed collaborative online tool, for three sessions. Progress points were measured and quantitatively analyzed before and after the intervention, thematic analysis was applied on the qualitative data. Participants showed improvements in overall performance when using the learning tool ACTIVATE. E-collaboration was found to be effective to children with ADHD and positively influencing their knowledge, experience, motivation and social skills.
Information theory has gained application in a wide range of disciplines, including statistical inference, natural language processing, cryptography and molecular biology. However, its usage is less pronounced in medical science. In this chapter, we illustrate a number of approaches that have been taken to applying concepts from information theory to enhance medical decision making. We start with an introduction to information theory itself, and the foundational concepts of information content and entropy. We then illustrate how relative entropy can be used to identify the most informative test at a particular stage in a diagnosis. In the case of a binary outcome from a test, Shannon entropy can be used to identify the range of values of test results over which that test provides useful information about the patient's state. This, of course, is not the only method that is available, but it can provide an easily interpretable visualization. The chapter then moves on to introduce the more advanced concepts of conditional entropy and mutual information and shows how these can be used to prioritise and identify redundancies in clinical tests. Finally, we discuss the experience gained so far and conclude that there is value in providing an informed foundation for the broad application of information theory to medical decision making.
In the IoT+Fog+Cloud architecture, with the ever increasing growth of IoT devices, allocation of IoT devices to the Fog service providers will be challenging and needs to be addressed properly in both the strategic and non-strategic settings. In this paper, we have addressed this allocation problem in strategic settings. The framework is developed under the consideration that the IoT devices (e.g. wearable devices) collecting data (e.g. health statistics) are deploying it to the Fog service providers for some meaningful processing free of cost. Truthful and Pareto optimal mechanisms are developed for this framework and are validated with some simulations.