The decision support systems that have been developed to assist physicians in the diagnostic process often are based on static data which may be out of date. We present a comprehensive analysis of artificial intelligent methods which could be applied to documents encoded by SNOMED CT. By mining information directly from SNOMED CT encoded documents, a decision support system could contain timely updated diagnostic information, which is of significant value in fast changing situations such as minimally understood emerging diseases and epidemics. Through a high level comparison of many AI methods it is found that a TAN-Bayesian method could be the most suitable to apply to SNOMED CT data.
This research paper presents the implementation of an adaptive learning algorithm from artificial intelligence known as Reinforcement Learning in a robot that must deal with more than one reward where the rewards may come into conflict with each other. A case that practically illustrates this problem is in a service industry environment where a robot is implemented to pick up objects for cleaning and/or attending to clients. An example of conflicting rewards for this application would be achieving a high reward for quickly picking up objects which typically conflicts with minimizing any damage inflicted on the object during the picking up process. The innovative component of our research is that we will use multiple competing rewards for some states in contrast to the single reward per state method traditionally used in Reinforcement Learning. We compare these two approaches through the implementation of a Lego Mindstorm robot that has been programmed with both learning methods. The objective of our robot is to pick up objects quickly without damaging the object. We illustrate the conditions under which it is advantageous to use a single state for competing rewards over a multi-state approach through practical comparisons on efficiency for our Lego robot. The objective of this research is to broaden the adaptability of learning robots. The impact on service industries, such as hotel and restaurant service, of this research would be to increase the acceptability of adaptable robots into fields of manual labor that have traditionally been limited due to the inflexibility of robots in dealing with dynamic situations.
This paper illustrates an automated system that replicates the investigative operation of human fraud auditors. Human fraud auditors often utilize fraud detection methods that exploit structure in database tables to uncover outliers that may be part of a fraud case. From the uncovered outliers, an auditor will build a case of fraud by searching data related to the outlier possibly across many different databases and tables within these different databases. This paper illustrates an industrial implementation of an adaptive fraud case building system that uses machine learning to conduct the search and decision-making process with an automated outlier detection component. This system was successfully applied to uncover fraud cases in real marketing data.
Adaptive Benford's Law [1] is a digital analysis technique that specifies the probabilistic distribution of digits for many commonly occurring phenomena, even for incomplete data records. We combine this digital analysis technique with a reinforcement learning technique to create a new fraud discovery approach. When applied to records of naturally occurring phenomena, our adaptive fraud detection method uses deviations from the expected Benford's Law distributions as an indicators of anomalous behaviour that are strong indicators of fraud. Through the exploration component of our reinforcement learning method we search for the underlying attributes producing the anomalous behaviour. In a blind test of our approach, using real health and auto insurance data, our Adaptive Fraud Detection method successfully identified actual fraudsters among the test data.
Benford's Law [1] specifies the probabilistic distribution of digits for many commonly occurring phenomena, ideally when we have complete data of the phenomena. We enhance this digital analysis technique with an unsupervised learning method to handle situations where data is incomplete. We apply this method to the detection of fraud and abuse in health insurance claims using real health insurance data. We demonstrate improved precision over the traditional Benford approach in detecting anomalous data indicative of fraud and illustrate some of the challenges to the analysis of healthcare claims fraud.
With the advent of Kearns & Singh's (2000) rigorous upper bound on the error of temporal difference estimators, we derive the first rigorous error bound for the maximum likelihood policy evaluation method as well as deriving a Monte Carlo matrix inversion policy evaluation error bound. We provide, the first direct comparison between the error bounds of the maximum likelihood (ML), Monte Carlo matrix inversion (MCMI) and temporal difference (TD) estimation methods for policy evaluation. We use these bounds to confirm generally held notions of the superior accuracy of the model-based estimation methods of ML and MCMI over the model-free method of TD. With our error bounds, we are also able to specify parameters and conditions that affect each method's estimation accuracy.
In the area of unsupervised reinforcement learning, where we model the environment as a Markov Decision Process, the goal is to gather information about this system to produce an optimal strategy for navigating this environment. The states of the environment have rewards associated with entering each state. An optimal policy for navigating this system will optimize the long-term rewards. Producing an optimal policy often requires the intermediate step of policy evaluation. A variety of methods have been used to perform policy evaluation, the most popular of which is Temporal Differencing. Temporal Differencing is an efficient model-free method that saves on storage space since no model of the environment is explicitly stored. Model-based methods of policy evaluation have generally been less popular because of perceived slower execution times and greater storage costs, especially as the state space size grows. This thesis counter-acts those limitations by demonstrating efficient model-based policy evaluation approaches with a linear in state space size storage cost and more accurate value estimates of policies than Temporal Difference methods. The thesis uses two model-based approaches, a maximum likelihood method for sparsely and densely connected networks and a matrix inversion approach for intermediate cases. As state space size grows, a least-squares approximation may be applied to these model-based methods.
In 1950, Forsythe and Leibler (1950) introduced a statistical technique for finding the inverse of a matrix by characterizing the elements of the matrix inverse as expected values of a sequence of random walks. Barto and Duff (1994) subsequently showed relations between this technique and standard dynamic programming and temporal differencing methods. The advantage of the Monte Carlo matrix inversion (MCMI) approach is that it scales better with respect to state-space size than alternative techniques. In this paper, we introduce an algorithm for performing reinforcement learning policy evaluation using MCMI. We demonstrate that MCMI possesses accuracy similar to a maximum likelihood model-based policy evaluation approach but avoids ML's slow execution time. In fact, we show that MCMI executes at a similar runtime to temporal differencing (TD). We then illustrate a least-squares generalization technique for scaling up MCMI to large state spaces. We compare this leastsquares Monte Carlo matrix inversion (LS-MCMI) technique to the least-squares temporal differencing (LSTD) approach introduced by Bradtke and Barto (1996) demonstrating that both LS-MCMI and LSTD have similar runtime.
The study of value estimation in Markov reward processes has been dominated by research on temporal difference methods since the introduction of TD(0) in 1988. Temporal difference methods are often contrasted with a maximum likelihood approach where the transition matrix and reward vector are estimated explicitly and converted into a value estimate by solving a matrix equation. It is often asserted that maximum likelihood estimation yields more accurate values, but the temporal difference...
The wedging action fixture to which the invention relates has a wedging action means, a pair of pressing members, wedging action releasing means and other members for limiting the relative movement and transmitting forces between these parts. This fixture can be secured to any desired portion of an elongated supporting member by a wedging action performed by the wedging action means. The external force to be borne is applied to the wedging action means so as to obtain a wedging force which is transmitted to a pair of pressing members. The pressing members acts on both surface of the elongated supporting member so as to cramp the latter therebetween or, alternatively, on both opposing walls of a channeled supporting member to urge the walls away from each other, thereby to bear the external force at any desired position on the elongated supporting member. The fixture can easily be unfastened simply by operating the wedging action releasing means, so that it can be easily moved to any desired position on the elongated supporting member or detached from the latter.