We analyze consumer defaults in a sample of 64 000 customers taking personal loans from a Korean bank. Applying a generalized additive modeling (GAM) framework, we show a nonlinear impact of loan and borrower characteristics. In particular, the likelihood of default is high for both low-income borrowers and high-income borrowers. Our results are robust to a range of different tests, and they highlight the usefulness of the GAM framework, especially the graphical presentation of nonlinearities.
Emergence of the Coronavirus 2019 Disease has highlighted further the need for timely support for clinicians as they manage severely ill patients. We combine Semantic Web technologies with Deep Learning for Natural Language Processing with the aim of converting human-readable best evidence/ practice for COVID-19 into that which is computer-interpretable. We present the results of experiments with 1212 clinical ideas (medical terms and expressions) from two UK national healthcare services specialty guides for COVID-19 and three versions of two BMJ Best Practice documents for COVID-19. The paper seeks to recognise and categorise clinical ideas, performing a Named Entity Recognition (NER) task, with an ontology providing extra terms as context and describing the intended meaning of categories understandable by clinicians. The paper investigates: 1) the performance of classical NER using MetaMap versus NER with fine-tuned BERT models;2) the integration of both NER approaches using a lightweight ontology developed in close collaboration with senior doctors;and 3) the easy interpretation by junior doctors of the main classes from the ontology once populated with NER results. We report the NER performance and the observed agreement for human audits. Copyright © 2022 for this paper by its authors.
Background How to treat a disease remains to be the most common type of clinical question. Obtaining evidence-based answers from biomedical literature is difficult. Analogical reasoning with embeddings from deep learning (embedding analogies) may extract such biomedical facts, although the state-of-the-art focuses on pair-based proportional (pairwise) analogies such as man:woman::king:queen (“queen = −man +king +woman”). Objective This study aimed to systematically extract disease treatment statements with a Semantic Deep Learning (SemDeep) approach underpinned by prior knowledge and another type of 4-term analogy (other than pairwise). Methods As preliminaries, we investigated Continuous Bag-of-Words (CBOW) embedding analogies in a common-English corpus with five lines of text and observed a type of 4-term analogy (not pairwise) applying the 3CosAdd formula and relating the semantic fields person and death: “dagger = −Romeo +die +died” (search query: −Romeo +die +died). Our SemDeep approach worked with pre-existing items of knowledge (what is known) to make inferences sanctioned by a 4-term analogy (search query −x +z1 +z2) from CBOW and Skip-gram embeddings created with a PubMed systematic reviews subset (PMSB dataset). Stage1: Knowledge acquisition. Obtaining a set of terms, candidate y, from embeddings using vector arithmetic. Some n-gram pairs from the cosine and validated with evidence (prior knowledge) are the input for the 3cosAdd, seeking a type of 4-term analogy relating the semantic fields disease and treatment. Stage 2: Knowledge organization. Identification of candidates sanctioned by the analogy belonging to the semantic field treatment and mapping these candidates to unified medical language system Metathesaurus concepts with MetaMap. A concept pair is a brief disease treatment statement (biomedical fact). Stage 3: Knowledge validation. An evidence-based evaluation followed by human validation of biomedical facts potentially useful for clinicians. Results We obtained 5352 n-gram pairs from 446 search queries by applying the 3CosAdd. The microaveraging performance of MetaMap for candidate y belonging to the semantic field treatment was F-measure=80.00% (precision=77.00%, recall=83.25%). We developed an empirical heuristic with some predictive power for clinical winners, that is, search queries bringing candidate y with evidence of a therapeutic intent for target disease x. The search queries -asthma +inhaled_corticosteroids +inhaled_corticosteroid and -epilepsy +valproate +antiepileptic_drug were clinical winners, finding eight evidence-based beneficial treatments. Conclusions Extracting treatments with therapeutic intent by analogical reasoning from embeddings (423K n-grams from the PMSB dataset) is an ambitious goal. Our SemDeep approach is knowledge-based, underpinned by embedding analogies that exploit prior knowledge. Biomedical facts from embedding analogies (4-term type, not pairwise) are potentially useful for clinicians. The heuristic offers a practical way to discover beneficial treatments for well-known diseases. Learning from deep learning models does not require a massive amount of data. Embedding analogies are not limited to pairwise analogies; hence, analogical reasoning with embeddings is underexploited.
Previous research has shown that people who are in poverty live in deprived neighbourhoods. Ethnic minority groups are more likely than the White majority to be poor and live in such areas. The likelihood of being poor may be reduced by having access to mixed social networks. But, for those living in deprived neighbourhoods there may be neither opportunities nor resources to form and maintain social networks that are mixed in terms of their ethnic or geographic composition. This paper tests this contention, for ethnic groups in the UK. Specifically, we use the UK's largest household survey to examine the relationship between deprivation and mixing by investigating the following research questions: (1) Does neighbourhood deprivation alter the influence of mixed social network on poverty status? and (2) Is the influence of neighbourhood deprivation and social networks on poverty status equivalent for all ethnic minority groups? Our results suggest that high neighbourhood deprivation tends to over-ride the positive associations of geographically mixed social networks. Moreover, while this result is strong for the White British majority, there is only weak evidence that it holds for ethnic minority groups. This may imply that resource constraints restrict social network benefits, particularly for ethnic minorities.
Background Deep Learning opens up opportunities for routinely scanning large bodies of biomedical literature and clinical narratives to represent the meaning of biomedical and clinical terms. However, the validation and integration of this knowledge on a scale requires cross checking with ground truths (i.e. evidence-based resources) that are unavailable in an actionable or computable form. In this paper we explore how to turn information about diagnoses, prognoses, therapies and other clinical concepts into computable knowledge using free-text data about human and animal health. We used a Semantic Deep Learning approach that combines the Semantic Web technologies and Deep Learning to acquire and validate knowledge about 11 well-known medical conditions mined from two sets of unstructured free-text data: 300 K PubMed Systematic Review articles (the PMSB dataset) and 2.5 M veterinary clinical notes (the VetCN dataset). For each target condition we obtained 20 related clinical concepts using two deep learning methods applied separately on the two datasets, resulting in 880 term pairs (target term, candidate term). Each concept, represented by an n-gram, is mapped to UMLS using MetaMap; we also developed a bespoke method for mapping short forms (e.g. abbreviations and acronyms). Existing ontologies were used to formally represent associations. We also create ontological modules and illustrate how the extracted knowledge can be queried. The evaluation was performed using the content within BMJ Best Practice. Results MetaMap achieves an F measure of 88% (precision 85%, recall 91%) when applied directly to the total of 613 unique candidate terms for the 880 term pairs. When the processing of short forms is included, MetaMap achieves an F measure of 94% (precision 92%, recall 96%). Validation of the term pairs with BMJ Best Practice yields precision between 98 and 99%. Conclusions The Semantic Deep Learning approach can transform neural embeddings built from unstructured free-text data into reliable and reusable One Health knowledge using ontologies and content from BMJ Best Practice.
The bootstrap appears to be well suited for performing accurate inference in medical cost-effectiveness analysis, where data are generally drawn from a randomized experiment. However, the literature to date has given priority to asymptotic approximations. Cost data, though, often exhibit heavy tails and this presents challenges to asymptotic methods. In this paper, we study the performance of these methods for cost effectiveness analysis in medical research and compare them with inference based on the standard bootstrap. We find both methods fail to deliver reliable inference in the presence of heavy tails. We therefore consider two alternative resampling schemes: the wild bootstrap and randomization inference. Both methods provide accurate inference, providing the underlying statistic has finite expected value. In general, these methods provide inference with little size distortion and power above that achievable with asymptotic methods or the standard bootstrap.
This study extends experimental tests of (cumulative) prospect theory (PT) over prospects with more than three outcomes and tests second-order stochastic dominance principles (Levy and Levy, Management Science 48:1334–1349, 2002; Baucells and Heukamp, Management Science 52:1409–1423, 2006). It considers choice behavior of people facing prospects of three different types: gain prospects (losing is not possible), loss prospects (gaining is not possible), and mixed prospects (both gaining and losing are possible). The data supports the distinction of risk behavior into these three categories of prospects, Further, probability weighting and diminishing sensitivity of utility as predicted by PT are observed. Loss aversion is, however, less pronounced, except for choices where one prospect is degenerate. The data suggests that the probability of losing may be relevant for loss aversion.
A number of measures of intertemporal poverty have recently been proposed in the theoretical literature. In this paper, we apply two of these measures to analyse intertemporal poverty in Great Britain during the period 1991-2005, using data from the British Household Panel Survey. Previous studies on poverty using this data-set have employed static measures of poverty. We illustrate how the use of intertemporal poverty measures makes it possible to analyse aspects of poverty which cannot be captured by static, annual, measures of poverty. We then model the determinants of intertemporal poverty, conditional upon being poor, using a Heckman two-step selection model. Keywords: Intertemporal poverty measurement, BHPS, Great Britain JEL Classi cations: D31, I32.
This paper aims to illustrate the potential benefits for academic end-users of integrating existing efforts around describing, building, and using Grid infrastructures. It shows that UK e-Science and e-Social Science projects, among others, can be documented with different levels of user abstraction to facilitate understanding and sharing of the expertise acquired when using Grids and developing Grid-based e-Science and e-Social Science applications. The research study presented uses three existing service-oriented approaches to test the viability of capturing and abstracting the Grid services used in a particular project, and thereby, going beyond document-centric approaches. Each of the three approaches exhibits different levels of abstraction and formalisation and is illustrated by a Grid-based application from the Social Sciences. This example is used to underpin the proposal that it is time to move towards the creation of a collective Knowledge Base that goes beyond presenting projects solely in the scientific literature.
One of the difficulties in processing the results of an empirical quantitative data modeling process is in post-analysis filtering or presenting the information produced. Things are a little easier if the data has a geospatial dimension: one can then use GIS software, specialized statistical modeling software, or produce "mash-ups" for web-based tools. These are all heavyweight solutions. This article presents a lightweight solution which is useful for Grid based or standalone implementation of a problem. Introduction Statistical applications that use social science microdata, whether cross-sectional or longitudinal (panel) studies, suffer from a lack of easy-to-use presentational and exploratory tools. Economic applications, for example, tend to result in a myriad of results tables and/or basic graphics. If such applications are Grid based, in the sense that data hosting and computation are orchestrated via the appropriate middleware, then finishing off the workflow by producing "difficult to digest" hardcopy seems to be a step backwards. However, if the application has a spatial dimension then a natural approach is to use a choropleth map style presentational and exploratory display. Tools of this type are common in the GIS (Geographical Information Systems) domain, but not necessarily easy, or cheap, to deploy. This article presents a lightweight version of one of these tools, which was originally developed for an e-Social Science pilot demonstrator that investigated UK ethnic minority welfare. The original project, entitled Grid Enabled Microeconometric Data Analysis (Peters et al., 2006,2007), Grid enabled two different microdata sources, the British Household Panel Survey (BHPS) and the 1991 Census Sample of Anonymised Records (SARs), and performed calculation of poverty measures and associated statistics using a high performance computing node . The nature of the UK Census at the time meant that the microdata's geography included the UK, its regions, and an artificial local authority area (a SARs area). The project's visualization tool allowed display and investigation of the application's results, by the geography and category of interest (ethnic minority and gender in this case), using a map interface with linked graphics or tables. This used the GeoTools open source GIS Java library (Codehaus, 2006) and was deployed via the project's web interface. Work on the tool has moved on since the aforementioned project's conclusion to address other issues including: 1) data display, 2) tool deployment, and 3) other microdata applications. For issue 1), the original map colour scheme was improved, and the issue of producing output for formal publication was addressed. The nature of the application also suggested enhancements to the visualization process to allow results filtering by sample size and/or statistical precision (p-value for a hypothesis test, standard error for an estimate). Sample size filtering is particularity pertinent as it alludes to the types of data disclosure restrictions imposed upon microdata by governmental and other data owners, a topic which leads into the issue of tool deployment. The original tool was deployed as part of an eResearch project that required authorised and authenticated access to the application's data sources. The tool still retains this characteristic, however, to repeat the application using 2001 UK data (the present Census currency) is not feasible as the data are only available in a secure data enclave. The code, therefore, needs taking to the remote location. The present version can be used standalone and is opensource. The latter point is important, as data owners are unwilling to deploy unvetted black box code on their secure servers. The third and final issue is the ability to use the tool for different applications. This requires both mapping information and microdata, the only requirement being that these are in a format suitable for the GeoTools library. For UK microdata applications that use the Census geographies, the mapping information is available from the Edina UK borders project under Athens authentication. This short paper discusses the present state of the tool, the details associated with the extensions discussed above and developments in progress. Deployment and Implementation Issues The core of the tool is based upon the GeoTools open source Java library that provides methods for the manipulation and viewing of geospatial data. As it is a Java library, associated applets or applications are platform independent. Users do not need to install it on their computers as the necessary parts are downloaded as jar files when the applet is loaded. The present implementation is fixed on GeoTools 2.0 to avoid re-writing the user interface every time GeoTools is updated. The tool itself can be used both as a Java applet embedded in a Grid based application or as a standalone Java application, the latter being particularly useful for special environments such as secure data enclaves. To use it in a Grid based setting the user must have appropriate authentication and authorisation for the service, such as an e-certificate for the NGS (the UK's National Grid Service) or EGEE (The EU's Enabling Grids for E-SciencE). It is designed to sit on top of a statistical analysis that has been deployed on a Grid, and is dissociated from the middleware. This isolates development of the tool from issues (fashion, sustainability) related to the evolution of the Grid. A hypothetical implementation using the P-GRADE Portal (Sipos and Kacsuk, 2006), which can replace certain of the bespoke elements of Peters et al (2006, 2007), is presented in Figure 1 below. Both the Grid and standalone usages may require extra security to permit use of the mapping data. This will depend upon the user's mapping file provider. For example, a user has to agree to the UK Borders license agreement for the embedded maps required for working with UK Census microdata. In practice this requires a user to have an Athens username and password. These maps are downloaded in shape format, the supported format used in the GeoTools library. The correct geographical levels (regional maps and SARs area maps) for the currency of the data sources are readily available, however, some modifications are needed to produce a UK wide map as England, Scotland, and Wales are obtained as separate distinct mapping files. These have to be combined. At the regional level, distinguishing the regions 'South East', 'Outer London', and 'Inner London' also required manipulation of the shape data. Once the final mapping file is available then it can be linked with an appropriately formatted data file. Most data analyses will produce flat result files which can be visualized after some filtering. This requires an application to convert them into the dbf structure required by the shape file. The combination of shape file and dbf file then become the input for the visualisation applet. Figure 1 The Tool as an Add-on to General Middleware User Interface Features In certain social sciences, such as the economics discipline area, viewing the results of any microdata based modelling process in anything other than a table is still a relatively new experience. If the results of the analyses are related to specific geographies then they can be viewed using the aforementioned visualisation tool. Basic interaction (zooming, panning, etc.) within the map is possible, as well as viewing relevant statistical plots linked to a specific area of a map. The user can choose between categories pertinent to the analysis and the geographical level of the map. Our example deals with poverty measures that are produced for an ethnic group and gender category at UK regional or local area (a SARs area) geography. Feedback from a series of peer conferences and workshops about the original visualization applet produced a number of criticisms and constructive proposals for extensions. The three main criticisms concerned: a) the chosen colour map; it was not suitable for the colour blind, b) print journal cost effectiveness; academic articles are still published in journals using greyscale by default1 and c) extra functionality. 1 Publishers will produce colour plates in academic journals by request. The cost for this, however, is somewhat high. P-GRADE Portal 1) stage data extraction job 2) stage compute job 3) stage results processing job HPC computing nodes, on either the NGS or EGEE. Data Host Files for the tool. Session Initiation Visualization Tool Map data
This article discusses the use of Grid technology to integrate the data, computation and presentation elements of an empirical economic modelling process. We achieve this by using a form of statistical data fusion developed in the poverty mapping literature to address a substantive issue: determining United Kingdom ethnic minority welfare. Elements of this methodology appear to be well suited to such grid-enablement, and we present and illustrate our implementation using the context of this microdata application.
This article presents and discusses the motivation, methodology and implementation of an e-Social Science pilot demonstrator project entitled: Grid Enabled Micro-econometric Data Analysis (GEMEDA). This used the National Grid Service (NGS) to investigate a policy relevant Social Science issue: the welfare of ethnic minority groups in the United Kingdom. The underlying problem is that of a statistical analysis that uses quantitative data from more than one source. The application of grid technology to this problem allows one to integrate elements of the required empirical modelling process: data extraction, data transfer, statistical computation and results presentation, in a manner that is transparent to a casual user.
The User Requirements and Web Based Access for eResearch Workshop, organized jointly by NeSC and NCeSS, was held on 19 May 2006. The aim was to identify lessons learned from e-Science projects that would contribute to our capacity to make Grid infrastructures and tools usable and accessible for diverse user communities. Its focus was on providing an opportunity for a pragmatic discussion between e-Science end users and tool builders in order to understand usability challenges, technological options, community-specific content and needs, and methodologies for design and development. We invited members of six UK e-Science projects and one US project, trying as far as possible to pair a user and developer from each project in order to discuss their contrasting perspectives and experiences. Three breakout group sessions covered the topics of user-developer relations, commodification, and functionality. There was also extensive post-meeting discussion, summarized here. Additional information on the workshop, including the agenda, participant list, and talk slides, can be found online at http://www.nesc.ac.uk/esi/events/685/ Reference: NeSC report UKeS-2006-07 available from http://www.nesc.ac.uk/technical_papers/UKeS-2006-07.pdf
Using a sample of male and female workers from the 1992 Employment in Britain survey, we estimate a generalised grouped zero-inflated Poisson regression model of employees' self-reported lateness. Lateness is higher for males, private sector workers and in service industries. Reflecting theoretical predictions from both psychology and economics, we model lateness as a function of incentives, the monitoring of, and sanctions for, lateness within the workplace, job satisfaction and attitudes to work. Various aspects of workplace incentive and disciplinary policies turn out to affect lateness; however, controlling for these, an important role for job satisfaction remains.
The classical trinity of tests is used to check for the presence of atremble in economic experiments in which the response variable is binary. A tremble is said to occur when an agent makes a decision completely at random, without regard to the values taken by the explanatory variables. The properties of the tests are discussed, and an extension of the methodology is used to test for the presence of a tremble in binary panel data from a well-known economic experiment.
This article notes that it is now practical to use the method of enumerationto analyse the performance of estimators and hypothesis tests of fullyparametric binary data models. The general method is presented and thenemployed to investigate the power performance of a common misspecificationtest for the Probit model. The advantages, disadvantages and limitations ofenumeration compared with standard Monte Carlo simulation are thendiscussed. Finally, an example from experimental economics is used todemonstrate that the methodology can also be used in small empirical studies.
. Various count data models are applied to data collected from a sample of Norfolk young persons, who were asked how many times they had had sexual intercourse during the previous two-week period. The models take account of the fact that the data are “grouped”, meaning that for some observations, the count is not known exactly but is known to fall in a particular range. Using a formal testing procedure, we find overwhelming evidence of the presence of excess zeros, and this we attribute to the fact that, at any time, a certain proportion of the population are sexually inactive. Our final model contains two equations, the first being the participation equation which determines whether an individual is sexually active, and the second being the frequency equation which determines the count, conditional on being active. Age, gender, salary, occupational status, marital status and type of living environment all have interesting effects on either participation or frequency. Since a significant proportion of the original sample declined to reveal coital frequency, we address the potential problem of selection bias by including a Heckman-type correction term during the model selection process.
When values of regressors are symmetrically disposed, many M-estimators in a wide class of models have a reflection property, namely, that as the signs of the coefficients on regressors are reversed, their estimators' sampling distribution is reflected about the origin. When the coefficients are zero, sign reversal can have no effect. So in this case, the sampling distribution of regression coefficient estimators is symmetric about zero, the estimators are median unbiased and, when moments exist, the estimators are exactly uncorrelated with estimators of other parameters. The result is unusual in that it does not require response variates to have symmetric conditional distributions. It demonstrates the potential importance of covariate design in determining the distributions of estimators, and it is useful in designing and interpreting Monte Carlo experiments. The result is illustrated by a Monte Carlo experiment in which maximum likelihood and symmetrically censored least-squares estimators are calculated for small samples from a censored normal linear regression, Tobit, model.
Score or Lagrange multiplier versions of a Hausman test for distributional specification are presented for limited dependent variable models. These tests are based on the first derivatives of semiparametric criterion functions associated with robust estimation of such models and only require the computation of the maximum likelihood estimator. Various regression forms of the statistic are also presented. Monte Carlo results for the Tobit model indicate that such tests may be efficacious in detecting situations in which the maximum likelihood estimator is seriously biased due to incorrect distributional specification.