Occupational Health Informatics is the field concerned with the optimal use of information to maximize occupational health and well-being, achieved by multidisciplinary professionals using appropriate technologies. The occupational health community should advocate for the recording of occupation in electronic health record systems. Data linkages can be created between occupation, disease, function and exposures to inform occupational functionality, recognition and recording of occupational diseases, disability accommodation and work as a health outcome.
Objectives Automatic job coding tools were developed to reduce the laborious task of manually assigning job codes based on free-text job descriptions in census and survey data sources, including large occupational health studies. The objective of this study is to provide a case study of comparative performance of job coding and JEM (Job-Exposure Matrix)-assigned exposures agreement using existing coding tools. Methods We compared three automatic job coding tools [AUTONOC, CASCOT (Computer-Assisted Structured Coding Tool), and LabourR], which were selected based on availability, coding of English free-text into coding systems closely related to the 1988 version of the International Standard Classification of Occupations (ISCO-88), and capability to perform batch coding. We used manually coded job histories from the AsiaLymph case-control study that were translated into English prior to auto-coding to assess their performance. We applied two general population JEMs to assess agreement at exposure level. Percent agreement and PABAK (Prevalence-Adjusted Bias-Adjusted Kappa) were used to compare the agreement of results from manual coders and automatic coding tools. Results The coding per cent agreement among the three tools ranged from 17.7 to 26.0% for exact matches at the most detailed 4-digit ISCO-88 level. The agreement was better at a more general level of job coding (e.g. 43.8-58.1% in 1-digit ISCO-88), and in exposure assignments (median values of PABAK coefficient ranging 0.69-0.78 across 12 JEM-assigned exposures). Based on our testing data, CASCOT was found to outperform others in terms of better agreement in both job coding (26% 4-digit agreement) and exposure assignment (median kappa 0.61). Conclusions In this study, we observed that agreement on job coding was generally low for the three tools but noted a higher degree of agreement in assigned exposures. The results indicate the need for study-specific evaluations prior to their automatic use in general population studies, as well as improvements in the evaluated automatic coding tools.
Introduction Occupational data in prospective cohort studies is often underutilized due to the human and financial resources required to code open-ended text, such as job titles. Recognizing the value of occupational data in health research, as well as potential errors associated with manual coding, an Automated Coding Algorithm (ACA)-NOC algorithm was developed utilizing a Natural Language Processing approach. Objectives We tested the ACA-NOC algorithm on two regional cohorts of a pan-Canadian cohort study, which represents the largest dataset an algorithm of this kind has been applied to. This process will harmonize and greatly expand the utility of the occupational data, enrich the research platforms, and further refine the efficiency of the algorithm. Methods The ACA-NOC algorithm was tested on data from the Canadian Partnership for Tomorrow’s Health (CanPath), a longitudinal cohort examining the role of genetic, environmental, lifestyle, and behavioural factors in the development of cancer and chronic disease. Using an iterative and interactive approach, the algorithm was applied to job title data from 111,000 questionnaires from two regional cohorts, coding the data to the Canadian National Occupation Classification (NOC) system. The algorithm was further refined based on each round of analysis, increasing the quantity of accurately coded data. Results Results from this research demonstrate the ability to refine the ACA-NOC algorithm with a 10% overall improvement in exact matching from the baseline algorithm. There were also instances where the algorithm performance was superior to the manual coding. The utilization of the algorithm offers significant savings in time, human resources and cost compared to a singular manual coding approach. Conclusions The coding and harmonization of this multi-cohort data demonstrates the value of the ACA-NOC algorithm, while increasing the utility of the CanPath data and research related to occupational health. Future research may involve comparisons between CanPath and international cohorts.
IntroductionOngoing studies into the use of algorithms for the automated coding of job titles to the Canadian National Occupation Classification have performance accuracy which are at least equivalent to manual coding accuracy. Moreover automated coding provides significant time savings. These studies have identified that both natural language processing and machine learning algorithms are effective for auto coding. Whereas NLP based and machine learning approaches both rely on bespoke rules, and existing data sets, machine learning models can proliferate bias from training data if not corrected.ObjectivesThe goal of the study is to explore the impact of altering sex/gender ratios in training data sets on overall performance of the machine learning based prediction of NOC codes using patient provided job titles.MethodsUsing data participant patient data provided by Atlantic PATH, training data sets were prepared for 100 4-digit NOC categories. The data sets were prepared with sex/gender ratios of 50/50 30/70, 70/30. The data sets were used to train ENENOC machine learning platform and tested on a set of manually coded job titles provided by Atlantic PATH CanPATH . Performance levels were contrasted for all 4-digit NOC categories used in the study.ResultsInitial results in this preliminary study have identified that sex and gender are variables that can influence auto coding performance, however the extent to which overall coding accuracy is impacted is relative minor. Further studies are required with larger training sets to fully explore the extent of sex and gender as contributing variables to bias to ENENOC.ConclusionWe initiated studies to investigate the impact of sex and gender bias on performance of the ENENOC algorithm. Together, the ENENOC contributed training and test sets provide a suitable framework for ongoing work in this area.
Many research studies seek to identify the social determinants of health and occupation is an important predictor, both at the level of the individual as well as for populations. Whereas job titles are usually solicited during interviews or by questionnaire, before being able to use this information the responses need to be categorized using a coding system, such as the Canadian National Occupational Classification (NOC). Manual coding is the usual method, which is a time-consuming and error-prone activity with variable or inconsistent outcomes from teams of coders. In recent work the ACA-NOC algorithm1 was developed to perform automated coding based on matching job title text with the NOC’s job titles and textual descriptions. This algorithm was benchmarked on a small sample manually coded data set with subject matter experts subsequent review of coding discrepancies to facilitate functional improvements to the algorithm. Performance levels achieved illustrated the viability of the approach albeit larger benchmarking data sets were required. CanPATH2has collected data from approximately 330,000 volunteer Canadians, including information about health, lifestyle, occupation, environment and behavior. We report on the further benchmarking and further development of this algorithm in CanPATH funded project using over 60,000 manually coded job titles from the constituent Alberta Tomorrow Project. The algorithm was also applied to over 100,000 un-coded job titles from Atlantic PATH, including the Core questionnaire and occupational history data. The core outcome of the project identified that auto-coding results are comparable to manual coding in accuracy and superior in speed e.g. 2 years of manual coding (64,000 records) can be auto coded in 72 hours. The algorithm was considered ready for deployment in operational settings: point of care, decision support for manual coders. Additional insights gained during the project revealed that (i) NOC and ATP data sets have a distribution bias where some NOC categories were over or under-represented and numerous non-standard lexical features were found in job titles and NOC job descriptions, (ii) benchmarking datasets from ATP included coding errors that were corrected by expert coders leading to the creation of gold standard test sets for further algorithm improvement studies, (iii) a study on 17 categories of occupations initially difficult to code, identified some job categories with near 90% coding accuracy. Automated coding of job titles to the NOC has been shown to be both practicable to good levels of accuracy and shown to significantly accelerate manual coding efforts from years to autocoding in a matter of hours without decrease in accuracy. Autocoding can replace costly, error prone manual labor with accurate point-of-care auto-coding such that patient occupation information during healthcare encounters could now supplement existing administrative data sets in electronic health record systems. This data can be used better to understand the socioeconomic consequences of health conditions, advise patients about returning to work with a health condition, recognizing occupations at risk of disease e.g. as in the COVID-19 pandemic.ReferencesBao H, Baker CJO, Adisesh A. Occupation coding of job titles: iterative development of an automated coding algorithm for the canadian national occupation classification (ACA-NOC). JMIR Form Res 2020 Aug 5;4(8):e16422. doi:10.2196/16422 CanPath the Canadian Partnership for Tomorrow Project. https://canpath.ca/ accessed: 01.09.2021
IntroductionInformation about occupations and job attributes is available in siloed databases. The lack of integrated data precludes ad-hoc querying and research investigating occupational determinants of health e.g. COVID-19 or stress. Core to integration of occupation data is taxonomic representation of job categories. In North America the official occupational taxonomies are the Canadian National Occupational Classification (NOC) and the United States Standard Occupational Classification (SOC).ObjectivesThis study aimed to integrate job attribute data from the Canadian Career Handbook (CH) and O*NET database to facilitate cross-classification query capabilities and to prototype the creation of metrics for comparing occupations based on job attributes.MethodsThe integrated database was completedhierarchical structures of both occupational taxonomies were represented; job attributes were selected from the CH and O*NET-SOC; the database was populated with occupational descriptions; occupational codes from the CH and O*NET-SOC were linked using the Brookfield Institute NOC to O*NET-SOC crosswalk.ResultsThe database consists of 1679 rows with unique occupations and 181 columns with occupational attributes. Rows contain a unique combination of hierarchical structures from the CH and O*NET-SOC. Rows also contain detailed occupational descriptions from CH and O*NET-SOC. We queried the integrated data checking O*NET-SOC to CH equivalence and cross-taxonomy selection of job attributes, e.g. Retrieve all or selected attributes for an occupation by CH code or equivalent code in O*NET-SOC. We ran queries for targeted scenarios to retrieve occupations: i) where work is done in physical proximity to others, ii) where incumbents are exposed to disease or infections, iii) at risk of back pain due to physical work factors, iv) where incumbents experience high work-related stressors.ConclusionWe report a database combining selected information from the CH and O*NET-SOC that facilitates complex occupational health queries. Further we investigated work-related stressors on low back pain risk by occupation.
Scientific data analyses often combine several computational tools in automated pipelines, or workflows. Thousands of such workflows have been used in the life sciences, though their composition has remained a cumbersome manual process due to a lack of standards for annotation, assembly, and implementation. Recent technological advances have returned the long-standing vision of automated workflow composition into focus. This article summarizes a recent Lorentz Center workshop dedicated to automated composition of workflows in the life sciences. We survey previous initiatives to automate the composition process, and discuss the current state of the art and future perspectives. We start by drawing the “big picture” of the scientific workflow development life cycle, before surveying and discussing current methods, technologies and practices for semantic domain modelling, automation in workflow development, and workflow assessment. Finally, we derive a roadmap of individual and community-based actions to work toward the vision of automated workflow development in the forthcoming years. A central outcome of the workshop is a general description of the workflow life cycle in six stages: 1) scientific question or hypothesis, 2) conceptual workflow, 3) abstract workflow, 4) concrete workflow, 5) production workflow, and 6) scientific results. The transitions between stages are facilitated by diverse tools and methods, usually incorporating domain knowledge in some form. Formal semantic domain modelling is hard and often a bottleneck for the application of semantic technologies. However, life science communities have made considerable progress here in recent years and are continuously improving, renewing interest in the application of semantic technologies for workflow exploration, composition and instantiation. Combined with systematic benchmarking with reference data and large-scale deployment of production-stage workflows, such technologies enable a more systematic process of workflow development than we know today. We believe that this can lead to more robust, reusable, and sustainable workflows in the future.
IntroductionOccupational encoding is a technique that allows job titles provided by study participants to be categorized according to their role in the labor force. Encoding has primarily been a slow error-prone manual process which is ripe for automation.ObjectivesOur goals was to design and test an automated coding prototype using machine learning techniques.MethodsThe prototype classification system ENENOC (the ENsemble Encoder for the National Occupational Classification) is comprised of series of steps involving data cleaning, exact match search, multi classifier ensembling, hierarchical classification, and multiple output selection. In the absence of exact matching between job title input and NOC category descriptions, the input data is embedded using the TF-IDF algorithm and Doc2Vec. The embeddings are fed into a hierarchical, ensemble classifier that uses classical machine learning techniques: Random Forests, Support Vector Machine and K-Nearest Neighbour. Ensemble encoding is achieved using a majority-voting system. The hierarchical two tier classification methodology first predicts the first digit of the NOC code followed while the second tier predicts the second third and fourth digit of the NOC code for the input data. The combined approach produces a single, 4-digit code as a top choice, as well as four alternate NOC codes, that serve as additional ranked choice based on the Doc2Vec model.ResultsThe prototype was benchmarked on a manually annotated data set comprising of 64,000 records. It produced a top-1 Per-Digit Macro F1-Score of 0.65 and a top-5 Per-Digit Macro F1-Score of 0.76, both of which are highly within published accuracy ranges for manual coding (44% to 89% inter-annotator agreement). ENENOC coded 30,000 job titles in 3 hours.ConclusionThe ENENOC prototype is a sophisticated ENsemble Encoder for the National Occupational Classification which has state of the art performance accuracy with significant speed improvements over manual coding.
BACKGROUND:In many research studies, the identification of social determinants is an important activity, in particular, information about occupations is frequently added to existing patient data. Such information is usually solicited during interviews with open-ended questions such as "What is your job?" and "What industry sector do you work in?" Before being able to use this information for further analysis, the responses need to be categorized using a coding system, such as the Canadian National Occupational Classification (NOC). Manual coding is the usual method, which is a time-consuming and error-prone activity, suitable for automation.OBJECTIVE:This study aims to facilitate automated coding by introducing a rigorous algorithm that will be able to identify the NOC (2016) codes using only job title and industry information as input. Using manually coded data sets, we sought to benchmark and iteratively improve the performance of the algorithm.METHODS:We developed the ACA-NOC algorithm based on the NOC (2016), which allowed users to match NOC codes with job and industry titles. We employed several different search strategies in the ACA-NOC algorithm to find the best match, including exact search, minor exact search, like search, near (same order) search, near (different order) search, any search, and weak match search. In addition, a filtering step based on the hierarchical structure of the NOC data was applied to the algorithm to select the best matching codes.RESULTS:The ACA-NOC was applied to over 500 manually coded job and industry titles. The accuracy rate at the four-digit NOC code level was 58.7% (332/566) and improved when broader job categories were considered (65.0% at the three-digit NOC code level, 72.3% at the two-digit NOC code level, and 81.6% at the one-digit NOC code level).CONCLUSIONS:The ACA-NOC is a rigorous algorithm for automatically coding the Canadian NOC system and has been evaluated using real-world data. It allows researchers to code moderate-sized data sets with occupation in a timely and cost-efficient manner such that further analytics are possible. Initial assessments indicate that it has state-of-the-art performance and is readily extensible upon further benchmarking on larger data sets.
This paper reports on the early-stage development of an analytics framework to support the semantic integration of dynamic surveillance data across multiple scales to inform decision making for malaria eradication. We propose using the Semantic Web of Things (SWoT), a combination of Internet of Things (IoT) and semantic web technologies, to support the evolution and integration of dynamic malaria data sources and improve interoperability between different datasets generated through relevant IoT assets (e.g. computers, sensors, persons, and other smart objects and devices).
Global health surveillance and pandemic intelligence rely on the systematic collection and integration of data from diverse distributed and heterogeneous sources at various levels of granularity. These sources include data from multiple disciplines represented in different formats, languages, and structures posing significant integration challenges This article provides an overview of challenges in data driven surveillance. Using Malaria surveillance as a use case we highlight the contribution made by emerging semantic data federation technologies that offer enhanced interoperability, interpretability and explainability through the adoption of ontologies. The paper concludes with a focus on the relevance of these technologies for ongoing pandemic preparedness initiatives.
. PSOA RuleML is a rule language which introduces positional-slotted, object-applicative terms in generalized rules, permitting relation applications with optional object identifiers and positional or slotted arguments. This paper describes an open-source PSOA RuleML API, whose functionality facilitates factory-based syntactic object creation and manipulation. The API parses an XML-based concrete syntax of PSOA RuleML, creates abstract syntax objects, and uses these objects for translation into a RIF-like presentation syntax. The availability of such an API will benefit PSOA rule-based research and applications.
Mouse embryonic stem cells (mESCs) cultured in the presence of LIF occupy a ground state with highly active pluripotency-associated transcriptional and epigenetic circuitry. However, ground state pluripotency in some inbred strain backgrounds is unstable in the absence of ERK1/2 and GSK3 inhibition. Using an unbiased genetic approach, we dissect the basis of this divergent response to extracellular cues by profiling gene expression and chromatin accessibility in 170 genetically heterogeneous mESCs. We map thousands of loci affecting chromatin accessibility and/or transcript abundance, including 10 QTL hotspots where genetic variation at a single locus coordinates the regulation of genes throughout the genome. For one hotspot, we identify a single enhancer variant ∼10 kb upstream of Lifr associated with chromatin accessibility and mediating a cascade of molecular events affecting pluripotency. We validate causation through reciprocal allele swaps, demonstrating the functional consequences of noncoding variation in gene regulatory networks that stabilize pluripotent states in vitro.
Occupation is an explanatory variable in health research that is used to identify the degree to which exposures to environmental hazards and working conditions are correlated with disease. Moreover disease and functional impairment can limit employment options open to patients. Despite the importance of these issues many essential data sets have yet to be integrated. In the current study we defined an integrated semantic model and populated coded patient data representing disease (ICD), functional impairment (ICF), occupation (NOC), and job attributes (NOC Career Handbook). Automated NOC coding of patient responses to “What is your job” were coded by a custom algorithm developed in previous work. To validate the utility of the model, SPARQL queries and outputs were prepared and discussed in the context of authentic physician and case worker activities.
Mouse embryonic stem cells (mESCs) cultured under controlled conditions occupy a stable ground state where pluripotency-associated transcriptional and epigenetic circuitry are highly active. However, mESCs from some genetic backgrounds exhibit metastability, where ground state pluripotency is lost in the absence of ERK1/2 and GSK3 inhibition. We dissected the genetic basis of metastability by profiling gene expression and chromatin accessibility in 185 genetically heterogeneous mESCs. We mapped thousands of loci affecting chromatin accessibility and/or transcript abundance, including eleven instances where distant QTL co-localized in clusters. For one cluster we identified Lifr transcript abundance as the causal intermediate regulating 122 distant genes enriched for roles in maintenance of pluripotency. Joint mediation analysis implicated a single enhancer variant ~10kb upstream of Lifr that alters chromatin accessibility and precipitates a cascade of molecular events affecting maintenance of pluripotency. We validated this hypothesis using reciprocal allele swaps, revealing mechanistic details underlying variability in ground state metastability in mESCs.
Threatened freshwater ecosystems urgently require improved tools for effective management. Food web analysis is currently under-utilized, yet can be used to generate metrics to support biomonitoring assessments by measuring the stability and robustness of ecosystems. Using a previously developed analysis pipeline, we combined taxonomic outputs from DNA metabarcoding with a text-mining routine to extract trait information directly from the literature. This pipeline allowed us to generate heuristic food webs for sites within the lower Saint John/Wolastoq River and the Grand Lake Meadows (hereafter called the “GLM complex”), Atlantic Canada's largest freshwater wetland. While these food webs are derived from empirical traits and their structure has been shown to discriminate sites both spatially and temporally, the accuracy of their properties have not been assessed against other methods of trophic analysis. We explored two approaches to validate the utility of heuristic food webs. First, we qualitatively compared how well-trophic position derived from heuristic food webs recovered spatial and temporal differences across the GLM complex in comparison to traditional stable isotope approaches. Second, we explored how the trophic position of invertebrates, derived from heuristic food webs, predicted trophic position measured from δ 15N values. In general, both heuristic food webs and stable isotopes were able to detect seasonal changes in maximum trophic position in the GLM complex. Samples from the entire GLM complex demonstrated that prey-averaged trophic position measured from heuristic food webs strongly predicted trophic position inferred from stable isotopes (R 2 = 0.60), and even stronger relationships were observed for some individual models (R 2 = 0.78 for best model). Beyond their areas of congruence, heuristic food web and stable isotope analyses also appear to complement one another, suggesting a surprising degree of independence between community trophic niche width (assessed from stable isotopes) and food web size and complexity (assessed from heuristic food webs). Collectively, these analyses indicate that trait-based networks have properties that correspond to those of actual food webs, supporting the routine adoption of food web metrics for ecosystem biomonitoring.
Ecological networks are powerful tools for visualizing biodiversity data and assessing ecosystem health and function. Constructing these networks requires considerable empirical efforts, and this remains highly challenging due to sampling limitations and the laborious and notoriously limited, error-prone process of traditional taxonomic identification. Recent advancements in high-throughput gene sequencing and high-performance computing provide new ways to address these challenges. DNA metabarcoding, a method of bulk taxonomic identification from DNA extracted from environmental samples, can generate detailed biodiversity information through a standardizable analytical pipeline for species detection. When this biodiversity information is annotated with prior knowledge on taxon interactions, body size, and trophic position, it is possible to generate trait-based networks, which we call "heuristic food webs". Although curating trait matrices for constructing heuristic food webs is a laborious, often intractable process using manual literature surveys, it can be greatly accelerated via text mining, allowing knowledge of relevant traits to be gathered across large databases. To explore this possibility, we employed a General Architecture for Text Engineering (GATE) system to create a hybrid text-mining pipeline combining rule-based and machine-learning modules. This pipeline was then used to query online repositories of published papers for missing data on a key trait, body size, that could not be gathered from existing trophic link libraries of freshwater benthic macroinvertebrates. Combining text-mined body size information with feeding information from existing sources allowed us to generate a database of over 20,000 pairwise trophic interactions. Next, we developed a pipeline that uses taxa lists generated from DNA metabarcoding and annotates this matrix with trophic information from existing databases and text-mined body size data. In this way, we generated heuristic food webs for wetland sites within a large delta complex formed by the confluence of the Peace and Athabasca rivers in northern Alberta: the Peace Athabasca delta. Finally, we used these putative food webs and their network properties to resolve spatial and temporal differences between the benthic subwebs of wetlands in the Peace and Athabasca sectors of the delta complex. Specifically, we asked two questions. (1) How do food web properties (e.g. number of links, linkage density, trophic height) differ between the wetlands of the Peace and Athabasca deltas? (2) How do food web properties change temporally in wetlands of the two deltas? We discuss using DNA-generated, trait-based food webs as a powerful tool for rapid bioassessment, assess the limitations of our current approach, and outline a path forward to make this powerful tool more widely available for land managers and conservation biologists.
Background According to the World Health Organization, malaria surveillance is weakest in countries and regions with the highest malaria burden. A core obstacle is that the data required to perform malaria surveillance are fragmented in multiple data silos distributed across geographic regions. Furthermore, consistent integrated malaria data sources are few, and a low degree of interoperability exists between them. As a result, it is difficult to identify disease trends and to plan for effective interventions. Objective We propose the Semantics, Interoperability, and Evolution for Malaria Analytics (SIEMA) platform for use in malaria surveillance based on semantic data federation. Using this approach, it is possible to access distributed data, extend and preserve interoperability between multiple dynamic distributed malaria sources, and facilitate detection of system changes that can interrupt mission-critical global surveillance activities. Methods We used Semantic Automated Discovery and Integration (SADI) Semantic Web Services to enable data access and improve interoperability, and the graphical user interface-enabled semantic query engine HYDRA to implement the target queries typical of malaria programs. We implemented a custom algorithm to detect changes to community-developed terminologies, data sources, and services that are core to SIEMA. This algorithm reports to a dashboard. Valet SADI is used to mitigate the impact of changes by rebuilding affected services. Results We developed a prototype surveillance and change management platform from a combination of third-party tools, community-developed terminologies, and custom algorithms. We illustrated a methodology and core infrastructure to facilitate interoperable access to distributed data sources using SADI Semantic Web services. This degree of access makes it possible to implement complex queries needed by our user community with minimal technical skill. We implemented a dashboard that reports on terminology changes that can render the services inactive, jeopardizing system interoperability. Using this information, end users can control and reactively rebuild services to preserve interoperability and minimize service downtime. Conclusions We introduce a framework suitable for use in malaria surveillance that supports the creation of flexible surveillance queries across distributed data resources. The platform provides interoperable access to target data sources, is domain agnostic, and with updates to core terminological resources is readily transferable to other surveillance activities. A dashboard enables users to review changes to the infrastructure and invoke system updates. The platform significantly extends the range of functionalities offered by malaria information systems, beyond the state-of-the-art.
Informational needs of agricultural consultants are increasingly complex. Advising farmers on the appropriate measures for optimizing cropping yields demands access to custom data archives and analytics tools. In line with the increasing number of archives, the expertise required of consultants goes beyond the capabilities of these non-technical agri-specialists. These end users have diverse ad-hoc query needs and require tools that provide simple access to distributed data silos and easy ways to integrate relevant information. In this article, the authors report on a pilot deployment of Semantic Automated Discovery and Integration (SADI) Web services for the federation and computation of agricultural data. A registry of 9 SADI Web services was deployed to expose data from a variety of different data resources in support of a defined set of query needs. The authors demonstrate that the deployment of these services facilitates the ad-hoc creation and execution of mission critical workflows targeting use cases in agricultural operations management. Using HYDRA, a semantic query engine for SADI Web services with a custom built graphical user interface, agricultural consultants can identify optimal crop varieties, and compute profit margins of each variety using a complex cost model.
Malaria is an infectious disease affecting people across tropical countries. In order to devise efficient interventions, surveillance experts need to be able to answer increasingly complex queries integrating information coming from repositories distributed all over the globe. This, in turn, requires extraordinary coding abilities that cannot be expected from non-technical surveillance experts. In this paper, we present a deployment of Semantic Automated Discovery and Integration (SADI) Web services for the federation and querying of malaria data. More than 10 services were created to answer an example query requiring data coming from various sources. Our method assists surveillance experts in formulating their queries and gaining access to the answers they need.
Harold Boley合作论文数Semantic Web Laboratory;Faculty of Computer Science;University of New Brunswick6