Hepatitis E virus (HEV) is the most common cause of acute viral hepatitis worldwide. An increased risk for HEV infection has been reported in organ-transplant recipients, mainly from Europe. Prospective data on HEV prevalence in the United States (U.S.) organ transplant population are limited. To determine the prevalence and factors associated with HEV infection among solid organ transplant-recipients, we conducted a prospective, cross-sectional, multicentre study among transplant-recipients and age- and organ-matched waitlist patients. Participants answered a risk-exposure questionnaire and were tested for HEV-RNA (in-house PCR), HEV-IgG, and IgM (ELISA, Wantai). Among 456 participants, 224 were transplant-recipients, and 232 were waitlist patients. The mean age was 58 years, 35% female, and 74% White. HEV seroprevalence of the entire cohort was 20.2% and associated with older age (p < 0.0001) and organ transplantation (p = 0.02). The HEV seropositivity was significantly higher among transplant-recipients compared with waitlist patients (24% vs. 16.4%, p = 0.042). Among transplant recipients, relative-risk of being HEV seropositive increased with older age (RR = 3.4 [1.07-10.74] in patients >70 years compared with ≤50 years, p = 0.037); history of graft hepatitis (2.2 [1.27-3.72], p = 0.005); calcineurin inhibitor use (RR = 1.9 [1.03-3.34], p = 0.02); and kidney transplantation (2.4 [1.15-5.16], p = 0.02). HEV-RNA, genotype 3 was detected in only two patients (0.4%), both transplant-recipients. HEV seroprevalence was higher among transplant-recipients than waitlist patients. HEV should be considered in transplant-recipients presenting with graft hepatitis. Detection of HEV-RNA was rare, suggesting that progression to chronic HEV infection is uncommon in transplant-recipients in the U.S.
Awareness of data and information quality issues has grown rapidly in light of the critical role played by the quality of information in our data-intensive, knowledge-based economy. Research in the past two decades has produced a large body of data quality knowledge and has expanded our ability to solve many data and information quality problems. We present an overview of the evolution and current landscape of data and information quality research. We introduce an integrative framework to characterize the research along two dimensions: topics and methods. Representative papers are cited for purposes of illustrating the issues addressed and the methods used. We also identify and discuss challenges to be addressed in future research.
Data curation is the process of acquiring multiple sources of data, assessing and improving data quality, standardizing, and integrating the data into a usable information product, and eventually disposing of the data. The research describes the building of a proof-of-concept for an unsupervised data curation process addressing a basic form of data cleansing in the form of identifying redundant records through entity resolution and spelling corrections. The novelty of the approach is to use ER as the first step using an unsupervised blocking and stop word scheme based on token frequency. A scoring matrix is used for linking unstandardized references, and an unsupervised process for evaluating linking results based on cluster entropy. The ER process is iterative, and in each iteration, the match threshold is increased. The prototype was tested on 18 fully-annotated test samples of primarily synthetic person data varied in two different ways, good data quality versus poor data quality, and a single record layout versus two different record layouts. In samples with good data quality and using both single and mixed layouts, the final clusters had an average F-measure of 0.91, precision of 0.96, and recall of 0.87 outcomes comparable to results from a supervised ER process. In samples with poor data quality whether mixed or single layout, the average F-measure was 0.78, precision 0.74, and recall 0.83 showing that data quality assessment and improvement is still a critical component of successful data curation. The results demonstrate the feasibility of building an unsupervised ER engine to support data integration for good quality references while avoiding the time and effort to standardize reference sources to a common layout, design, and test matching rules, design blocking keys, or test blocking alignment. Also, the paper proposes how unsupervised data quality improvement processes could also be incorporated into the design allowing the model to address an even broader range of data curation applications.
Tension between value creation and value appropriation can arise when firms enter relationships with a platform. While creating value collectively, these relationships strengthen the network effects which increase the platform’s ability to appropriate value. This shift in bargaining position could restrain cooperation between the firms and the platform, thereby diminishing joint value creation. We analyze e-book data to study product offerings as a strategic mechanism that publishers use when managing relationships with Amazon Kindle. We find that publishers make product offering choices that increase value creation while alleviate appropriation risks. Compared to small publishers, large publishers make product decisions that are more protective of their own interests and less conducive to Kindle’s success. We discuss our findings in relation to theory and future research.
Renewable energy systems need to be able to make frequent and rapid adjustments to address shifting solar and wind production. This requires increasingly sophisticated industrial control systems (ICS). But, that also increases the potential risks from cyber-attacks. Despite increasing attention to technical aspects (i.e., software and hardware) of cybersecurity, many professionals and scholars pay little or no attention to its organizational aspects, particularly to stakeholders’ perceptions of the status of cybersecurity within organizations. Given that cybersecurity decisions and policies are mainly made based on stakeholders’ perceived needs and security views, it is critical to measure such perceptions. In this paper, we introduce a methodology for analyzing differences in perceptions of cybersecurity among organizational stakeholders. To measure these perceptions, we first designed House of Security (HoS) as a framework that includes eight constructs of security: confidentiality, integrity, availability, technology resources, financial resources, business strategy, policy and procedures, and culture. We then developed a survey instrument to analyze stakeholders’ perceptions based on these eight constructs. In a pilot study, we used the survey with people in various functional areas and levels of management in two energy and ICS organizations, and conducted a gap analysis to uncover differences in cybersecurity perceptions. This paper introduces the HoS and describes the survey instrument, as well as some of the preliminary findings.
In an era of big data, organizations are undergoing a revolution in the ways in which data is being visualized and transformed. Successfully navigating this process and completing the transformation requires that organizations not only shift their systems, structures, and human resource, but they must also develop an organizational culture that shares their vision. In this study we bring in three theories: upper echelons, contingency, and organizational adaption, to discuss what conditions affect an organization's decision to hire a Chief Data Officer (CDO), and where a CDO would be more likely to come from. In this research we pair and propose 12 propositions. We postulate that organization size, industry dynamism, diversification strategy, top management team (TMT) functional heterogeneity, Chief Executive Officer (CEO) tenure, and firm performance are six antecedents that have the greatest impact on CDO appointment and origin.
Enhancing student academic performance and transdisciplinary ability is challenging, but the time and effort put into accomplishing this ambitious feat is priceless. We develop secure privacy preserving across Personal Health Data (PHD) repository and single-cell genomics research for building an Innovative Systematic Pedagogy for Integrated Research - Education (INSPIRE) (http://americancse.org/events/csce2017/csce17_awards). In this paper we further build a novel, eclectic, and insightful framework based on classical and popular machine learning approaches to help us meet the educational challenge. Our framework focuses on using integrative research technologies to help solve “Education’s Performance Prediction Data Mining Crisis” (EPPDMC), by putting to rest issues associated with mining and making best use of big data for educational enhancement, such as multi-source education acquisition, data fusion, and unstructured data analysis. We exploit the uses of deep learning, text classification, and semi-supervised learning approaches to solve challenging problems that educators face when analyzing multiplatform big data involved in education, research and training students. Based on new machine learning approached we developed for genomic big-data research and in combination with machine learning methods (http://americancse.org/events/csce2017/keynotes_lectures/yang_talk) and the vast availability of education data available to us, not only can we utilize structured, unstructured, and even multi-media data, but while engaging in leaning intelligent thinking along the way, we can also maximize the utilization of big data by studying the motion and performance of these data. Hence we build the INSPIRE model that can further incorporate Student Face Expression in Class (SFEiC) to help educators and managers make further improvements as they become involved in the teaching-learning process. This research further facilitates the effectiveness of the INSPIRE model.
A new breed of executive, the chief data officer (CDO), is emerging as a key leader in the organization. We provide a three-dimensional cubic framework that describes the role of the CDO. The three dimensions are: (1) Collaboration Direction (inwards vs. outwards), (2) Data Space (traditional data vs. big data) and (3) Value Impact (service vs. strategy). We illustrate the framework with examples from early adopters of the CDO role and provide recommendations to help organizations assess and strategize the establishment of their own CDOs.
Data Quality provides an expos of research and practice in the data quality field for technically oriented readers. It is based on the research conducted at the MIT Total Data Quality Management (TDQM) program and work from other leading research institutions. This book is intended primarily for researchers, practitioners, educators and graduate students in the fields of Computer Science, Information Technology, and other interdisciplinary areas. It forms a theoretical foundation that is both rigorous and relevant for dealing with advanced issues related to data quality. Written with the goal to provide an overview of the cumulated research results from the MIT TDQM research perspective as it relates to database research, this book is an excellent introduction to Ph.D. who wish to further pursue their research in the data quality area. It is also an excellent theoretical introduction to IT professionals who wish to gain insight into theoretical results in the technically-oriented data quality area, and apply some of the key concepts to their practice.
Organizations have increasingly invested in technology and human resources to collect, store, and process vast quantities of data. Even so, they often find themselves stymied in their efforts to translate this data into meaningful insights that they can use to improve business processes, make smart decisions, and create strategic advantages. Issues surrounding the quality of data and information that cause these difficulties range in nature from the technical (e.g., integration of data from disparate sources) to the nontechnical (e.g., lack of a cohesive strategy across an organization ensuring the right stakeholders have the right information in the right format at the right place and time). Although there has been no consensus about the distinction between data quality and information quality, there is a tendency to use data quality (DQ) to refer to technical issues and information quality (IQ) to refer to nontechnical issues. In this chapter, we do not make such distinction and use the term data quality to refer to the full range of issues. More importantly, we advocate interdisciplinary approaches to conducting research in this area. This interdisciplinary nature of research demands that we integrate and introduce research results, regardless of technical and nontechnical in terms of the 16
This is a sound textbook for Information Technology and MIS undergraduate students, and MBA graduate students and all professionals looking to grasp a fundamental understanding of information quality. The authors performed an extensive literature search to determine the Fundamental Topics of Data Quality in Information Systems. They reviewed these topics via a survey of data quality experts at the International Conference on Information Quality held at MIT. The concept of data quality is assuming increased importance. Poor data quality affects operational, tactical and strategic decision-making, and yet error rates of up to 70%, with 30% typical are found in practice (Redman). Data that is deficient leads to misinformed people, who in turn make bad decisions. Poor quality data impedes activities such as re-engineering business processes and implementing business strategies. Poor data quality has contributed to major disasters in the federal government, NASA, Information Systems, Federal Bureau of Investigation, and most busineses. The diverse uses of data and the increased sharing of data that has arisen as a result of the widespread introduction of data warehouses have exacerbated deficiencies with the quality of data (Ballou). In addition, up to half the cost of creating a data warehouse is attributable to poor data quality. The management of data quality so as to ensure the quality of information products is examined in Wang. The purpose of this book is to alert our IT-MIS-Business professionals to the pervasiveness and criticality of data quality problems. The secondary agenda is to begin to arm the students with approaches and the commitment to overcome these problems. The current authors have a combined list of over 200 published papers on data and information quality.
Certain polychlorinated biphenyls (PCB) have long half-lives and, despite the regulatory bans on the industrial pollutants that expose humans to PCB, are detectable in human serum. However, many of them are not detectable because of the small quantities that may be present in body fluids. For this reason, attempts have been made to estimate the total concentration of PCB (SigmaPCB) using the relationship between SigmaPCB and the concentrations of a few of the PCB congeners which can be reliably measured at detectable levels. PCB 153 or a combination of PCB 153,138, and 180 have previously been used for this purpose. However, because of the unique populations investigated in these studies, the results are not necessarily applicable to the racially/ethnically heterogeneous US population. We defined SigmaPCB as the sum of the concentrations of 12 PCB congeners, and sum of 33 PCB congeners for NHANES 2001-2002 and 2003-2004 respectively. We built regression models in a step-wise fashion using SigmaPCB as the dependent variable and age, race/ethnicity, and gender as the covariates for both whole-weight and lipid-adjusted data. In addition, concentration of PCB 153 was used as the continuous independent variable for 2001-2002 models, and PCB 153 and PCB 180 for 2003-2004 models respectively. R(2) for both models for NHANES 2001-2002 was >86%. The R(2) for both NHANES 2003-2004 models was >81%. Thus, the estimate of SigmaPCB for the general US population can be improved by considering common demographic variables, such as race/ethnicity, and selected congeners.