OBJECTIVE:Poor clinical data quality might affect clinical decision making and patient treatment. This study identifies quality defects in clinical data collected automatically by bedside monitoring devices in the Intensive Care Unit (ICU) and examines their effect on clinical decisions.METHODS:Real-world data collected from 7688 patients admitted to the general ICU in a tertiary referral hospital over seven years was retrospectively analyzed. Data quality defect detection methods that use time-series analysis techniques identified two types of data quality defects: (a) completeness: the extent of non-missing values, and (b) validity: the extent of non-extreme values within the continuous range of values. Data quality defects were compared to five scenarios of medication and procedure prescriptions that are common in ICU settings: Blood-pressure reduction, blood-pressure elevation, anesthesia medications, intubation procedures, and muscle relaxant medications.RESULTS:Results from a logistic regression revealed a strong connection between data quality and the clinical interventions examined: lower validity level increased the likelihood of prescription decisions for all five scenarios, and lower completeness level increased the likelihood of prescription decisions for some scenarios.DISCUSSION:The results highlight the possible effect of data quality defects on physicians' decisions. Lower validity of certain key clinical parameters, and in some scenarios lower completeness, correlated with stronger tendency to prescribe medications or perform invasive procedures.CONCLUSIONS:Data quality defects in clinical data affect decision making even without practitioners' awareness. Thus, it is important to emphasize these effects to ICU staff, as well as to medical device manufacturers.
Visualization tools are critical components of cyber security systems allowing analyzers to better understand, detect and prevent security breaches. Security administrators need to understand which users accessed the database and what operations were performed in order to detect irregularities. The current work compares the Sankey diagram with the more commonly used node-link diagram as an alternative visualization technique for cyber security tasks in a controlled experiment. The results indicate, that the Sankey tool showed a consistent advantage in task completion time and was more effective (measured by the percent of correct answers) in synoptic tasks, while the Node-link diagram was more effective in basic, elementary tasks. Further results revealed that performance had only a small effect on user satisfaction and preferences. Our results suggest that the Sankey tool may be a viable option for cyber security visualization tools and strengthens the need to provide personalized visualization tools based on user preferences.
This chapter describes the implications for managing metadata, a higher-level abstraction of data that exists within repositories, applications, systems and organizations. Metadata is a key factor for the successful implementation of complex decision environments. Managing metadata offers significant benefits and poses several challenges, due to the complex nature of metadata. The complexity is demonstrated by reviewing different functions that metadata serves in decision environments. To fully reap the benefits of metadata, it is necessary to manage metadata in an integrated manner. Crucial gaps for integrating metadata are identified by comparing the requirements for managing metadata with the capabilities offered by commercial software products designed for managing it. The chapter then proposes a conceptual architecture for the design of an integrated metadata repository that attempts to redress these gaps. The chapter concludes with a review of emerging research directions that explore the contribution of metadata in decision environments.
Decision making is often supported by decision models. This study suggests that the negative impact of poor data quality (DQ) on decision making is often mediated by biased model estimation. To highlight this perspective, we develop an analytical framework that links three quality levels - data, model, and decision. The general framework is first developed at a high-level, and then extended further toward understanding the effect of incomplete datasets on Linear Discriminant Analysis (LDA) classifiers. The interplay between the three quality levels is evaluated analytically - initially for a one-dimensional case, and then for multiple dimensions. The impact is then further analyzed through several simulative experiments with artificial and real-world datasets. The experiment results support the analytical development and reveal nearly-exponential decline in the decision error as the completeness level increases. To conclude, we discuss the framework and the empirical findings, elaborate on the implications of our model on the data quality management, and the use of data for decision-models estimation.
The integration of smart devices and sensors together with extensive data processing capabilities into the electrical power grid is a fundamental infrastructure of Smart Grid's environments. Such capabilities introduce major Data Quality challenges: The optimization of electricity production and consumption in such environments relies on collecting and analyzing vast amounts of sensor-based data samples in real time. Degradation in the quality of such data might hinder its analysis and result in sub-optimal Smart Grid configuration. This study aims at exploring the effect of two Data Quality determinants in Smart Grid's environments — sampling frequency, reflecting the temporal distribution, and sampling density, reflecting spatial distribution. Beyond technical aspects, sampling density and frequency have economic implications, which must affect their optimal configuration. This study contributes to further conceptualization of these Data Quality determinants and assessing their impact in Smart Grid's environments, by developing an analytical model that links their configuration to cost-benefit tradeoffs. This manuscript presents the model development, and its preliminary evaluation with a large-scale datasets that reflects energy consumption in a real-world environment.
Data currency declines, caused by recorded data values becoming outdated, can damage the usability and accountability of data resources. Detecting and updating outdated values may improve data currency and reduce the associated damage, but such efforts may be costly and cannot always be justified. This study models currency decline scenarios using a continuous-time Markov chain stochastic process with a finite number of states, each reflecting a valid data value. The model considers state transition probabilities, transition time distributions, and the tradeoff between the damage associated with outdated data and the cost of reacquisition. The proposed formulation permits the currency level to be estimated without having to rely on a baseline for comparison, as well as the prediction of future currency declines, assessment of their accumulated damage, and optimization of the timing of cost-effective data auditing and reacquisition. The study introduces a comprehensive evaluation of the proposed model, using a large real-world dataset relating to the handling of insurance claims over multiple time periods. The evaluation results highlight the applicability of the model, and its potential contribution to proactive data quality management and cost-effective handling of currency declines.
The Performance Measurement System (PMS) explored in the study was implemented by public police forces, using advanced Business Intelligence (BI) technologies. The study examines the impact of enhancing that PMS, through analysis of metric results over an 8-year period that covered a transition between two major system versions. The analysis results indeed show a significant impact of transitioning to the new PMS in most (75 %) performance metrics. A noticeable impact of the transition is the temporary performance decline, followed by some improvement that can be attributed in part to the redefinition of some metrics. Further, the results confirmed the preliminary assumptions that the improvement in the measured performance is positively and significantly associated with human-resource allocation; however, with some mediation effects of the crime category and the organization unit.
The Smart Grid (SG) concept reflects the integration of Information and Communication Technologies (ICT) together with Internet of Things (IoT) technologies into the electrical power grid, which enables devices to produce data regarding their energy consumption and, by that, to manage the electricity consumption and production in a smarter way. The vast amounts of data generated and processed in SG environments raise the issue of data quality (DQ) management. The “Quality of Context” approach observes DQ from a business-value perspective, focusing more on data contents and use and less on its physical characteristics. Accordingly, this study explores two SG-relevant DQ aspects - sampling frequency and density. The former reflect the high sampling rate needed for ubiquitous computing environments such as the SG, while the latter reflects real-world limitations on sensor infrastructures. As the study progresses, our goal is to further conceptualize these DQ dimensions and evaluate their impact in both simulated and real-world SG environments, toward defining mechanism for their optimal configuration.
Overcrowding at EDs is a world known problem which negatively affects the quality of medical care. It is evident, for example, by long patients’ waiting times. EDs frequently use average patients’ length of stay (LOS) as a performance indicator. Long LOS is commonly correlated with overcrowding. In this research, a prototype of an electronic online digital dashboard, termed operational BI, was developed following the Design Science Research methodology. The system is targeted to be used by the ED staff. Simulation was used to assess if such a system, which displays ED’s critical data in a dashboard visualisation, can decrease LOS, thereby ease EDs’ overcrowding. Six scenarios were simulated, depicting various use patterns. The results show a potential decrease of 34–44% in average LOS, depending on the use pattern. These are promising results that call for further, real-world examination of the system.
With the aim of bridging the gap between well-established research on information technology (IT) value creation and the emergent study of business intelligence (BI), this study develops and tests a model of BI value creation that is firmly anchored in both streams of research. The analysis draws on the resource based view and on conceptualizations of organizational learning to hypothesize about the paths by which BI assets and BI capabilities create business value. The research model is first assessed in an exploratory analysis of data collected through interviews in three firms and then tested in a confirmatory analysis of data collected through a survey. (C) 2016 Elsevier B.V. All rights reserved.
The Smart Grid (SG) concept reflects the integration of Information and Communication Technologies (ICT) together with Internet of Things (IoT) technologies into the electrical power grid, which enables devices to produce data regarding their energy consumption and, by that, to manage the electricity consumption and production in a smarter way. The vast amounts of data generated and processed in SG environments raise the issue of data quality (DQ) management. The "Quality of Context" approach observes DQ from a business-value perspective, focusing more on data contents and use and less on its physical characteristics. Accordingly, this study explores two SG-relevant DQ aspects - sampling frequency and density. The former reflect the high sampling rate needed for ubiquitous computing environments such as the SG, while the latter reflects real-world limitations on sensor infrastructures. As the study progresses, our goal is to further conceptualize these DQ dimensions and evaluate their impact in both simulated and real-world SG environments, toward defining mechanism for their optimal configuration.
Scenarios, in which real-world state transitions are documented by a sequence of data records, introduce unique data quality (DQ) challenges. Due to time and workload constraints, one might choose to update the values only for a subset of the required attributes, while replicating previous values of others. As a result, the record sequence might reflect the current state incorrectly and fail to capture critical transitions. Our model addresses such scenarios by evaluating record sequences and alerting on high likelihood of erroneous replication. The metrics consider attribute characteristics, the distances between consecutive values, and the likelihood of value transition. The potential contribution of the model is demonstrated with a preliminary evaluation of 200 real-world records collected in an Obstetrics unit in a large hospital. A model trained with 100 records reached accuracy level of 85% with detecting data acquisition flaws in the other 100. The manuscript introduces the model development, describes the preliminary evaluation with real-world data, and highlights directions for future research progress. • Information Systems Database Management Systems Data Cleaning • Medical Information Policy Medical Records
This paper presents the vision, construction, and implementation of an Integrated Manufacturing Technology (IMT) laboratory for higher education. The laboratory introduces integration of Information Technology (IT) in manufacturing processes and provides hands-on learning experience in an authentic, holistic, and automated production environment, not commonly found in academic teaching. The manufacturing process includes fully-automated, semi-automated, and manual stations and is supported by a heterogeneous IT infrastructure, including manufacturing databases, shop-floor control, data warehousing, business intelligence, and enterprise systems. The IMT laboratory is used to facilitate learning in a databases course and a business intelligence course, resulting in positive feedback from the students.
Performance Measurement Systems (PMS) have long captured the attention of organizational behavior and information systems (IS) research. The PMS in the study was implemented by public police forces, using advanced Business Intelligence (BI) technologies. The study examines the impact of enhancing that PMS, through analysis of the metric results over an 8-year time period that covered a transition between two major system versions. The analysis results indeed show a significant impact of transitioning to the new PMS in most (75%) performance metrics. A noticeable impact of the transition is the temporary performance decline, followed by some improvement that can be attributed in part to the redefinition of some metrics. Further, the results confirmed the preliminary assumptions that the improvement in the measured performance is positively and significantly associated with human-resource allocation; however, with some mediation effects of the crime category and the organization unit.
When there is a disparity in the value of different data records and fields, there is a need for an optimization of data resources.Not all data necessarily contribute the same value.It depends on the usage of the data, as well as a variety of other factors.This paper presents models for optimizing data management in the presence of a disparity between the values contributed by different data.We expound on what disparity of data value represents and illustrate models to derive a numerical measure of such disparity.We then use real-world data from a large data resource used to manage alumni relations, and demonstrate our optimization methods and results.We then discuss the tradeoffs involved between value and cost, and the implications for data management, both in this real-world context and in general.
Information systems aid managers in their marketing decision making as they enable marketers to understand their customers better and come up with effective solutions to satisfy their needs, and provide the computer-based tools, models and methods to analyze the data. In addition, information systems can also be central to the provision of customer-relationship management tools. Thus, we suggest that the design and configuration of these systems must reflect assessments of cost-benefit tradeoffs beside satisfying technical and functional requirements. In this study, we develop a framework for an economics-driven assessment of alternatives for designing marketing information systems. The framework defines different design strategies that consider uncertainties in utilization, cost and performance differentials between technologies, and penalties resulting from delaying implementation. Our evaluation highlights conditions under which one design strategy will outperform others.
Data quality (DQ) might degrade over time, due to changes in real-world entities or behaviors that are not reflected correctly in datasets that describe them. This study presents a continuous-time Markov-Chain model that reflects DQ as a dynamic process. The model may help assessing and predicting accuracy degradation over time. Taking into account cost-benefit tradeoffs, it can also be used to recommend an economically-optimal point in time at which data values should be evaluated and possibly reacquired. The model addresses data-acquisition scenarios that reflect real-world processes with a finite number of states, each described by certain data-attribute values. It takes into account state-transition probabilities, the distribution of time spent in each state, the damage associated with incorrect data that fails to reflect the real-world state, and the cost of data reacquisition. Given current state and the time passed since the last transition, the model estimates the expected damage of a data record and recommends whether or not to correct it, by comparing the potential benefits of correction (elimination of potential damage), versus reacquisition cost. Following common design science research guidelines, the applicability and the potential contribution of the model is demonstrated with a real-world dataset that reflects a process of handling insurance claims. Insurants’ status must be kept up-to-date, to avoid potential monetary damages; however, contacting an insurant for status update is costly and time consuming. Currently the contact decision is guided by some heuristics that are based on employees’ experience. The evaluation shows that applying the model has major cost-saving potential, compared to the current state.
The banking industry is a strong leader in the use of Internet technologies for revolutionizing customer services. This study explores demographic and financial factors that may affect the use of Internet-Banking IB services provided by a large bank. The study analyzed certain IB activities for a large sample of the bank's customers. The analysis highlights some usage characteristics and patterns that have evolved around the more traditional IB services, such as account-status inquiries and fund transfers. However, with a newly-developed Business-Intelligence BI application, such patterns have not evolved yet. This can be explained by the different nature of this novel BI application, and by the time required for end-users to assimilate and adopt such an innovative application. The findings can help understanding customers' IB needs, detecting customer-segments that use the IB services differently, and help developing and personalizing advanced IB services, such as the online BI tool.
Performance measurement, as an effective tool for implementing organizational strategy and assisting ongoing control and surveillance, is broadly adopted today. The performance measurement system (PMS) explored in this case study was implemented, using business intelligence (BI) technologies, for a public police force. The system lets police commanders view and analyze the performance scores of their own units and get feedback on the success of their activities. The study examines the system's impact, through analysis of the metric results over a time period of five years. The results show that the vast majority of the metrics examined indeed improved. Further, the results underscore the moderation effect of relative metrics weights, as well as the different behavior of metrics that reflect activity versus those that reflect outcomes. The study underscores both the positive and the negative aspects of those results, and discusses their implications for future PMS implementation with BI technologies.
Data stream mining (DSM) deals with continuous online processing and evaluation of fastaccumulating data, in cases where storing and evaluating large historical datasets is neither feasible nor efficient. This research introduces the Multiple Sliding Windows (MSW) algorithm, and demonstrates its application for a DSM scenario with discrete independent variables and a continuous dependent variable. The MSW development emerged from the need to dynamically allocate computational resources that are shared by many tasks, and predicts the required resources per task. The algorithm was evaluated with a large real-world dataset that reflects resource allocation at Intel's global data servers cloud. The evaluation assesses three MSW treatments: the use of multiple slidingwindows, a novel iterative mechanism for feature selection, and adaptive detection of concept drifts. The evaluation showed positive and significant results in terms of prediction quality and the ability to adapt to swift and/or graduate changes in data stream characteristics. Following the successful evaluation, the adoption of the proposed MSW solution by Intel led to cost savings estimated in millions of dollars annually. While evaluated in a specific context, the generic and modular definition of the MSW permits implementation in other domains that deal with DSM problems of similar nature.