NOTE: The first page of text has been automatically extracted and included below in lieu of an abstract Technology and Information Management Program Abstract This paper describes a new graduate program in Technology and Information Management (TIM) being developed by the Jack Baskin School of Engineering at the University of California, Santa Cruz. As a University of California graduate program, it proposes1 to offer both the M.S. and Ph.D. degrees, with the M.S. intended to prepare its graduates for careers in “high-tech” firms of Silicon Valley, California, and elsewhere. We view TIM as a new and distinct discipline within engineering, combining technology management, systems engineering, and information technology. As an engineering program, TIM addresses both the Management of Technology (MOT) and the Technology of Management (TOM). In MOT, initial emphasis is on the development of theory, analytical results, methods and tools that more closely couple economic factors into engineering and product decisions of firms. This includes studies of the role of information technology in the management of complex systems of both technology and people. In TOM, the emphasis is on development of both theory and software to enable organizations to manage large collections of data in a way that preserves and enhances the information and knowledge that data represents, as well as enabling people in an organization to retrieve that information in a timely and comprehensible way, in areas from manufacturing to sales to services, and across the enterprise functions of analysis, planning and operations. In summary, the domain of the TIM program is: 1) the management of technology and innovation, with emphasis on analytic approaches to complex problems whose solutions have both technological and financial components, and 2) the development of technology of management. Information technology, and information systems and services are core components of both. Background The central technology in most complex systems today is information technology, and the most rapidly changing business environments are in the areas of information technology and complex system design. The design and management of these complex systems presents challenges to enterprises and to individual executives and managers. This is especially true in organizations which are in, or critically depend upon, rapidly changing technologies and rapid introduction of new products and services in competitive environments (a definition of “high tech” enterprises). By using, synthesizing and extending ideas from traditional fields such as computer science, economics, and business management, TIM studies and teaches its students how the use of information technology can lead to more effective management of enterprises (to achieve “competitive advantage”), and more generally, addresses problems in the design and management of complex systems involving people, technology and organizations. Research and teaching programs combining technology, systems, and management are much needed especially today, as firms deal with more complex decisions in a global environment, and address such topics as “out-sourcing” and “off-shoring” of many of the traditional technology enterprise functions and engineering and support tasks to lower cost regions and countries such as India and China. The challenges faced in “high tech” enterprises require the integration of technology and business understanding to solve complex interdisciplinary problems. Management of the
We assessed the performance of Kalman Filtering-Smoothing and Expectation Maximization (EM-KF) in gap filling snow-depth sensor data. To this end, hourly snow-depth data from three spatially dense wireless-sensor networks were randomly removed and imputed using EM-KF. Maximum gap size was 40 + hours and differences between artificially removed and gap-filled data larger than 20 cm were removed before computing Root Mean Square Error and Bias. These differences are spurious over- and underestimations that can generally be identified through visual inspection, likely due to instability in the gap-filling process. The expected accuracy of EM-KF in this initial, controlled proof of concept is close to measurement uncertainty (1 to 2 cm for ultrasonic depth sensors). Compared to regressing missing data against nearby sensors, a frequently used strategy in the field, EM-KF tends to yield smaller errors in networks with a comparatively large number of co-located sensors (nine and eight as opposed to four in a third network). In these data-rich networks, maximum differences in daily Root Mean Square Errors between EM-KF and a regression are up to 6 to 8 cm at a daily time scale, with peaks in winter and in particular during snowfalls. EM-KF yields superior results particularly during snowfalls, likely because it exploits the temporal structure and uncertainty in the data through a state-space model. In the third network with fewer co-located sensors, differences in accuracy between EM-KF and a multilinear regression were inconsistent. This implies that the performance of EM-KF benefits from increasing the amount of available information and with increasing dependency of data across nodes. Temporally and spatially dense snow-depth data are being increasingly collected in operational contexts: EM-KF may support supervised filling of gaps in these data - particularly during snowfall events - and thus provide continuous-time information for avalanche or water-resources forecasting in snow-dominated regions. Automatic, unsupervised gap-filling using EM-KF will necessarily need more research to identify reasons of spurious over- and underestimations. Future work should also upscale this proof of concept to operational sensor networks spanning large water basins to bring conclusions closer to real-world applications.
The prevalence of type 2 Diabetes Mellitus (T2DM) has reached critical proportions globally over the past few years. Diabetes can cause devastating personal suffering and its treatment represents a major economic burden for every country around the world. To property guide effective actions and measures, the present study aims to examine the profile of the diabetic population in Mexico. We used the Karhunen-Loève transform which is a form of principal component analysis, to identify the factors that contribute to T2DM. The results revealed a unique profile of patients who cannot control this disease. Results also demonstrated that compared to young patients, old patients tend to have better glycemic control. Statistical analysis reveals patient profiles and their health results and identify the variables that measure overlapping health issues as reported in the database (i.e. collinearity).
Online Display Advertising’s importance as a marketing channel is partially due to its ability to attribute conversions to campaigns. Current industry practice to measure ad effectiveness is to run randomized experiments using placebo ads, assuming external validity for future exposures. We identify two different effects, i.e., a strategic effect of the campaign presence in marketplaces, and a selection effect due to user targeting; these are confounded in current practices. We propose two novel randomized designs to: (1) estimate the overall campaign attribution without placebo ads, (2) disaggregate the campaign presence and ad effects. Using the Potential Outcomes Causal Model, we address the selection effect by estimating the probability of selecting influenceable users. We show the ex-ante value of continuing evaluation to enhance the user selection for ad exposure mid-flight. We analyze two performance-based (CPA) and one Cost-Per-Impression (CPM) campaigns with 20 million users each. We estimate a negative CPM campaign presence effect due to cross product spillovers. Experimental evidence suggests that CPA campaigns incentivize selection of converting users regardless of the ad, up to 96% more than CPM campaigns, thus challenging the standard practice of targeting most likely converting users. Data, as supplemental material, are available at http://dx.doi.org/10.1287/mksc.2016.0982 .
In this paper, we present a method to dynamically predict the failure of physiological subsystems from patients admitted to the Intensive Care Unit (ICU) using heterogeneous data. We model the probability of failure in each subsystem as a latent state that evolves over time. We propose a method using Generalized Linear Dynamic models to model this latent state which is updated each time new patient data is observed. Then, we estimate the probability of patient mortality as a combination of the estimated probability of failure for different physiological subsystems. We use noun phrase extraction and statistical Topic Models to extract discriminative features which capture the patient health context that can not be obtained when only numerical features are used. We proposed a method of imputing missing values using the non-ignorable nature of the patient data. We test our proposed approach using 15,000 Electronic Medical Records (EMRs) obtained from the MIMIC II public dataset. Experimental results show that the proposed model allows us to predict subsystem failure and mortality probability with high sensitivity and specificity and detect an increase in the probability of mortality.
Patients often search for information on the web about treatments and diseases after they are discharged from the hospital. However, searching for medical information on the web poses challenges due to related terms and synonymy for the same disease and treatment. In this paper, we present a method that combines Statistical Topics Models, Language Models and Natural Language Processing to retrieve healthcare related documents. In addition, we test if the incorporation of terms extracted from the patient's discharge summary improves the retrieval performance. We show that the proposed framework outperformed the winner of the retrieval CLEF eHealth 2013 challenge by 68% in the MAP measure (0:5226 vs 0:3108), and by 13% in NDCG (0:5202 vs 0:3637). Compared with standard language models, we obtain an improvement of 92% in MAP (0:2666) and 45% in NDCG. (0:3637).
In this paper, we present a method to dynamically estimate the probability of mortality inside the Intensive Care Unit (ICU) by combining heterogeneous data. We propose a method based on Generalized Linear Dynamic Models that models the probability of mortality as a latent state that evolves over time. This framework allows us to combine different types of features (lab results, vital signs readings, doctor and nurse notes, etc) into a single state, which is updated each time new patient data is observed. In addition, we include the use of text features, based on medical noun phrase extraction and Statistical Topic Models. These features provide context about the patient that cannot be captured when only numerical features are used. We fill out the missing values using a Regularized Expectation Maximization based method assuming temporal data. We test our proposed approach using 15,000 Electronic Medical Records (EMRs) obtained from the MIMIC II public dataset. Experimental results show that the proposed model allows us to detect an increase in the probability of mortality before it occurs. We report an AUC 0.8657. Our proposed model clearly outperforms other methods of the literature in terms of sensitivity with 0.7885 compared to 0.6559 of Naive Bayes and F-score with 0.5929 compared to 0.4662 of Apache III score after 24 hours.
We propose a user targeting simulator for online display advertising. Based on the response of 37 million visiting users (targeted and non-targeted) and their features, we simulate different user targeting policies. We provide evidence that the standard conversion optimization policy shows similar effectiveness to that of a random targeting, and significantly inferior to other causally optimized targeting policies.
We analyze the causal effect of online ads on the conversion probability of the users who click on the ad (clickers). We show that designing a randomized experiment to find this effect is infeasible, and propose a method to find the local effect on the clicker conversions. This method is developed in the Potential Outcomes causal model, via Principal Stratification to model non-ignorable post-treatment (or endogenous) variables such as user clicks, and is validated with simulated data. Based on two large-scale randomized experiments, performed for 7.16 million users and 22.7 million users to evaluate ad exposures, a pessimistic analysis for this effect shows a minimum increase of the campaigns effect on the clicker conversion probability of 75% with respect to the non-clickers. This finding contradicts a recent belief that clicks are not indicative of campaign success, and provides guidance in the user targeting task. In addition, we find a larger number of converting users attributed to the overall campaign than those attributed based on the click-to-conversion (C2C) standard business model. This evidence challenges the well-accepted belief that C2C attribution model over-estimates the value of the campaign.
In this paper, we propose a framework to dynamically estimate the probability that a patient is readmitted after he is discharged from the ICU and transferred to a lower level care. We model this probability as a latent state which evolves over time using Dynamical Linear Models (DLM). We use as an input a combination of numerical and text features obtained from the patient Electronic Medical Records (EMRs). We process the text from the EMRs to capture different diseases, symptoms and treatments by means of noun phrases and ontologies. We also capture the global context of each text entry using Statistical Topic Models. We fill out the missing values using a Expectation Maximization based method (EM). Experimental results show that our method outperforms other methods in the literature terms of AUC, sensitivity and specificity. In addition, we show that the combination of different features (numerical and text) increases the prediction performance of the proposed approach.
Machine learning classifiers have recently emerged as a way to predict the introduction of bugs in changes made to source code files. The classifier is first trained on software history, and then used to predict if an impending change causes a bug. Drawbacks of existing classifier-based bug prediction techniques are insufficient performance for practical use and slow prediction times due to a large number of machine learned features. This paper investigates multiple feature selection techniques that are generally applicable to classification-based bug prediction methods. The techniques discard less important features until optimal classification performance is reached. The total number of features used for training is substantially reduced, often to less than 10 percent of the original. The performance of Naive Bayes and Support Vector Machine (SVM) classifiers when using this technique is characterized on 11 software projects. Naive Bayes using feature selection provides significant improvement in buggy F-measure (21 percent improvement) over prior change classification bug prediction results (by the second and fourth authors [28]). The SVM's improvement in buggy F-measure is 9 percent. Interestingly, an analysis of performance for varying numbers of features shows that strong performance is achieved at even 1 percent of the original number of features.
We perform a randomized experiment to estimate the effects of a display advertising campaign on online user conversions. We present a time series approach using Dynamic Linear Models to decompose the daily aggregated conversions into seasonal and trend components. We attribute the difference between control and study trends to the campaign. We test the method using two real campaigns run for 28 and 21 days respectively from the Advertising.com ad network.
In this paper, we develop a time series approach, based on Dynamic Linear Models (DLM), to estimate the impact of ad impressions on the daily number of commercial actions when no user tracking is possible. The proposed method uses aggregate data, and hence it is simple to implement without expensive infrastructure. Specifically, we model the impact of daily number of ad impressions on daily number of commercial actions. We incorporate persistence of campaign effects on actions assuming a decay factor. We relax the assumption of a linear impact of ads on actions using the logtransformation. We also account for outliers with long-tailed distributions fitted and estimated automatically without a pre-defined threshold. This is applied to observational data post-campaign and does not require an experimental set-up. We apply the method to data from the Advertising.com ad network on 2,885 campaigns for 1,251 products during six months, to calibrate and perform model selection. We set up a randomized experiment for two campaigns where user tracking is feasible. We find that the output of the proposed method is consistent with the results of A/B testing with similar confidence intervals. Finally, we validate our model with a synthetic public data set, PROMO, and identify 84% of effective campaigns correctly.
In this paper we describe the design, development, and evaluation of a general human-machine interaction search system, and its potential and use in the context of a collaboration project with SAP and Saffron. The objective of a specialized version of the system is to provide medical and healthcare information services to users via interactive search for personalized patient needs. Patients usually have questions regarding healthcare, including those which concern illness symptoms, duration and types of treatment, possible drug effects, and more. Authorized personnel would often be ideal in responding to such needs; however they could potentially be very expensive, and not easy to support and maintain. If patients could have access to information at their home, by means of i-phone or online access, this could save time, doctor office visit expenses, as well as valuable and restricted medical time. What is more, information concerning other anonymized and similar patient cases provides knowledge and perspective on a wide range of patient issues. From the doctors' perspective, they typically need to spend time on differential analysis about new patient cases: study symptoms, research possible causes, rank results by emergency priority and treat them accordingly. A search system that would direct a doctor (or patient/user) to similar patient cases would save significant amount of manual search time. The powerful new feature of this system is the storage and mining of past patient cases knowledge, to create metadata to be used in the subsequent retrieval of relevant documents. Finally, the interactive search system would speed up identification of rare cases; for instance, symptoms that do not appear commonly in past cases may require special treatment or expert referral. We build a model which dynamically learns medical needs of interacting MDs and patients. The model works on free or unstructured text, allowing disambiguation of vague words and flexibility in describing medical needs. In addition, both experts with an advanced knowledge of medical terminology, and beginning users using basic medical terms, can achieve high search relevance. Furthermore, our approach obviates the need for the assignment of tags or labels, such as treatment, symptoms, causes, to documents, to respond effectively to user queries. In particular, we build a temporal difference algorithm to predict user's information needs by incorporating both current and predicted knowledge into learning the user profile. Our source of information about the user consists of submitted queries and feedback on the returned results. We tested our system on publicly available medical data (OhsuMed TREC dataset 2002) and we achieved a significant improvement in retrieval accuracy, compared to the literature. We provide quantitative results as well as demonstration screenshots which illustrate a) the value of interaction (user time spent with system versus results accuracy), b) the value of using medical terminology understanding, when compared with simple general words, and c) the value of allowing the maximum number of feedback submissions to vary.
Most of the relevance feedback algorithms only use document terms as feedback (local features) in order to update the query and re-rank the documents to show to the user. This approach is limited by the terms of those documents without any global context. We propose to use statistical topic modeling techniques in relevance feedback to incorporate a better estimate of context by including global information about the document. This is particularly helpful for difficult queries where learning the context from the interactions with the user is crucial. We propose to use the topic mixture information obtained to characterize the documents and learn their topics. Then, we rank documents incorporating positive and negative feedback by fitting a latent distribution for each class of documents online and combining all the features using Bayesian Logistic Regression. We show results using the OHSUMED dataset for 3 different variants and obtain higher performance, up to 12.5% in Mean Average Precision (MAP).
In this paper, we develop a time series approach, based on Dynamic Linear Models (DLM), to estimate the impact of ad impressions on the daily number of commercial actions when no user tracking is possible. The proposed method uses aggregate data, and hence it is simple to implement without expensive infrastructure. Specifically, we model the impact of daily number of ad impressions in daily number of commercial actions. We incorporate persistence of campaign effects on actions assuming a decay factor. We relax the assumption of a linear impact of ads on actions using the log-transformation. We also account for outliers with long-tailed distributions fitted and estimated automatically without a pre-defined threshold. This is applied to observational data post-campaign and does not require an experimental set-up. We apply the method to data from one commercial ad network on 2,885 campaigns for 1,251 products during six months, to calibrate and perform model selection. We set up a randomized experiment for two campaigns where user tracking is feasible. We find that the output of the proposed method is consistent with the results of A/B testing with similar confidence intervals.
We develop a descriptive method to estimate the impact of ad impressions on commercial actions dynamically without tracking cookies. We analyze 2,885 campaigns for 1,251 products from the Advertising.com ad network. We compare our method with A/B testing for 2 campaigns, and with a public synthetic dataset.
In this paper, we develop an experimental analysis to estimate the causal effect of online marketing campaigns as a whole, and not just the media ad design. We analyze the causal effects on user conversion probability. We run experiments based on A/B testing to perform this evaluation. We also estimate the causal effect of the media ad design given this randomization approach. We discuss the framework of a marketing campaign in the context of targeted display advertising, and incorporate the main elements of this framework in the evaluation. We consider budget constraints, the auction process, and the targeting engine in the analysis and the experimental set up. For the effects of this evaluation, we assume the targeting engine to be a black box that incorporates the impression delivery policy, the budget constraints, and the bidding process. Our method to disaggregate the campaign causal analysis is inspired on randomized experiments with imperfect compliance and the intention-to-treat (ITT) analysis. In this framework, individuals assigned randomly to the study group might refuse to take the treatment. For estimation, we present a Bayesian approach and provide credible intervals for the causal estimates. We analyze the effects of 2 independent campaigns for different products from the Advertising.com ad network for 20M+ users each campaign.