An agent-based simulation model hierarchy emulating disease states and behaviors critical to progression of diabetes type 2 was designed and implemented in the DEVS framework. This model was built to approximately reproduce some essential findings that were previously reported for a rather complex model of diabetes progression. Our models are translations of basicelements of this previously reported system dynamics model of diabetes. The system dynamics model, which mimics diabetes progression over an aggregated US population, was disaggregated and reconstructed bottom-up at the individual (agent) level. Four levels of model complexity were defined in order to systematically evaluate which parameters are needed to mimic outputs of the system dynamics model. The four estimated models attempted to replicate stock counts representing disease states in the system dynamics model while estimating impacts of an elderliness factor, obesity factor and health-related behavioral parameters. Health-related behavior was modeled as a simple realization of the Theory of Planned Behavior, a joint function of individual attitude and diffusion of social norms that spread over each agent’s social network. Although the most complex agent-based simulation model contained 31 adjustable parameters, all models were considerably less complex than the system dynamics model which required numerous time series inputs to make its predictions. All three elaborations of the baseline model provided significantly improved fits to the output of the system dynamics model, although behavioral factors appeared to contribute more than the elderliness factor. The results illustrate a promising approach to translate complex system dynamics models into agent-based model alternatives that are both conceptually simpler and capable of capturing main effects of complex local agent-agent interactions.
System dynamics models are usually used to investigate aggregate level behavior, but these models can be decomposed into agents that have more detailed individual behaviors. Here we develop a simple model of the STEM workforce to illuminate the impacts that arise from the disaggregation and refinement of system dynamics models via agent-based modeling. Particularly, alteration of Poisson assumptions, adding heterogeneity to decision-making processes of agents, and discrete-time formulation are investigated and their impacts are illustrated. The goal is to demonstrate both the promise and danger of agent-based modeling in the context of a relatively simple model and to delineate the importance of modeling decisions that are often overlooked.
he role of big data in addressing the needs of the present healthcare system in US and rest of the world has been echoed by government, private, and academic sectors. There has been a growing emphasis to explore the promise of big data analytics in tapping the potential of the massive healthcare data emanating from private and government health insurance providers. While the domain implications of such collaboration are well known, this type of data has been explored to a limited extent in the data mining community. The objective of this paper is two fold: first, we introduce the emerging domain of "big" healthcare claims data to the KDD community, and second, we describe the success and challenges that we encountered in analyzing this data using state of art analytics for massive data. Specifically, we translate the problem of analyzing healthcare data into some of the most well-known analysis problems in the data mining community, social network analysis, text mining, and temporal analysis and higher order feature construction, and describe how advances within each of these areas can be leveraged to understand the domain of healthcare. Each case study illustrates a unique intersection of data mining and healthcare with a common objective of improving the cost-care ratio by mining for opportunities to improve healthcare operations and reducing what seems to fall under fraud, waste, and abuse.
of need (CON) state. Results: There were 4135 radiation oncologists who received a total of $1,499,625,803 in payments from Medicare in 2012. Seventy-five percent of radiation oncologists were male. The median reimbursement was $146,453. The code with the highest total reimbursement was 77418 (radiation treatment delivery intensity modulated radiation therapy [IMRT]). The most commonly billed evaluation and management (E/M) code for new visits was 99205 (49%). The most commonly billed E/M code for established visits was 99213 (54%). Forty percent of providers billed none of their new office visits using 99205 (the highest E/M billing code), whereas 34% of providers billed all of their new office visits using 99205. For the 1510 radiation oncologists (37%) who billed technical services, median Medicare reimbursement was $606,008, compared with $93,921 for all other radiation oncologists (P<.001). On multivariate analysis, technical services billing (P<.001), male sex (P<.001), and rural location (P=.007) were predictive of higher Medicare reimbursement. Conclusions: The billing of technical services, with their high capital and labor overhead requirements, limits any comparison in reimbursement between individual radiation oncologists or between radiation oncologists and other specialists. Male sex and rural practice location are independent predictors of higher total Medicare reimbursements.
The authors propose a new online collaborative tool for visually understanding national health indicators, which facilitates the full spectrum of investigation of indicators, from an overview of all the correlation coefficients between variables, to investigation of subsets of selected variables, and to individual data element analysis. this tool is publicly accessible at http://cda.ornl.gov/heat/heatmap.html. In this paper, they discuss the key issues regarding the interface design and implementation. They also illustrate how to use their interface for analyzing the health indicator dataset by showing some key system views. In the end, they introduce and discuss some ongoing research efforts extending this work.
Fault tree analysis is a method for evaluating reliability and availability in terms of equipment system “states”, but this method does not lend itself easily to the evaluation of equipment interactions through time. This makes fault trees difficult to use for the analysis of systems whose reliability and availability depend on complex interactions between its subsystems. This difficulty is overcome by combining fault trees with discrete event simulation methods. The new TRAM methodology combines models and techniques for the analysis of throughput, availability, reliability, and maintainability into a single approach. This paper describes the TRAM methodology and illustrates it with an application to a chemical processing plant. TRAM combines fault tree analysis at a low level of the system description and discrete event simulation at a higher level to create a new method for analyzing the availability and throughput capacity of material processing plants. Failure and repair data is modeled stochastically by a very flexible type of finite mixture distribution that allows the analyst to separate the effects of different repair strategies, such as the reliance on procurement of off-site (vs. on-site) spare parts. An important application of the TRAM method is to facilitate the design of a plant that tolerates outages of its subsystems in the most efficient way possible. Mitigation strategies including in-process storage, alternate work-flows, availability of spare parts, and design for over-production: all of these can be assessed using the TRAM approach, and it thereby facilitates the design of more robust manufacturing systems. The TRAM methodology enables sophisticated “what-if” analyses of alternative designs, e.g. equipment sets, capacities (tanks sizes), shift schedules, spare parts, etc. to optimize plant design and operation. It is a stochastic, time dependent process that provides probabilities of success (or failure) and confidence bounds on availability and throughput. Finally, the TRAM methodology can help plant managers and owners to focus on the plant production metrics by which they are compensated, and not solely on abstract metrics such as availability. Accordingly, TRAM is potentially a more influential tool in the industry than conventional RAM methods. The TRAM method is based on the discrete event formalism developed by Zeigler et al. [1], and explained further in [2]. In TRAM the plant model is completely separated from the simulation engine and can be specified by input data contained in an XML file. Alternatively, the user can construct connections between subsystem components using a graphical user interface. The GUI is very useful in supporting the verification of the correct mass balance in the model.
The KDD community has described a multitude of methods for knowledge discovery on large datasets. We consider some of these methods and integrate them into an analyst s workflow that proceeds from the data-centric descriptive level to the model-centric causal level. Examples of the workflow are shown for the Health Indicators Warehouse, which is a public database for community health information that is a potent resource for conducting data science on a medium scale. We demonstrate the potential of HIW as a source of serious visual analytics efforts by showing correlation matrix visualizations, multivariate outlier analysis, multiple linear regression of Medicare costs, and scatterplot matrices for a broad set of health indicators. We conclude by sketching the first steps toward a causal dependence hypothesis.
The knowledge management community has introduced a multitude of methods for knowledge discovery on large datasets. In the context of public health intelligence, we integrated and incorporated some of these methods into an analyst's workflow that proceeds from the data-centric descriptive level of analysis to the model-centric causal level of reasoning. We show several case studies of the proposed analyst's workflow as applied to the US Health Indicators Warehouse (HIW), which is a medium scale, public dataset regarding community health information as collected by the US federal government. In our case studies, we demonstrate a series of visual analytics efforts targeted at the HIW, including visual analysis according to correlation matrices, multivariate outlier analysis, multiple linear regression of Medicare costs, confirmatory factor analysis, and hybrid scatterplot and heatmap visualization for distributions of a group of health indicators. We conclude by sketching a preliminary framework for examining causal dependence hypotheses for future data science research in public health.
Climate and human mobility are essentially interconnected and interdependent. Our mobility through ground transportation system powered by fossil fuel is one of the primary forces behind the two major global crises of today’s society, namely energy scarcity and climate change. On the other hand, long term change in climate and frequency of climate extreme events, such as hurricanes, floods, snow and ice storms can have both short term mobility challenges and cause long term human migrations as an adaptation phenomenon. Human settlements develop around stable environments where shelter and sustenance are found, and on a higher level, economies can be built. Climate change is expected to shift these stable regions and consequently induce large scale migrations, which in turn can result in famine, cultural conflict, disease propagation, and stress on natural resources and critical infrastructures. In the near term, to reduce oil dependence, environmental impacts, and congestion, a number of alternative energy supply, distribution, and end-use transportation systems, technologies and policies are presently being explored. However, it is still unclear when and in what precise combination these sources and technologies will emerge as successful and sustainable solutions. Ideally, future plausible development and implementation strategies for alternative energy resources and technologies will secure and support a societal system in which energy, environment, and mobility interests are simultaneously optimized. In the longer term, under climate change scenarios it is plausible to expect displacement, migration, and resettlement as an interactive consequence of climate change and its impacts on water resources, land cover and land use. Given the intertwined nature of such a system across wide geographic scales, assessing the effectiveness of possible planning strategies and discovering their unanticipated consequences require data collection, modeling, and simulation at the finest data, process, and societal response levels coupled with the system’s behavior over large spatial and temporal scales.
The system performance metric “availability” is a central concept with respect to the concerns of a plant’s operators and owners, yet it can be abstract enough to resist explanation at system levels. Hence, there is a need for a system-level metric more closely aligned with a plant’s (or, more generally, a system’s) raison d’être. Historically, availability of repairable systems – intrinsic, operational, or otherwise – has been defined as a ratio of times. This paper introduces a new concept of availability, called endogenous availability, defined in terms of a ratio of quantities of product yield. Endogenous availability can be evaluated using a discrete event simulation analysis methodology. A simulation example shows that endogenous availability reduces to conventional availability in a simple series system with different processing rates and without intermediate storage capacity, but diverges from conventional availability when storage capacity is progressively increased. It is shown that conventional availability tends to be conservative when a design includes features, such as in – process storage, that partially decouple the components of a larger system.
Recent events highlight the need for efficient tools for anticipating the threat posed by terrorists, whether individual or groups. Antiterrorism includes fostering awareness of potential threats, deterring aggressors, developing security measures, planning for future events, halting an event in process, and ultimately mitigating and managing the consequences of an event. To analyze such components, one must understand various aspects of threat elements like physical assets and their economic and social impacts. To this aim, we developed a three-layer Bayesian belief network (BBN) model that takes into consideration the relative threat of an attack against a particular asset (physical layer) as well as the individual psychology and motivations that would induce a person to either act alone or join a terrorist group and commit terrorist acts (social and economic layers). After researching the many possible motivations to become a terrorist, the main factors are compiled and sorted into categories such as initial and personal indicators, exclusion factors, and predictive behaviors. Assessing such threats requires combining information from disparate data sources most of which involve uncertainties. BBN combines these data in a coherent, analytically defensible, and understandable manner. The developed BBN model takes into consideration the likelihood and consequence of a threat in order to draw inferences about the risk of a terrorist attack so that mitigation efforts can be optimally deployed. The model is constructed using a network engineering process that treats the probability distributions of all the BBN nodes within the broader context of the system development process.
Radical and contentious activism may or may not evolve into violent behavior depending on contextual factors related to social, political, cultural and infrastructural conditions. Significant theoretical advances have been made in understanding these contextual factors and the import of their interrelations. However, there has been relatively little progress in the development of processes and capabilities that leverage such theoretical advances to automate the anticipatory analysis of violent intent. In this paper, we describe a framework that implements such processes and capabilities, and discuss the implications of using the resulting system to assess the emergence of radicalization leading to violence.
In a complex environment, it can be difficult to assess the degree to which a continuous variable influences microbial community structure. We propose a method that involves using the community data to the value of the presumed dominant variable. The assumption is that in order to the variable the community composition must be sensitive to or affected by the variable in question. The concentration range over which the prediction is accurate should thus provide information on the concentrations that influence community structure. We explored this approach using T-RLFP data on at a site polluted by tannery wastes. We were able to use the microbial community structure measures to predict Cr concentration over a surprisingly wide range of concentration. Although, it appears from this work that this approach can give useful information about the relationships between microbial community structure and specific environmental conditions much further testing is required.