Organizations store data regarding their operations, employees, consumers, and suppliers in their databases. Some of the data are considered confidential, and by law, the organization is required to provide appropriate security measures in order to preserve privacy. Yet a number of companies have little or no security measures. The reason for this lack of security may, at least in part, be attributed to a lack of awareness and empirical evidence about the relative effectiveness of security mechanisms. This study investigates the effectiveness of different security mechanisms for protecting numerical database attributes. The trade-off between security, accessibility, and accuracy are examined. A comparison of different security mechanisms reveals that fixed data perturbation is preferred because it maximizes both security and accessibility. An investigation of the different approaches to fixed data perturbation indicates that multiplicative method best meets these criteria.
Conceptual data modelling (CDM) refers to the phase of the information systems development process that involves the abstraction and representation of the real world data pertinent to an organization. When CDM is properly and rigorously performed, the delivered system is expected to be functionally richer, less error-prone, more fully attuned to meet user needs, more able to adjust to changing user requirements and less expensive. However, there is little evidence that conceptual data modelling for the enterprise is actually conducted. There is the feeling that the 'corporate reality' is much different. In many organizations, CDM is never employed. In others, it is applied in a haphazard, project-to-project basis, thus leading to considerable redundancy. The academic community has mainly focused on proposing semantic data models but has not demonstrated a rigorous basis for conceptual data modelling. Specifically, the community has failed to show how a conceptual data model can map to an accurate logical data model. It is the purpose of this paper to discuss and compare the perspectives of academic and practitioner communities regarding the application of conceptual data modelling.
Statistical databases provide security of confidential data by preventing access to individual values. However, under certain situations, the security of individual data can be compromised by statistical functions alone. A number of approaches have been suggested to counter this problem. This paper addresses one such approach, namely, fixed data perturbation. The purpose of the research is to evaluate the effectiveness of different statistical distributions in perturbing different forms of database populations when employing an additive form of fixed data perturbation. Specifically, this paper evaluates the effectiveness of the Normal, Log-normal, Gamma and Uniform distributions in perturbing data sets with a defined distribution. The results of extensive Monte-Carlo experiments conducted on different database populations reveal that the Uniform distribution provides the best performance (high security and low bias) for all database populations except those described by the Log-normal distribution.
Conceptual and logical database design are complex tasks for non-expert designers. Currently, the popular data models for conceptual and logical database design are the entity–relationship (ER) and the relational model, respectively. Logical design methodologies for relational databases have relied on mathematically rigorous approaches which are impractical, or textbook approaches which do not provide the rich constructs to capture real applications. Consequently, designers have to use their intuition to develop their own rules and heuristics. There is a need, therefore, to develop practical rules and heuristics that can be used to handle the complexity of design in real applications. This paper proposes a realistic and detailed approach for conceptual design using the ER model for relational databases. The approach is based on four rules that specify the order in which various types of relationships must be modelled, three rules that pertain to detection of derived relationships, and three heuristics based on observation of constructs in real applications. The approach is illustrated by many examples.
Design aids can improve the quality of systems developed by end-users and non-expert designers. This paper reports a study undertaken to establish the concept validation of a design aid that is based on feedback to improve the quality of conceptual and logical relational databases. We describe the design of SERFER (Simulated ER based FEedback system for R elational databases) and test its effectiveness in a laboratory experiment using the "hidden operator" method. The results show that feedback can help users detect and correct certain types of database design errors in modeling ternary relationships. However, no improvement seems possible in the case of unary relationships. The experiment could not determine whether errors can be corrected in modeling binary relationships, since the subjects were reasonably adept and rarely committed serious errors in this case.
A framework is developed to explain human error behavior in modeling conceptual databases. The framework is based on the notion of directness distance or ‘gulf’ suggested in recent literature. It specifies four aspects of ‘gulf’ in the context of conceptual database design - syntax, mapping, rules, and consistency. Based on the model, six types of errors are suggested - syntactic, abstraction, simplification, overload, convergence, and divergence. These are then matched to errors found in four empirical studies on database representation. Four types of errors - convergence, abstraction, simplification and overload - were typically found in these studies. The paper provides design guidelines to prevent these errors.
Our objective in this paper is to provide a thorough understanding of the usability of data management environments with an end to conducting research in this area. We do this by synthesizing the existing literature that pertains to (i) data modelling as a representation medium and (ii) query interface evaluation in the context of data management. We were motivated by several trends that are prevalent in the current computing context. First, while there seems to be a proliferation of new modelling ideas that have been proposed in the literature, commensurate experimental evaluation of these ideas is lacking. Second, there appears to exist a significant user population that is quite adept at working in certain computing environments (e.g. spreadsheets) with a limited amount of computing skills. Finally, the choices in terms of technological platforms that are now available to implement new software designs allow us to deal with the implementation issue more effectively. The outcomes of this paper include a delineation of what constitutes an appropriate conceptualization of this area and a specification of research issues that tend to dominate the design of a research agenda.
Planners and public policy makers in recent years have become increasingly concerned with issues related to nuclear waste transportation. Rational planning and policy for nuclear waste transportation depends upon systematically assimilating into a coherent and meaningful whole, diverse information at several levels of analysis. The authors describe and demonstrate the rudiments of a deep knowledge architecture for evaluating alternative nuclear waste transshipment possibilities.