Institutional databases can be instrumental in understanding a business process, but additional databases are also needed to broaden the empirical perspective on the investigation. We present a few data mining principles by which a business process can be analyzed in quantitative details and new process components can be postulated. Sequential and parallel process decomposition can apply, guided by human understanding of the investigated process and the results of data, mining. In a repeated cycle, human operators formulate open questions, use queries to get relevant data, use quests that invoke automated search, and interpret the discovered knowledge. As an example we use mining for knowledge about student enrollment, which is an essential part of the university educational process. The target of discovery has been quantitative knowledge useful in understanding the university enrollment. Many discoveries have been made. The particularly surprising findings have been presented to the university administrators and affected the institutional policies.
Knowledge representation which is internal to computer lacks empirical meaning so it is insufficient for the investigation of the external world. Operational definitions are necessary to provide empirical meaning of concepts, but they were largely ignored by the research on automation of discovery and in AI. Operational definitions can be viewed as algorithms that operate in the real world. Each provides a mapping from objects to numbers. Such mappings prompt an analogy with geometry of differential manifolds. We discuss philosophical foundations of the analogy between geometry and object descriptions by many operationally defined concepts. Many operational definitions that are needed for one concept are analogous to many local maps in an atlas on a differentiable manifold. Conceptual framework of differential manifolds can be used to define the notion of procedure equivalence, as well as the requirement that all operational definitions of the same concept must form a coherent set. No set of operational definitions is complete so that expansion of operational definitions is one of the key tasks. Among many possible expansions, only a very special few lead to a satisfactory growth of scientific knowledge. The selection criteria can be represented in the framework of differential manifolds.
Many reported discovery systems build discrete models of hidden structure, properties, or processes in the diverse fields of biology, chemistry, and physics. We show that the search spaces underlying many well-known systems are remarkably similar when re-interpreted as search in matrix spaces. A small number of matrix types are used to represent the input data and output models. Most of the constraints can be represented as matrix constraints; most notably, conservation laws and their analogues can be represented as matrix equations. Typically, one or more matrix dimensions grow as these systems consider more complex models after simpler models fail, and we introduce a notation to express this. The novel framework of matrix-space search serves to unify previous systems and suggests how at least two of them can be integrated. Our analysis constitutes an advance toward a generalized account of model-building in science.
Operational definitions link scientific attributes to experimental situations, prescribing for the experimenter the actions and measurements needed to measure or control attribute values. While very important in real science, operational procedures have been neglected in machine discovery. We argue that in the preparatory stage of the empirical discovery process each operational definition must be adjusted to the experimental task at hand. This is done in the interest of error reduction and repeatability of measurements. Both small error and high repeatability are instrumental in theory formation. We demonstrate that operational procedure refinement is a discovery process that resembles the discovery of scientific laws. We demonstrate how the discovery task can be reduced to an application of the FAHRENHEIT discovery system. A new type of independent variables, the experiment refinement variables, have been introduced to make the application of FAHRENHEIT theoretically valid. This new extension to FAHRENHEIT uses simple operational procedures, as well as the system's experimentation and theory formation capabilities to collect real data in a science laboratory and to build theories of error and repeatability that are used to refine the operational procedures. We present the application of FAHRENHEIT in the context of dispensing liquids in a chemistry laboratory.
Using un-repeatable data and forgetting about measurement error are two cardinal sins in empirical sciences. A machine discovery system must be able to handle both before attempting serious discoveries. We describe an application of the discovery system FAHRENHEIT in a science laboratory, focused on the preparatory stage of the empirical discovery process, i.e. the investigation of repeatability and the measurement of error. To cope with real-world empirical discovery, FAHRENHEIT has been reorganized as a distributed multi-process system and a robotic component including external manipulators and measuring instruments. We present the application of FAHRENHEIT to an experiment in which the system discovers repeatability conditions and error in the context of dispensing liquids in a chemistry laboratory. We then present the theory of the process. Many quantitative discovery systems distinguish between dependent and independent control variables. We argue that to handle repeatability and error, independent variables should be divided further into two categories: theory formation variables and experiment refinement variables. The former have been used by BACON, FAHRENHEIT and other empirical discovery systems to re-discover scientific laws. The latter are used to determine the repeatability and error, prior to the system discovery of the main theory using theory formation variables.
Systems that discover empirical equations from data require large scale testing to become a reliable research tool. In the central part of this paper we discuss two convergence tests for large scale evaluation of equation finders and we demonstrate that our system, which we introduce earlier, has the desired convergence properties. Our system can detect a broad range of equations useful in different sciences, and can be easily expanded by addition of new variable transformations. Previous systems, such as BACON or ABACUS, disregarded or oversimplified the problems of error analysis and error propagation, leading to paradoxical results and impeding the true world applications. Our system treats experimental error in a systematic and statistically sound manner. It propagates error to the transformed variables and assigns error to parameters in equations. It uses errors in weighted least squares fitting, in the evaluation of equations, including their acceptance, rejection and ranking, and uses parameter error to eliminate spurious parameters. The system detects equivalent terms (variables) and equations, and it removes the repetitions. This is important for convergence tests and system efficiency. Thanks to the modular structure, our system can be easily expanded, modified, and used to simulate other equation finders.
We describe an application of the discovery system FAHRENHEIT in a chemistry laboratory. Our emphasis is on automation of the discovery process as oposed to human intervention and on computer control over real experiments and data collection as opposed to the use of simulation. FAHRENHEIT performs automatically many cycles of experimentation, data collection and theory formation. We report on electrochemistry experiments of several hour duration, in which FAHRENHEIT has developed empirical equations (quantitative regularities) equivalent to those developed by an analytical chemist working on the same problem. The theoretical capabilities of FAHRENHEIT have been expanded, allowing the system to find maxima in a dataset, evaluate error for all concepts, and determine reproducibility of results. After minor adjustments FAHRENHEIT has been able to discover regularities in maxima locations and heights, and to analyse repeatability of measurements by the same mechanism, adapted from BACON, by which all numerical regularities are detected.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
In this paper we track the development of research in empirical discovery. We focus on four machine discovery systems that share a number of features: the use of data-driven heuristics to constrain the search for numeric laws; a reliance on theoretical terms; and the recursive application of a few general discovery methods. We examine each system in light of the innovations it introduced over its predecessors, providing some insight into the conceptual progress that has occurred in machine discovery. Finally, we reexamine this research from the perspectives of the history and philosophy of science.