It is often assumed that Big Data resources are too large and complex for human comprehension. The analysis of Big Data is best left to software programs. Not so. When data analysts go straight to complex calculations, before they perform simple estimations, they will find themselves accepting wildly ridiculous conclusions. There is nothing quite like a simple, intuitive estimate to yank an overly-eager analyst back to reality. Often the simple act of looking at a stripped-down version of the problem opens a new approach that can drastically reduce computation time. In some situations analysts will find that a point is reached when higher refinements in methods yield diminishing returns. Advanced predictive algorithms may offer little improvement over a thoughtfully chosen, simple estimator. This chapter suggests a few fast and simple methods for exploring large and complex data sets.
This chapter examines seven future directions for the field of precision medicine: (1) A hypersurveillance future, wherein advances in personal data collection techniques allow us to micromanage wellness; (2) A do-it-yourself future, wherein improved access to the medical literature and to laboratory tests allows patients to diagnose and manage their own diseases; (3) A eugenics future, wherein individuals employ Precision Medicine with the intention of improving the human gene pool; (4) A public health future, in which Precision Medicine will serve whole populations; (5) A data analysis future, in which a new generation of scientists will devote their careers to Precision Medicine data analysis; (6) A clinical trials future, in which trials will be radically redesigned and streamlined; and (7) the future of animal experimentation, in which we come to understand the limitations and benefits of animal models.
Standards are sometimes touted as the solution to every data integration issue. Unfortunately, the utility of data standards has been undermined by the proliferation of standards, the frequent versioning of data standards, the intellectual property encumbrances placed upon standards, the tendency for standards to increase in complexity over time, and the idiosyncratic ways in which standards are implemented by data managers. Data specifications are related to data standards, but differ in several important ways. Data specifications do not dictate an exact way of describing every type of data; and specifications provide the general form and descriptive features of a well-described data object. This chapter will compare the two related topics: data standards and data specifications, indicating how each has been used in Big Data efforts. Big Data resources may benefit by using a standard specification for data objects and by developing a specific set of protocols for converting well-specified data into a conformed standard format, as needed. The importance of data integration among diverse Big Data resources is discussed.
Data identification is one of the most under-appreciated and least understood issues in data science. Measurements, annotations, properties, and classes of information have no informational meaning unless they are attached to an identifier that distinguishes one data object from all other data objects and that links together all of the information that has been or will be associated with the identified data object. The method of identification and the selection of objects and classes to be identified relates fundamentally to the organizational model of complex data. If the simplifying step of data identification is ignored or implemented improperly, data cannot be shared, and conclusions drawn from the data cannot be believed. All well-designed information systems are, at their heart, identification systems: ways of naming data objects so that they can be retrieved. The concept of data identification is of such overriding importance that simplified data sets should be envisioned as collections of unique identifiers to which data is attached. Once data objects have been properly identified, they can be deidentified and, under some circumstances, reidentified. The ability to deidentify data objects confers enormous advantages when issues of confidentiality, privacy, and intellectual property emerge. Data reidentification will be discussed in the context of error detection, error correction, and data validation. This chapter discusses methods for identifying data and explains the catastrophic consequences of inadequate identification.
The convergence of three historically new conditions have greatly increased the utility of data repurposing projects. These are: (i) the availability of new computational algorithms that can be applied to preexisting data to produce new findings that could not have been discovered when the data was originally collected; (ii) the availability of massive data sets produced by the aggregation of multiple preexisting data sets, allowing data analysts to test the assertions drawn from earlier, smaller sets of data, and (iii) the availability of large, complex sets of data, collected and organized from diverse preexisting data sources, that permit us to find relationships and solve problems that could not be addressed by analyzing data sources that are restricted to one type of data. This chapter provides many examples of data repurposing projects that could not have been undertaken in 10 or 20 years ago.
Rare diseases are biologically different from common diseases. They occur in a younger patient population, display a Mendelian pattern of inheritance, typically arise as multi-organ syndromes, and there are many more rare diseases than there are common diseases. From these simple observations, we can infer a great deal about the underlying biology of both the rare diseases and the common diseases.
Data integration occurs when a query proceeds through multiple data sets, thereby relating diverse data extracted from different data sources. Data integration is particularly important to biomedical researchers since data obtained from experiments on human tissue specimens have little applied value unless they can be combined with medical data (i.e., pathologic and clinical information). In the past, research data were correlated with medical data by manually retrieving, reading, assembling and abstracting patient charts, pathology reports, radiology reports and the results of special tests and procedures. Manual annotation of research data is impractical when experiments involve hundreds or thousands of tissue specimens resulting in large, complex data collections. The purpose of this paper is to review how XML (eXtensible Markup Language) provides the fundamental tools that support biomedical data integration. The article also discusses some of the most important challenges that block the widespread availability of annotated biomedical data sets.