For more than a half-century, scientists have been developing a tool for linear unmixing utilizing collections of algorithms and computer programs that is appropriate for many types of data commonly encountered in the geologic and other science disciplines. Applications include the analysis of particle size data, Fourier shape coefficients and related spectrum, biologic morphology and fossil assemblage information, environmental data, petrographic image analysis, unmixing igneous and metamorphic petrographic variable and the unmixing and determination of oil sources, to name a few. Each of these studies used algorithms that were designed to use data whose row sums are constant. Non-constant sum data comprise what is a larger set of data that permeates many of our sciences. Many times, these data can be modeled as mixtures even though the row sums do not sum to the same value for all samples in the data. This occurs when different quantities of one or more end-member are present in the data. Use of the constant sum approach for these data can produce confusing and inaccurate results especially when the end-members need to be defined away from the data cloud. The approach to deal with these non-constant sum data is defined and called Hyperplanar Vector Analysis (HVA). Without abandoning over 50 years of experience, HVA merges the concepts developed over this time and extends the linear unmixing approach to more types of data. The basis for this development involves a translation and rotation of the raw data that conserves information (variability). It will also be shown that HVA is a more appropriate name for both the previous constant sum algorithms and future programs algorithms as well.
The problem exists in which the harmonic(s) have the greatest potential for a clear, unambiguous solution upon subsequent shape analysis. In addition, shape data must be acquired economically, allowing analysis of hundreds of particles per sample and tens or hundreds of samples in any particular investigation. Amplitude spectra of a finite Fourier series in closed form are used as shape descriptors of each particle. The mixing proportions and end member compositions have proven to be of value with respect to a variety of sedimentological investigations using shape and size frequency data. The lower harmonics are a measure of gross shape while the higher harmonics measure increasingly fine-scaled surface features. The width of the intervals can be unequal and depends only on the shape of the distribution of the entire pooled data set. The lower the relative entropy of a data set, the more contrast exists between samples contained in that set.
This chapter provides an overview of quantitative, multivariate exploratory data analysis methods in use in environmental forensics and chemical fingerprinting. Given multiple contaminant sources and empirical chemical data from a study area, environmental scientists often wish to unravel the contributions of multiple sources with overlapping spatial and temporal distributions. Such problems may be addressed through a multivariate approach, such as principal components analysis (PCA) and self-training receptor-modeling techniques. These approaches are related in that they are exploratory data analysis methods—techniques applied when one wishes not to assume a priori knowledge of contributing sources and their fingerprints. PCA is widely used in environmental chemistry as well as many scientific disciplines, and can be implemented using commercial software packages. Four self-training receptor model methods are discussed, compared, and applied herein. There are advantages and disadvantages to each of these methods, but regardless of the method, success depends on the data analyst’s experience, sensitivity to chemical data structure, and sampling plan design. Finally, one must be aware that the compositions of original source profiles may be altered after a release (even for recalcitrant chemicals such as PCBs). In interpreting these models, the data analyst should have some knowledge of alteration, weathering, and degradation mechanisms.
The purpose of the study is to (1) identify the sources of sediment in various environments, (2) define the history and transport processes of the sediments, and (3) better understand the erosion and potential replenishment of the local beaches along the southern coast of the Baja California peninsula. For the purpose of this study, six naturally defined areas were studied separately: El Cardonal, El Arco, San Lucas, El Tiburon, El Tule, and San Jose.Two main sedimentary provinces were identified via Fourier grain-shape analysis, El Medano and Los Cabos. El Medano sedimentary province includes the El Cardonal and El Arco areas, which are influenced by the dynamics associated with the Pacific Ocean dominated by northwesterly winds, waves, and longshore transport. Beaches from this province have a source mostly from marine material from the shallow shelf, and they are dominantly affected by longshore transport. Secondarily, they are dominated by old and recent aeolian material dissected by intermittent arroyos and local arroyo material from intrusive rocks. The Los Cabos sedimentary province includes the other four areas, and it is influenced by the dynamics of the Gulf of California. In this province, dominant southerly waves are present. Sediment transport occurs along the coast from southwest to northeast; although, some beaches contain material from northern areas, probably related to the direction of waves and sediment transport direction during meteoric events such as hurricanes. Beaches from this province have a source mostly from local arroyo material from intrusive rocks. Other beach material results from longshore transport and some material comes from the El Medano sedimentary province in the El Arco boundary area.Grain-shape data and the information associated with elongation (harmonic 2) show that marine samples (beach, shallow, and deep inner continental shelf) from Los Medanos sedimentary province contain high frequencies of grains with low elongation, opposite of the arroyo samples. This suggests that the low elongation grain source may be farther north of this province. In the Los Cabos sedimentary province, the local arroyos and the longshore transport have been identified as the major factors that nourish and distribute the beach material along the coast. The results of this study parallel those found in similar geographic regions where storms rather than steady currents dominant.
A computer model was generated that relates the waves impacting a beach in the western part of Central Italy to the longshore sediment transport and coastal features from S. Felice Circeo on the west to Terracina on the east. The model analyzed spectra of wave properties as they underwent transformations shoreward of the 100 m isobath. Longshore sediment transport was calculated at 5 profile locations along the 15.9 km beach. Most of the waves with potential for sediment transport have little net effect on annual sediment transport as they generally maintain sediment equilibrium. The model predicted sand erosion from the western portion of the study area and suggested sediment bypassing along the eastern end. Although onshore and offshore sediment transport was not modeled, inferences based on bottom shear stress and wave dynamics suggest potential onsho- re sediment transport at the western end of the study area.
The wide range of morphologies seen among taxa, as well as intraspecific variability, when combined with the diversity of information carried by morphology, ensures that no “one size fits all” method exists for morphologic characterization. In some cases, interest resides in determining how a single species' form has changed in response to environmental parameters; the organism being used solely as a proxy indicator. In other cases, shape components linked to such environmental factors will be excluded in order to determine the nature and rates of change of the genome as reflected in morphology. In still other cases, the morphology may be of interest in terms of some sort of functional efficiency with respect to locomotion or musculature.
Fourier shape analysis has been applied to quartz sand grains collected along the coastal regions around Port Stephens (150 km north of Sydney), with the intent to determine the dispersal patterns of the fluvial sand along this coast. By using a video-digitizing micro-processor controlled computer system, the outline of the quartz particles were defined in polar coordinates and a Fourier series in closed form was calculated. Twenty-four harmonics were obtained, each representing the contribution of elementary forms to the real shape of the grain. Selected distributions of harmonic amplitude were defined utilizing the maximum entropy method and chi-square statistic. The combined 15-24th harmonic amplitude distributions were submitted to unmixing analysis, allowing the definition of three main end-members to be made, each one related to smooth, angular and intermediate shaped grains. The proportion of these end-members in each sample has suggested the provenance and dispersion patterns for the study system. Three major sand types have been identified: flood reworked coastal plain sands, related to a mixture of river and mature coastal plain particles; fluvial-marine coastal sands, related to a sand mixture of fluvial and marine origin, found close to the area of the Hunter River mouth, and marine coastal sands in the northern portion of the study area, which show higher maturity than the other two types of sediment. Sands associated with headland erosion display shape characteristics similar to those of the fluvial material. The shape data suggests that the Hunter River supplies sediment at least in the southwest sector of the Newcastle Bight modern beach system, probably during higher discharge flood events. Additionally, this study indicates that textural characteristics of the river sediment are subjected to rapid changes in the coastal active zone, and hence generally make non-definitive indicators of source and dispersion of the sediment.
Fine-silt to medium-sand grain-size spectra of the carbonate-free fractions of 115 sea bed samples from the northern North Atlantic, the Labrador Sea and Baffin Bay can be explained as mixtures of five distinct end-members. The five end-members resemble combinations of grain-size modes which are characteristic of till-matrix (Dreimanis and Vagners, 1972). Therefore the ultimate source of most of the terrigenous deep-sea sediments in the study area is probably the veneer of glacial comminution products on the surrounding continents.
ABSTRACT Data sets are often analyzed in the form of collections of frequency tables (or percentiles derived from equivalent cumulative frequency distributions). Decisions concerning the number of intervals and interval width obviously affect the quality of the data in subsequent analysis. Relying on the basic concepts of information theory, a procedure is presented which evaluates the relative information content of a set of frequency data when subdivided in various manners. Maximum information is always preserved when “maximum entropy” histograms (with unequal class intervals) are used. Evaluation of several schemes of frequency table subdivision (phi-based arithmetic, log arithmetic, Z-score, log Z-score, maximum entropy) indicates that, surprisingly, collections of equal interval phi-based frequency tables contain the least information. Additionally, the concept of the relative entropy of a given collection of frequency tables is defined. The relative entropy is useful as a feature extractor wherein s...
Many data sets can be viewed as a collection of samples representing mixtures of a relatively small number of end members. When end members are present in the sample set, the algorithm QMODEL by Klovan and Miesch can efficiently determine proportionate contributions. EXTENDED QMODEL by Full, Ehrlich, and Klovan was designed to deduce the composition of realistic end members when the end members are not represented by samples. However, in the presence of high levels of random variation or outliers not belonging to the system of interest, EXTENDED QMODEL may not be reliable inasmuch as it is largely dependent on extreme values for definition of an initial mixing polyhedron. FUZZY QMODEL utilizes the fuzzy c-means algorithm of Bezdek to provide an alternative initial mixing polyhedron. This algorithm utilizes the collective property of all the data rather than outliers and so can produce suitable solutions in the presence of noisy or “messy” data points.
The ability to test for similarities and differences among families of shapes by closed-form Fourier expansion is greatly enhanced by the concept of homology. Underlying this concept is the assumption that each term of a Fourier series, when compared to the same term in another series, represents the “same thing”. A method that ensures homology is one which minimizes the “centering error,” as reflected in the first harmonic term of the Fourier expansion. The problem is to chose a set of edge points derived from a much larger, but variable, number of edge points such that a valid homologous Fourier series can be calculated. Methods are reviewed and criteria given to define a “proper” solution. An algorithm is presented which takes advantage of the fact that minimization of the “error term” can be accomplished by minimizing the distance between the origin of the polar coordinate system in the calculation of the Fourier series and the shape centroid. The use of this algorithm has produced higher quality solutions for quartz grain provenance studies.
Analysis of empirical data considered to be mixtures of a finite number of end members has been a topic of increasing interest recently. The algorithms EXTENDED CABFAC and QMODEL by Klovan and Miesch (1976) represent a satisfactory solution to this problem if pure end members are captured within the data set or if the composition of “true” end members are known a priori. Where neither condition is satisfied, the composition of “external” end members can, under certain conditions, be deduced from the structure of the data. Described herein is an algorithm termed EXTENDED QMODEL which defines feasible end members which are “closest” to the data envelope.