A collection of 53 wheat varieties’ root-growth trajectories was clustered into seven distinct morphological growth patterns. Across more than 30K DNA variants, we used Categorical Exploratory Data Analysis (CEDA) to examine all 21 pairwise contrasts between these patterns. Short single- and multi-locus genotype combinations were evaluated by entropy reduction and a finite-sample reliability check based on alternative and null ensembles of contingency tables. Bipartite networks between wheat varieties and retained genotype combinations were displayed as presence–absence heatmaps to reveal pair-specific block structures. As an out-of-pair negative control, the genotype combinations derived from each focal pair were examined, without feature reselection, in varieties belonging to the other five phenotype groups. These excluded groups did not reproduce the focal-pair block structure, supporting the pair specificity of the displayed information within the analyzed dataset. For methodological context, the same samples, markers, and phenotypes were also analyzed with mixed-model association, interaction-detection, and machine-learning methods. Among the four recurrent CEDA markers, marker-level support was strongest for RAC875_c60191_947, which had a FarmCPU PC1 Benjamini–Hochberg-adjusted value of $$q=0.0275$$ and ranked ninth by random-forest importance in the same dataset. These comparisons characterized complementary, method-specific outputs. CEDA thereby characterized genotype-pattern information associated with MRL trajectories in this wheat panel. Across the 21 pairwise contrasts, recurring category-associated genotypes highlighted three putative positional or functional candidates at two loci.
{Categorical Exploratory Data Analysis (CEDA) is used to uncover evolving risk dynamics in chronic disease across heterogeneous subpopulations.} Consider the bivariate dynamics of (Stroke(STK), heart-disease(HD)) to better represent the whole chronic disease in U.S. society. Our data-driven study analyzed the Year 2015 Kaggle-BRFSS database. The evolving disease-risk dynamics across 5 x 5 subpopulation-map defined via (age, general-health(GenHL)) heterogeneity-axis are precisely computed and graphically displayed. Due to the entirely categorical nature of this database, the computational paradigm called Categorical Exploratory Data Analysis (CEDA) is particularly suitable. Each positive or negative chronic disease risk is expressed as a categorical conditional distribution of (STK,HD) given a covariate feature-category presented by a binary vector of absence-presence statuses across subjects within a sub-population. Its finite sample confirmation via a minimum sum of Type-I and Type-II errors is evaluated as an overlapping area of the alternative and null entropy distributions. A subpopulation specific collective of all confirmed chronic disease risks constitutes a binary subject-vs-disease-risk bipartite network called heatmap. Five age-specific heatmap-series along the GenHL axis demonstrated evolving chronic disease risk dynamics with fundamental distinctions. The Age3 heatmap-series across the 5 GenHL stages assumed a pivotal role in societal chronic disease dynamics. As byproducts all predictive tasks are carried out based on the multiscale topological structures defined on the subject-axis of heatmaps.
Accurate detection of R-peaks in electrocardiogram (ECG) signals is essential for cardiac analysis and diagnosis. While existing algorithms often rely on complex transformations or machine learning, we revisit the classical philosophy of using simple and interpretable signal features. We propose a method based on two physiologically meaningful signal features: detrended relative height and local slope pattern. These features are integrated into a robust scoring system with adaptive thresholding. Evaluated on the MIT-BIH Arrhythmia Database, our algorithm uses only a single channel and achieves high sensitivity (99.69%) and specificity (99.82%), performing well within the class of interpretable and transparent methods. The results demonstrate that lightweight, transparent approaches can provide reliable performance and serve as effective baselines for further development.
We implement an analytic approach for ordinal measures and we use it to investigate the structure and the changes over time of self-worth in a sample of adolescents students in high school. We represent the variations in self-worth and its various sub-domains using entropy-based measures that capture the observed uncertainty. We then study the evolution of the entropy across four time points throughout a semester of high school. Our analytic approach yields information about the configuration of the various dimensions of the self together with time-related changes and associations among these dimensions. We represent the results using a network that depicts self-worth changes over time. This approach also identifies groups of adolescent students who show different patterns of associations, thus emphasizing the need to consider heterogeneity in the data.
We investigate the dynamic characteristics of Covid-19 daily infection rates in Taiwan during its initial surge period, focusing on 79 districts within the seven largest cities. By employing computational techniques, we extract 18 features from each district-specific curve, transforming unstructured data into structured data. Our analysis reveals distinct patterns of asymmetric growth and decline among the curves. Utilizing theoretical information measurements such as conditional entropy and mutual information, we identify major factors of order-1 and order-2 that influence the peak value and curvature at the peak of the curves, crucial features characterizing the infection rates. Additionally, we examine the impact of geographic and socioeconomic factors on the curves by encoding each of the 79 districts with two binary characteristics: North-vs-South and Urban-vs-Suburban. Furthermore, leveraging this data-driven understanding at the district level, we explore the fine-scale behavioral effects on disease spread by examining the similarity among 96 age-group-specific curves within urban districts of Taipei and suburban districts of New Taipei City, which collectively represent a substantial portion of the nation's population. Our findings highlight the implicit influence of human behaviors related to living, traveling, and working on the dynamics of Covid-19 transmission in Taiwan.
The Entropy-based Categorical Exploratory Data Analysis (CEDA) paradigm is elaborately refined to algorithmically explore the intricate high-order directional associative relational patterns within the heterogeneous chronical disease dynamics captured by Behavioral Risk Factor Surveillance System (BRFSS) database. Operating on this imbalanced categorical dataset represented fully by its metric-free high-dimensional histogram, our algorithms conduct data-driven computations to investigate chronic disease mechanisms across four sub-populations along the age-axis, culminating in comprehensive systemic understandings. Upon this categorical data-world, CEDA first recognizes the category-oriented 1D histogram as the simplest form of a piece of explainable information. Then, utilizing Kolmogorov's randomness-proper-based reliability check, CEDA identifies and confirms collectives of 1D histograms as major feature-categories of varying orders within each sub-population. These confirmed major feature-categories' binary memberships are then arranged into a subject-vs-feature-category bipartite network heatmap, revealing serial horizontal and vertical blocks framed by clusters of similar subjects characterized by individual-risk-landscapes (IRL) against clusters of structurally dependent major feature-categories. Based on such block-series, sub-population-specific disease mechanisms emerge as collective high-order interacting effects, elucidating directional associative relationships from study subjects' topological neighborhoods to response-categories. Notably, the topological individual-risk-landscape offers profound insights into complex system dynamics and simultaneously exposes atypical subjects as explainable errors across all Machine Learning classifiers.
Without imposing prior distributional knowledge underlying multivariate time series of interest, we propose a nonparametric change-point detection approach to estimate the number of change points and their locations along the temporal axis. We develop a structural subsampling procedure such that the observations are encoded into multiple sequences of Bernoulli variables. A maximum likelihood approach in conjunction with a newly developed searching algorithm is implemented to detect change points on each Bernoulli process separately. Then, aggregation statistics are proposed to collectively synthesize change-point results from all individual univariate time series into consistent and stable location estimations. We also study a weighting strategy to measure the degree of relevance for different subsampled groups. Simulation studies are conducted and shown that the proposed change-point methodology for multivariate time series has favorable performance comparing with currently available state-of-the-art nonparametric methods under various settings with different degrees of complexity. Real data analyses are finally performed on categorical, ordinal, and continuous time series taken from fields of genetics, climate, and finance.
Data analysis is a scientific endeavor of bottom-up data-driven engineering nature. This nature requires all employed conceptual criteria and algorithmic computations equipped with scientific interpretability. It must be free from top-down modeling via man-made structures and assumptions. We demonstrate data analysis of such nature on a critical disease in the real world. In the context of the Alzheimer’s Disease Neuroimaging Initiative (ADNI), we analyze time-to-event data transiting from mild cognitive impairment (MCI) to Alzheimer’s disease (AD) diagnosis. We first address issues related to non-informative censoring using conditional-vs-marginal entropies and the Redistribute-to-the-right algorithm. By employing Categorical Exploratory Data Analysis (CEDA) with 16 covariate variables, we identify a set of key factors, including the Mean of Composite Cognitive Score for Memory (V9) and 13-item-AD Assessment Scale-Cognitive Subscale at baseline (V8). For comparison purposes, this heavily censored data set is also analyzed using Cox’s proportional hazard (PH) modeling and partial likelihood-based approach. Due to complicated structural dependency among covariate features on a global scale, important factors, like V8, are missed in PH results. To further compare PH and CEDA results on locality scales, we subdivide the entire collection of 903 subjects respectively with respect to the four categories of V9 and V8 as a measure of handling induced heterogeneity. Through graphic displays featured with conditional entropy expansions, CEDA is seen to uncover and select more multi-scale informative feature-factors than PH results in all 8 sub-collections when accommodating covariate’s structural dependencies and heterogeneity.
The unknown multiscale structure hidden in large complex systems is explored bottom-up through discovered heterogeneity under structural dependency embedded within structured data sets. Via two real complex systems, we demonstrate computed hierarchical structures with broken symmetry constituting data’s information content. Through graphic displays, such information content indirectly, but efficiently resolves system-related scientific issues that are difficult to resolve directly. All bottom-up explorations and computations are based on conditional entropy and mutual information evaluated upon contingency table platforms after categorizing all quantitative features. Categorical Exploratory Data Analysis (CEDA) first extracts global major factors that share significant mutual information with the targeted response (Re) variable against many covariate (Co) features under the presence of structural dependency. Then each global major factor is taken as one perspective of heterogeneity to subdivide the entire data set according to its categories into sub-collections. This simple “de-associating” protocol significantly reduces structural dependency among the rest of the features such that another run of major factor selection performed on the sub-collection scale can precisely identify which feature sets could provide extra information beyond the global major factor. Finally, informative patterns collected from multiple perspectives of heterogeneity are displayed to explicitly resolve issues of prediction, classification, and detecting minute dynamic changes.
Individual subjects' ratings neither are metric nor have homogeneous meanings, consequently digital- labeled collections of subjects' ratings are intrinsically ordinal and categorical. However, in these situations, the literature privileges the use of measures conceived for numerical data. In this paper, we discuss the exploratory theme of employing conditional entropy to measure degrees of uncertainty in responding to self-rating questions and that of displaying the computed entropies along the ordinal axis for visible pattern recognition. We apply this theme to the study of an online dataset, which contains responses to the Rosenberg Self-Esteem Scale. We report three major findings. First, at the fine scale level, the resultant multiple ordinal-display of response-vs-covariate entropy measures reveals that the subjects on both extreme labels (high self-esteem and low self-esteem) show distinct degrees of uncertainty. Secondly, at the global scale level, in responding to positively posed questions, the degree of uncertainty decreases for increasing levels of self-esteem, while, in responding to negative questions, the degree of uncertainty increases. Thirdly, such entropy-based computed patterns are preserved across age groups. We provide a set of tools developed in R that are ready to implement for the analysis of rating data and for exploring pattern-based knowledge in related research.
We develop a computational protocol for mimicking personal gait dynamics with 12-dimensional time series derived from 4 accelerometer sensors found in the MAREA database and then explore its utilities in line with precision learning of human activities. The foundation of mimicking high dimensional rhythmic dynamics is explicitly established upon deterministic and stochastic structures found on structural representations of evolving biomechanical states hidden within all computed gait cycles. Such a technique enables practitioners to detect and confirm minute structural changes that could last for only a few cycles with high precision. Our computational developments are step-by-step illustrated via one subject’s data, while the other 8 subjects’ data are also analyzed and compared accordingly. A common cyclic composition of evolving biomechanical states of various temporal scales emerges from the 9 subjects’ comparisons. We conclude that mimicking an individual’s gait dynamics offers precise detections of potential multiscale minute differences against gait dynamics of different time periods or of different persons, and further offers clues of efficiency on personal walking activity. This mimicking-based capability is a cornerstone for the proof of concept: dynamics mimicking enables precision learning by improving the efficiency of learning and performing human activities in competitive sports, social dancing, and physical rehabilitation, among many others.
Purpose: The objective of this article is to review fun-damental differences between model-dependent and mod-el-free approaches to data analysis, and to explore the potential advantages of more open-ended machine learn-ing approaches in recovering complex behavioral patterns from precision livestock farming data streams.Sources: Case studies using simulated data were de-signed to mimic a real-world scenario. Data from a feeding trial in an organic dairy were reanalyzed using the Live-stock Informatics Toolkit.Synthesis: Case studies using simulated data are used to demonstrate how incomplete information about the management system can prohibit the development of an appropriate model for information compression, allow-ing aggregation bias to mask important behavioral indi-cators of compromised welfare. These hidden behavioral patterns are then recovered using unsupervised machine learning approaches that are able to leverage the intrinsic behavioral codependencies of group-housed animals. This simulated case study is then extended to demonstrate how model-based approaches can also overlook causes of com-promised welfare when the link between environmental factors and behavioral responses is strong but nonlinear, whereas model-free information-theoretic tools can easily recover and characterize such complex dynamics. Finally, in an empirical case study with data from a commercial organic dairy, the Livestock Informatics Toolkit is used to recover from milk parlor metadata complex associations between herd age structure, levels of milk production, and order of milking.Conclusions and Applications: Model-free machine learning algorithms provide a more open-ended approach to knowledge discovery that require fewer up-front as-sumptions about the management system. This can yield more comprehensive insights into large precision livestock farming data sets now commonly encountered in on-farm research trials and in applied data auditing scenarios.
Volatility is a measure of uncertainty or risk embedded within a stock's dynamics. Such risk has been received huge amounts of attention from diverse financial researchers. By following the concept of regime-switching model, we proposed a non-parametric approach, named encoding-and-decoding, to discover multiple volatility states embedded within a discrete time series of stock returns. The encoding is performed across the entire span of temporal time points for relatively extreme events with respect to a chosen quantile-based threshold. As such the return time series is transformed into Bernoulli-variable processes. In the decoding phase, we computationally seek for locations of change points via estimations based on a new searching algorithm in conjunction with the information criterion applied on the observed collection of recurrence times upon the binary process. Besides the independence required for building the Geometric distributional likelihood function, the proposed approach can functionally partition the entire return time series into a collection of homogeneous segments without any assumptions of dynamic structure and underlying distributions. In the numerical experiments, our approach is found favorably compared with parametric models like Hidden Markov Model. In the real data applications, we introduce the application of our approach in forecasting stock returns. Finally, volatility dynamic of every single stock of S&P500 is revealed, and a stock network is consequently established to represent dependency relations derived through concurrent volatility states among S&P500.
For a large ensemble of complex systems, a Many-System Problem (MSP) studies how heterogeneity constrains and hides structural mechanisms, and how to uncover and reveal hidden major factors from homogeneous parts. All member systems in an MSP share common governing principles of dynamics, but differ in idiosyncratic characteristics. A typical dynamic is found underlying response features with respect to covariate features of quantitative or qualitative data types. Neither all-system-as-one-whole nor individual system-specific functional structures are assumed in such response-vs-covariate (Re–Co) dynamics. We developed a computational protocol for identifying various collections of major factors of various orders underlying Re–Co dynamics. We first demonstrate the immanent effects of heterogeneity among member systems, which constrain compositions of major factors and even hide essential ones. Secondly, we show that fuller collections of major factors are discovered by breaking heterogeneity into many homogeneous parts. This process further realizes Anderson’s “More is Different” phenomenon. We employ the categorical nature of all features and develop a Categorical Exploratory Data Analysis (CEDA)-based major factor selection protocol. Information theoretical measurements—conditional mutual information and entropy—are heavily used in two selection criteria: C1—confirmable and C2—irreplaceable. All conditional entropies are evaluated through contingency tables with algorithmically computed reliability against the finite sample phenomenon. We study one artificially designed MSP and then two real collectives of Major League Baseball (MLB) pitching dynamics with 62 slider pitchers and 199 fastball pitchers, respectively. Finally, our MSP data analyzing techniques are applied to resolve a scientific issue related to the Rosenberg Self-Esteem Scale.
Large and densely sampled sensor datasets can contain a range of complex stochastic structures that are difficult to accommodate in conventional linear models. This can confound attempts to build a more complete picture of an animal’s behavior by aggregating information across multiple asynchronous sensor platforms. The Livestock Informatics Toolkit (LIT) has been developed in R to better facilitate knowledge discovery of complex behavioral patterns across Precision Livestock Farming (PLF) data streams using novel unsupervised machine learning and information theoretic approaches. The utility of this analytical pipeline is demonstrated using data from a 6-month feed trial conducted on a closed herd of 185 mix-parity organic dairy cows. Insights into the tradeoffs between behaviors in time budgets acquired from ear tag accelerometer records were improved by augmenting conventional hierarchical clustering techniques with a novel simulation-based approach designed to mimic the complex error structures of sensor data. These simulations were then repurposed to compress the information in this data stream into robust empirically-determined encodings using a novel pruning algorithm. Nonparametric and semiparametric tests using mutual and pointwise information subsequently revealed complex nonlinear associations between encodings of overall time budgets and the order that cows entered the parlor to be milked.
We reformulate and reframe a series of increasingly complex parametric statistical topics into a framework of response-vs.-covariate (Re-Co) dynamics that is described without any explicit functional structures. Then we resolve these topics’ data analysis tasks by discovering major factors underlying such Re-Co dynamics by only making use of data’s categorical nature. The major factor selection protocol at the heart of Categorical Exploratory Data Analysis (CEDA) paradigm is illustrated and carried out by employing Shannon’s conditional entropy (CE) and mutual information (I[Re;Co]) as the two key Information Theoretical measurements. Through the process of evaluating these two entropy-based measurements and resolving statistical tasks, we acquire several computational guidelines for carrying out the major factor selection protocol in a do-and-learn fashion. Specifically, practical guidelines are established for evaluating CE and I[Re;Co] in accordance with the criterion called [C1:confirmable]. Following the [C1:confirmable] criterion, we make no attempts on acquiring consistent estimations of these theoretical information measurements. All evaluations are carried out on a contingency table platform, upon which the practical guidelines also provide ways of lessening the effects of the curse of dimensionality. We explicitly carry out six examples of Re-Co dynamics, within each of which, several widely extended scenarios are also explored and discussed.
The notion of dominance is ubiquitous across the animal kingdom, wherein some species/groups such relationships are strictly hierarchical and others are not. Modern approaches for measuring dominance have emerged in recent years taking advantage of increased computational power. One such technique, named Percolation and Conductance (Perc), uses both direct and indirect information about the flow of dominance relationships to generate hierarchical rank order that makes no assumptions about the linearity of these relationships. It also provides a new metric, known as 'dominance certainty', which is a complimentary measure to dominance rank that assesses the degree of ambiguity of rank relationships at the individual, dyadic and group levels. In this focused review, we will (i) describe how Perc measures dominance rank while accounting for both nonlinear hierarchical structure as well as sparsity in data-here we also provide a metric of dominance certainty estimated by Perc, which can be used to compliment the information dominance rank supplies; (ii) summarize a series of studies by our research team reflecting the importance of 'dominance certainty' on individual and societal health in large captive rhesus macaque breeding groups; and (iii) provide some concluding remarks and suggestions for future directions for dominance hierarchy research. This article is part of the theme issue 'The centennial of the pecking order: current state and future prospects for the study of dominance hierarchies'.
Electroencephalography (EEG) is a brain imaging approach that has been widely used in neuroscience and clinical settings. The conventional EEG analyses usually require pre-defined frequency bands when characterizing neural oscillations and extracting features for classifying EEG signals. However, neural responses are naturally heterogeneous by showing variations in frequency bands of brainwaves and peak frequencies of oscillatory modes across individuals. Fail to account for such variations might result in information loss and classifiers with low accuracy but high variation across individuals. To address these issues, we present a systematic time-frequency analysis approach for analyzing scalp EEG signals. In particular, we propose a data-driven method to compute the subject-specific frequency bands for brain oscillations via Hilbert-Huang Transform, lifting the restriction of using fixed frequency bands for all subjects. Then, we propose two novel metrics to quantify the power and frequency aspects of brainwaves represented by sub-signals decomposed from the EEG signals. The effectiveness of the proposed metrics are tested on two scalp EEG datasets and compared with four commonly used features sets extracted from wavelet and Hilbert-Huang Transform. The validation results show that the proposed metrics are more discriminatory than other features leading to accuracies in the range of 94.93% to 99.84%. Besides classification, the proposed metrics show great potential in quantification of neural oscillations and serving as biomarkers in the neuroscience research.
Under any Multiclass Classification (MCC) setting defined by a collection of labeled point-cloud specified by a feature-set, we extract only stochastic partial orderings from all possible triplets of point-cloud without explicitly measuring the three cloud-to-cloud distances. We demonstrate that such a collective of partial ordering can efficiently compute a label embedding tree geometry on the Label-space. This tree in turn gives rise to a predictive graph, or a network with precisely weighted linkages. Such two multiscale geometries are taken as the coarse scale information content of MCC. They indeed jointly shed lights on explainable knowledge on why and how labeling comes about and facilitates error-free prediction with potential multiple candidate labels supported by data. For revealing within-label heterogeneity, we further undergo labeling naturally found clusters within each point-cloud, and likewise derive multiscale geometry as its fine-scale information content contained in data. This fine-scale endeavor shows that our computational proposal is indeed scalable to a MCC setting having a large label-space. Overall the computed multiscale collective of data-driven patterns and knowledge will serve as a basis for constructing visible and explainable subject matter intelligence regarding the system of interest.
We develop Categorical Exploratory Data Analysis (CEDA) with mimicking to explore and exhibit the complexity of information content that is contained within any data matrix: categorical, discrete, or continuous. Such complexity is shown through visible and explainable serial multiscale structural dependency with heterogeneity. CEDA is developed upon all features' categorical nature via histogram and it is guided by all features' associative patterns (order-2 dependence) in a mutual conditional entropy matrix. Higher-order structural dependency of k(≥3) features is exhibited through block patterns within heatmaps that are constructed by permuting contingency-kD-lattices of counts. By growing k, the resultant heatmap series contains global and large scales of structural dependency that constitute the data matrix's information content. When involving continuous features, the principal component analysis (PCA) extracts fine-scale information content from each block in the final heatmap. Our mimicking protocol coherently simulates this heatmap series by preserving global-to-fine scales structural dependency. Upon every step of mimicking process, each accepted simulated heatmap is subject to constraints with respect to all of the reliable observed categorical patterns. For reliability and robustness in sciences, CEDA with mimicking enhances data visualization by revealing deterministic and stochastic structures within each scale-specific structural dependency. For inferences in Machine Learning (ML) and Statistics, it clarifies, upon which scales, which covariate feature-groups have major-vs.-minor predictive powers on response features. For the social justice of Artificial Intelligence (AI) products, it checks whether a data matrix incompletely prescribes the targeted system.