Nonlinear science has evolved significantly over the 35 years since the launch of the journal Chaos. This Focus Issue, dedicated to the 80th Birthday of its founding editor-in-chief, David K. Campbell, brings together a selection of contributions on influential topics, many of which were advanced by Campbell's own research program and leadership role. The topics include new phenomena and method developments in the realms of network dynamics, machine learning, quantum and material systems, chaos and fractals, localized states, and living systems, with a good balance of literature review, original contributions, and perspectives for future research.
The acceleration in the field of data science is well known [see, e.g., D. Donoho, J. Comput. Graph. Stat. 26(4), 745-766 (2017) and references therein]. Improvements in technology for acquisition, storage, and processing have made unheard of amounts of data available to scientists; in parallel with that, the pace of methodological advance has also been rapid; with new techniques and packages becoming available, it seems, every day. With these affordances come many challenges, notably the volume and variety of the data [Fan et al., Natl. Sci. Rev. 1(2), 293-314 (2014)]. In this Perspective piece, we examine a different challenge-how to choose and use the right analysis method-and make an argument for the sharing of raw data.
A primary challenge in understanding collective behavior is characterizing the spatiotemporal dynamics of the group. We employ topological data analysis to explore the structure of honeybee aggregations that form during trophallaxis, which is the direct exchange of food among nestmates. From the positions of individual bees, we build topological summaries called CROCKER matrices to track the morphology of the group as a function of scale and time. Each column of a CROCKER matrix records the number of topological features, such as the number of components or holes, that exist in the data for a range of analysis scales at a given point in time. To detect important changes in the morphology of the group from this information, we first apply dimensionality reduction techniques to these matrices and then use classic clustering and change-point detection algorithms on the resulting scalar data. A test of this methodology on synthetic data from an agent-based model of honeybees and their trophallaxis behavior shows two distinct phases: a dispersed phase that occurs before food is introduced, followed by a food-exchange phase during which aggregations form. We then move to laboratory data, successfully detecting the same two phases across multiple experiments. Interestingly, our method reveals an additional phase change towards the end of the experiments, suggesting the possibility of another dispersed phase that follows the food-exchange phase.
The Madden–Julian Oscillation (MJO) is a dominant source of subseasonal atmospheric variability in the tropics and significantly impacts global weather and climate predictability. Changes in its activity and predictability due to human-induced global climate change have profound implications for future global weather prediction. Here we investigate changes in MJO predictability in reanalysis and climate model data and find that MJO predictability has increased over the past century. This increase can be attributed to anthropogenic warming and continues during the twenty-first century in projections. The increased predictability is accompanied by stronger MJO amplitude, more regular oscillation patterns and organized eastward propagation under global warming. Our results suggest that greenhouse warming will increase the predictability of the MJO, with far-reaching consequences for global weather prediction.
Context. Machine-learning methods for predicting solar flares typically employ physics-based features that have been carefully chosen by experts in order to capture the salient features of the photospheric magnetic fields of the Sun. Aims. Though the sophistication and complexity of these models have grown over time, there has been little evolution in the choice of feature sets, or any systematic study of whether the additional model complexity leads to higher predictive skill. Methods. This study compares the relative prediction performance of four different machine-learning based flare prediction models with increasing degrees of complexity. It evaluates three different feature sets as input to each model: a “traditional” physics-based feature set, a novel “shape-based” feature set derived from topological data analysis (TDA) of the solar magnetic field, and a combination of these two sets. A systematic hyperparameter tuning framework is employed in order to assure fair comparisons of the models across different feature sets. Finally, principal component analysis is used to study the effects of dimensionality reduction on these feature sets. Results. It is shown that simpler models with fewer free parameters perform better than the more complicated models on the canonical 24-h flare forecasting problem. In other words, more complex machine-learning architectures do not necessarily guarantee better prediction performance. In addition, it is found that shape-based feature sets contain just as much useful information as physics-based feature sets for the purpose of flare prediction, and that the dimension of these feature sets – particularly the shape-based one – can be greatly reduced without impacting predictive accuracy.
In 2022, the National Science Foundation (NSF) funded the Computing Research Association (CRA) to conduct a workshop to frame and scope a potential Convergence Accelerator research track on the topic of "Building Resilience to Climate-Driven Extreme Events with Computing Innovations". The CRA's research visioning committee, the Computing Community Consortium (CCC), took on this task, organizing a two-part community workshop series, beginning with a small, in-person brainstorming meeting in Denver, CO on 27-28 October 2022, followed by a virtual event on 10 November 2022. The overall objective was to develop ideas to facilitate convergence research on this critical topic and encourage collaboration among researchers across disciplines. Based on the CCC community white paper entitled Computing Research for the Climate Crisis, we initially focused on five impact areas (i.e. application domains that are both important to society and critically affected by climate change): Energy, Agriculture, Environmental Justice, Transportation, and Physical Infrastructure.
Reconstructing state-space dynamics from scalar data using time-delay embedding requires choosing values for the delay r and the dimension m. Both parameters are critical to the success of the procedure and neither is easy to formally validate. While embedding theorems do offer formal guidance for these choices, in practice one has to resort to heuristics, such as the average mutual information (AMI) method of Fraser & Swinney for r or the false near neighbor (FNN) method of Kennel et al. for m. Best practice suggests an iterative approach: one of these heuristics is used to make a good first guess for the corresponding free parameter and then an "asymptotic invariant"approach is then used to firm up its value by, e.g., computing the correlation dimension or Lyapunov exponent for a range of values and looking for convergence. This process can be subjective, as these computations often involve finding, and fitting a line to, a scaling region in a plot: a process that is generally done by eye and is not immune to confirmation bias. Moreover, most of these heuristics do not provide confidence intervals, making it difficult to say what "convergence"is. Here, we propose an approach that automates the first step, removing the subjectivity, and formalizes the second, offering a statistical test for convergence. Our approach rests upon a recently developed method for automated scaling-region selection that includes confidence intervals on the results. We demonstrate this methodology by selecting values for the embedding dimension for several real and simulated dynamical systems. We compare these results to those produced by FNN and validate them against known results-e.g., of the correlation dimension-where these are available. We note that this method extends to any free parameter in the theory or practice of delay reconstruction.(c) 2023 Elsevier B.V. All rights reserved.
Recently, there has been growing interest in the use of machine-learning methods for predicting solar flares. Initial efforts along these lines employed comparatively simple models, correlating features extracted from observations of sunspot active regions with known instances of flaring. Typically, these models have used physics-inspired features that have been carefully chosen by experts in order to capture the salient features of such magnetic field structures. Over time, the sophistication and complexity of the models involved has grown. However, there has been little evolution in the choice of feature sets, nor any systematic study of whether the additional model complexity is truly useful. Our goal is to address these issues. To that end, we compare the relative prediction performance of machine-learning-based, flare-forecasting models with varying degrees of complexity. We also revisit the feature set design, using topological data analysis to extract shape-based features from magnetic field images of the active regions. Using hyperparameter training for fair comparison of different machinelearning models across different feature sets, we show that simpler models with fewer free parameters generally perform better than more-complicated models, ie., powerful machinery does not necessarily guarantee better prediction performance. Secondly, we find that abstract, shape-based features contain just as much useful information, for the purposes of flare prediction, as the set of hand-crafted features developed by the solar-physics community over the years. Finally, we study the effects of dimensionality reduction, using principal component analysis, to show that streamlined feature sets, overall, perform just as well as the corresponding full-dimensional versions.
Climate change is an existential threat to the United States and the world. Inevitably, computing will play a key role in mitigation, adaptation, and resilience in response to this threat. The needs span all areas of computing, from devices and architectures (e.g., low-power sensor systems for wildfire monitoring) to algorithms (e.g., predicting impacts and evaluating mitigation), and robotics (e.g., autonomous UAVs for monitoring and actuation) -- as well as every level of the software stack, from data management systems and energy-aware operating systems to hardware/software co-design. The goal of this white paper is to highlight the role of computing research in addressing climate change-induced challenges. To that end, we outline six key impact areas in which these challenges will arise -- energy, environmental justice, transportation, infrastructure, agriculture, and environmental monitoring and forecasting -- then identify specific ways in which computing research can help address the associated problems. These impact areas will create a driving force behind, and enable, cross-cutting, system-level innovation. We further break down this information into four broad areas of computing research: devices&architectures, software, algorithms/AI/robotics, and sociotechnical computing. Additional contributions by: Ilkay Altintas (San Diego Supercomputer Center), Kyri Baker (University of Colorado Boulder), Sujata Banerjee (VMware), Andrew A. Chien (University of Chicago), Thomas Dietterich (Oregon State University), Ian Foster (Argonne National Labs), Carla P. Gomes (Cornell University), Chandra Krintz (University of California, Santa Barbara), Jessica Seddon (World Resources Institute), and Regan Zane (Utah State University).
Mixing of neighboring data points in a sequence is a common, but understudied, effect in physical experiments. This can occur in the measurement apparatus (if material from multiple time points is pulled into a measurement chamber simultaneously, for instance) or the system itself, e.g., via diffusion of isotopes in an ice sheet. We propose a model-free technique to detect this kind of local mixing in time-series data using an information-theoretic technique called permutation entropy. By varying the temporal resolution of the calculation and analyzing the patterns in the results, we can determine whether the data are mixed locally, and on what scale. This can be used by practitioners to choose appropriate lower bounds on scales at which to measure or report data. After validating this technique on several synthetic examples, we demonstrate its effectiveness on data from a chemistry experiment, methane records from Mauna Loa, and an Antarctic ice core.
Infectious diseases cause more than 13 million deaths a year, worldwide. Globalization, urbanization, climate change, and ecological pressures have significantly increased the risk of a global pandemic. The ongoing COVID-19 pandemic-the first since the H1N1 outbreak more than a decade ago and the worst since the 1918 influenza pandemic-illustrates these matters vividly. More than 47M confirmed infections and 1M deaths have been reported worldwide as of November 4, 2020 and the global markets have lost trillions of dollars. The pandemic will continue to have significant disruptive impacts upon the United States and the world for years; its secondary and tertiary impacts might be felt for more than a decade. An effective strategy to reduce the national and global burden of pandemics must: 1) detect timing and location of occurrence, taking into account the many interdependent driving factors; 2) anticipate public reaction to an outbreak, including panic behaviors that obstruct responders and spread contagion; 3) and develop actionable policies that enable targeted and effective responses.
The increase in variable renewable generators (VRGs) in power systems has altered the dynamics from a historical experience. VRGs introduce new sources of power oscillations, and the stabilizing response provided by synchronous generators (SGs, e.g., natural gas, coal, etc.), which help avoid some power fluctuations, will lessen as VRGs replace SGs. These changes have led to the need for new methods and metrics to quickly assess the likely oscillatory behavior for a particular network without performing computationally expensive simulations. This work studies the impact of a critical dynamical parameter—the inertia value—on the rest of a power system’s oscillatory response to representative VRG perturbations. We use a known localization metric in a novel way to quantify the number of nodes responding to a perturbation and the magnitude of those responses. This metric allows us to relate the spread and severity of a system’s power oscillations with inertia. We find that as inertia increases, the system response to node perturbations transitions from localized (only a few close nodes respond) to delocalized (many nodes across the network respond). We introduce a heuristic computed from the network Laplacian to relate this oscillatory transition to the network structure. We show that our heuristic accurately describes the spread of oscillations for a realistic power-system test case. Using a heuristic to determine the likely oscillatory behavior of a system given a set of parameters has wide applicability in power systems, and it could decrease the computational workload of planning and operation.
Solar flares are caused by magnetic eruptions in active regions (ARs) on the surface of the sun. These events can have significant impacts on human activity, many of which can be mitigated with enough advance warning from good forecasts. To date, machine learning-based flare-prediction methods have employed physics-based attributes of the AR images as features; more recently, there has been some work that uses features deduced automatically by deep learning methods (such as convolutional neural networks). We describe a suite of novel shape-based features extracted from magnetogram images of the Sun using the tools of computational topology and computational geometry. We evaluate these features in the context of a multi-layer perceptron (MLP) neural network and compare their performance against the traditional physics-based attributes. We show that these abstract shape-based features outperform the features chosen by the human experts, and that a combination of the two feature sets improves the forecasting capability even further.
Scaling regions-intervals on a graph where the dependent variable depends linearly on the independent variable-abound in dynamical systems, notably in calculations of invariants like the correlation dimension or a Lyapunov exponent. In these applications, scaling regions are generally selected by hand, a process that is subjective and often challenging due to noise, algorithmic effects, and confirmation bias. In this paper, we propose an automated technique for extracting and characterizing such regions. Starting with a two-dimensional plot-e.g., the values of the correlation integral, calculated using the Grassberger-Procaccia algorithm over a range of scales-we create an ensemble of intervals by considering all possible combinations of end points, generating a distribution of slopes from least squares fits weighted by the length of the fitting line and the inverse square of the fit error. The mode of this distribution gives an estimate of the slope of the scaling region (if it exists). The end points of the intervals that correspond to the mode provide an estimate for the extent of that region. When there is no scaling region, the distributions will be wide and the resulting error estimates for the slope will be large. We demonstrate this method for computations of dimension and Lyapunov exponent for several dynamical systems and show that it can be useful in selecting values for the parameters in time-delay reconstructions.
Trophallaxis is the mutual exchange and direct transfer of liquid food among eusocial insects such as ants, termites, wasps, and bees. This process allows efficient dissemination of nutrients and is crucial for the colony’s survival. In this paper, we present a data-driven agent-based model and use it to explore how the interactions of individual bees, following simple, local rules, affect the global food distribution. We design the rules in our model using laboratory experiments on honeybees. We validate its results via comparisons with the movement patterns in real bees. Using this model, we demonstrate that the efficiency of food distribution is affected by the density of the individuals, as well as the rules that govern their behavior: e.g., how they move and whether or not they aggregate. Specifically, food is distributed more efficiently when donor bees do not always feed their immediate neighbors, but instead prioritize longer motions, sharing their food with more-distant bees. This non-local pattern of food exchange enhances the overall probability that all of the bees, regardless of their position in the colony, will be fed efficiently. We also find that short-range attraction improves the efficiency of the food distribution in the simulations. Importantly, this model makes testable predictions about the effects of different bee densities, which can be validated in experiments. These findings can potentially contribute to the design of local rules for resource sharing in swarm robotic systems.
Current operational forecasts of solar eruptions are made by human experts using a combination of qualitative shape-based classification systems and historical data about flaring frequencies. In the past decade, there has been a great deal of interest in crafting machine-learning (ML) flare-prediction methods to extract underlying patterns from a training set – e.g. a set of solar magnetogram images, each characterized by features derived from the magnetic field and labeled as to whether it was an eruption precursor. These patterns, captured by various methods (neural nets, support vector machines, etc.), can then be used to classify new images. A major challenge with any ML method is the featurization of the data: pre-processing the raw images to extract higher-level properties, such as characteristics of the magnetic field, that can streamline the training and use of these methods. It is key to choose features that are informative, from the standpoint of the task at hand. To date, the majority of ML-based solar eruption methods have used physics-based magnetic and electric field features such as the total unsigned magnetic flux, the gradients of the fields, the vertical current density, etc. In this paper, we extend the relevant feature set to include characteristics of the magnetic field that are based purely on the geometry and topology of 2D magnetogram images and show that this improves the prediction accuracy of a neural-net based flare-prediction method.
Infectious diseases cause more than 13 million deaths a year, worldwide. Globalization, urbanization, climate change, and ecological pressures have significantly increased the risk of a global pandemic. The ongoing COVID-19 pandemic-the first since the H1N1 outbreak more than a decade ago and the worst since the 1918 influenza pandemic-illustrates these matters vividly. More than 47M confirmed infections and 1M deaths have been reported worldwide as of November 4, 2020 and the global markets have lost trillions of dollars. The pandemic will continue to have significant disruptive impacts upon the United States and the world for years; its secondary and tertiary impacts might be felt for more than a decade. An effective strategy to reduce the national and global burden of pandemics must: 1) detect timing and location of occurrence, taking into account the many interdependent driving factors; 2) anticipate public reaction to an outbreak, including panic behaviors that obstruct responders and spread contagion; 3) and develop actionable policies that enable targeted and effective responses.
Reinhard Stolle合作论文数BMW Car IT GmbH7