We consider a discrete-time continuous-space random walk, with a symmetric jump distribution, under stochastic resetting. Associated with the random walker are cost functions for jumps and resets, and we calculate the distribution of the total cost for the random walker up to the first passage to the target. By using the backward master equation approach we demonstrate that the distribution of the total cost up to the first passage to the target can be reduced to a Wiener-Hopf integral equation. The resulting integral equation can be exactly solved (in Laplace space) for arbitrary cost functions for the jump and selected functions for the reset cost. We show that the large cost behaviour is dominated by resetting or the jump distribution according to the choice of the jump distribution. In the important case of a Laplace jump distribution, which corresponds to run-and-tumble particle dynamics, and linear costs for jumps and resetting, the Wiener-Hopf equation simplifies to a differential equation which can easily be solved.
Language change is a cultural evolutionary process in which variants of linguistic variables change in frequency through processes analogous to mutation, selection and genetic drift. In this work, we apply a recently-introduced method to corpus data to quantify the strength of selection in specific instances of historical language change. We first demonstrate, in the context of English irregular verbs, that this method is more reliable and interpretable than similar methods that have previously been applied. We further extend this study to demonstrate that a bias towards phonological simplicity overrides that favouring grammatical simplicity when these are in conflict. Finally, with reference to Spanish spelling reforms, we show that the method can also detect points in time at which selection strengths change, a feature that is generically expected for socially-motivated language change. Together, these results indicate how hypotheses for mechanisms of language change can be tested quantitatively using historical corpus data.
We introduce the profligacy of a search process as a competition between its expected cost and the probability of finding the target. The arbiter of the competition is a parameter λ that represents how much a searcher invests into increasing the chance of success. Minimizing the profligacy with respect to the search strategy specifies the optimal search. We show that in the case of diffusion with stochastic resetting, the amount of resetting in the optimal strategy has a highly nontrivial dependence on model parameters resulting in classical continuous transitions, discontinuous transitions and tricritical points, as well as nonstandard discontinuous transitions exhibiting reentrant behavior and overhangs.
Neural Cellular Automata (NCA) are a powerful combination of machine learning and mechanistic modelling. We train NCA to learn complex dynamics from time series of images and Partial Differential Equation (PDE) trajectories. Our method is designed to identify underlying local rules that govern large scale dynamic emergent behaviours. Previous work on NCA focuses on learning rules that give stationary emergent structures. We extend NCA to capture both transient and stable structures within the same system, as well as learning rules that capture the dynamics of Turing pattern formation in nonlinear PDEs. We demonstrate that NCA can generalise very well beyond their PDE training data, we show how to constrain NCA to respect given symmetries, and we explore the effects of associated hyperparameters on model performance and stability. Being able to learn arbitrary dynamics gives NCA great potential as a data driven modelling framework, especially for modelling biological pattern formation.
Resetting a stochastic process has been shown to expedite the completion time of some complex tasks, such as finding a target for the first time. Here we consider the cost of resetting by associating to each reset a cost, which is a function of the distance travelled during the reset event. We compute the Laplace transform of the joint probability of first passage time tf , number of resets N and total resetting cost C, and use this to study the statistics of the total cost and also the time to completion . We show that in the limit of zero resetting rate, the mean total cost is finite for a linear cost function, vanishes for a sub-linear cost function and diverges for a super-linear cost function. This result contrasts with the case of no resetting where the cost is always zero. We also find that the resetting rate which optimizes the mean time to completion may be increased or decreased with respect to the case of no resetting cost according to the choice of cost function. For the case of an exponentially increasing cost function, we show that the mean total cost diverges at a finite resetting rate. We explain this by showing that the distribution of the cost has a power-law tail with a continuously varying exponent that depends on the resetting rate.
We construct a reliable estimation method for evolutionary parameters within the Wright-Fisher model, which describes changes in allele frequencies due to selection and genetic drift, from time-series data. Such data exist for biological populations, for example via artificial evolution experiments, and for the cultural evolution of behavior, such as linguistic corpora that document historical usage of different words with similar meanings. Our method of analysis builds on a Beta-with-Spikes approximation to the distribution of allele frequencies predicted by the Wright-Fisher model. We introduce a self-contained scheme for estimating parameters in the approximation, and demonstrate its robustness with synthetic data, especially in the strong-selection and near-extinction regimes where previous approaches fail. We further apply the method to allele frequency data for baker's yeast (Saccharomyces cerevisiae), finding a significant signal of selection in cases where independent evidence supports such a conclusion. We further demonstrate the possibility of detecting time points at which evolutionary parameters change in the context of a historical spelling reform in the Spanish language.
We consider a model system of persistent random walkers that can jam, pass through each other, or jump apart (recoil) on contact. In a continuum limit, where particle motion between stochastic changes in direction becomes deterministic, we find that the stationary interparticle distribution functions are governed by an inhomogeneous fourth-order differential equation. Our main focus is on determining the boundary conditions that these distribution functions should satisfy. We find that these do not arise naturally from physical considerations, but they need to be carefully matched to functional forms that arise from the analysis of an underlying discrete process. The interparticle distribution functions, or their first derivatives, are generically found to be discontinuous at the boundaries.
We consider the interplay between persistent motion, which is a generic property of active particles, and a recoil interaction which causes particles to jump apart on contact. The recoil interaction exemplifies an active contact interaction between particles, which is inelastic and is generated by the active nature of the constituents. It is inspired by the “shock” dynamics of certain microorganisms, such as Pyramimonas octopus , and always generates an effective repulsion between a pair of passive particles. Highly persistent particles can be attractive or repulsive, according to the shape of the recoil distribution. We show that the repulsive case admits an unexpected transition to attraction at intermediate persistence lengths, that originates in the advective effects of persistence. This allows active particles to fundamentally change the collective effect of active interactions amongst them, by varying their persistence length.
Background There have been over 30 million cases of COVID-19 in India and over 430,000 deaths. Transmission rates vary from region to region, and are influenced by many factors including population susceptibility, travel and uptake of preventive measures. To date there have been relatively few studies examining the impact of the pandemic in lower income, rural regions of India. We report on a study examining COVID-19 burden in a rural community in Tamil Nadu. Methods The study was undertaken in a population of approximately 130,000 people, served by the Rural Unit of Health and Social Affairs (RUHSA), a community health center of CMC, Vellore. We established and evaluated a COVID-19 PCR-testing programme for symptomatic patients—testing was offered to 350 individuals, and household members of test-positive cases were offered antibody testing. We also undertook two COVID-19 seroprevalence surveys in the same community, amongst 701 randomly-selected individuals. Results There were 182 positive tests in the symptomatic population (52.0%). Factors associated with test-positivity were older age, male gender, higher socioeconomic status (SES, as determined by occupation, education and housing), a history of diabetes, contact with a confirmed/suspected case and attending a gathering (such as a religious ceremony, festival or extended family gathering). Amongst test-positive cases, 3 (1.6%) died and 16 (8.8%) suffered a severe illness. Amongst 129 household contacts 40 (31.0%) tested positive. The two seroprevalence surveys showed positivity rates of 2.2% (July/Aug 2020) and 22.0% (Nov 2020). 40 tested positive (31.0%, 95% CI: 23.02 − 38.98). Our estimated infection-to-case ratio was 31.7. Conclusions A simple approach using community health workers and a community-based testing clinic can readily identify significant numbers of COVID-19 infections in Indian rural population. There appear, however, to be low rates of death and severe illness, although vulnerable groups may be under-represented in our sample. It’s vital these lower income, rural populations aren’t overlooked in ongoing pandemic monitoring and vaccine roll-out in India.
Languages emerge and change over time at the population level though interactions between individual speakers. It is, however, hard to directly observe how a single speaker's linguistic innovation precipitates a population-wide change in the language, and many theoretical proposals exist. We introduce a very general mathematical model that encompasses a wide variety of individual-level linguistic behaviours and provides statistical predictions for the population-level changes that result from them. This model allows us to compare the likelihood of empirically-attested changes in definite and indefinite articles in multiple languages under different assumptions on the way in which individuals learn and use language. We find that accounts of language change that appeal primarily to errors in childhood language acquisition are very weakly supported by the historical data, whereas those that allow speakers to change incrementally across the lifespan are more plausible, particularly when combined with social network effects.
Colexification refers to the phenomenon of multiple meanings sharing one word in a language. Cross-linguistic lexification patterns have been shown to be largely predictable, as similar concepts are often colexified. We test a recent claim that, beyond this general tendency, communicative needs play an important role in shaping colexification patterns. We approach this question by means of a series of human experiments, using an artificial language communication game paradigm. Our results across four experiments match the previous cross-linguistic findings: all other things being equal, speakers do prefer to colexify similar concepts. However, we also find evidence supporting the communicative need hypothesis: when faced with a frequent need to distinguish similar pairs of meanings, speakadjust their colexification preferences to maintain communicative efficiency and avoid colexifying those similar meanings which need to be distinguished in communication. This research provides further evidence to support the argument that languages are shaped by the needs and preferences of their speakers.
In a model of N volume-excluding spheres in a d-dimensional tube, we consider how differences between the drift velocities, diffusivities, and sizes of particles influence the steady-state distribution and axial particle current. We show that the model is exactly solvable when the geometrical constraints prevent any particle from overtaking all others—a notion we term quasi-one-dimensionality. Then, due to a ratchet effect, the current is biased towards the velocities of the least diffusive particles. We consider special cases of this model in one dimension, and derive the exact joint gap distribution for driven tracers in a passive bath. We describe the relationship between phase-space structure and irreversible drift that makes the quasi-one-dimensional (q1D) supposition key to the model’s solvability.
We consider the persistent exclusion process in which a set of persistent random walkers interact via hard-core exclusion on a hypercubic lattice in d dimensions. We work within the ballistic regime whereby particles continue to hop in the same direction over many lattice sites before reorienting. In the case of two particles, we find the mean first-passage time to a jammed state where the particles occupy adjacent sites and face each other. This is achieved within an approximation that amounts to embedding the one-dimensional system in a higher-dimensional reservoir. Numerical results demonstrate the validity of this approximation, even for small lattices. The results admit a straightforward generalization to dilute systems comprising more than two particles. A self-consistency condition on the validity of these results suggest that clusters may form at arbitrarily low densities in the ballistic regime, in contrast to what has been found in the diffusive limit.
We review various combinatorial interpretations and mappings of stationary-state probabilities of the totally asymmetric, partially asymmetric and symmetric simple exclusion processes (TASEP, PASEP, SSEP respectively). In these steady states, the statistical weight of a configuration is determined from a matrix product, which can be written explicitly in terms of generalised ladder operators. This lends a natural association to the enumeration of random walks with certain properties. Specifically, there is a one-to-many mapping of steady-state configurations to a larger state space of discrete paths, which themselves map to an even larger state space of number permutations. It is often the case that the configuration weights in the extended space are of a relatively simple form (e.g. a Boltzmann-like distribution). Meanwhile, various physical properties of the nonequilibrium steady state—such as the entropy—can be interpreted in terms of how this larger state space has been partitioned. These mappings sometimes allow physical results to be derived very simply, and conversely the physical approach allows some new combinatorial problems to be solved. This work brings together results and observations scattered in the combinatorics and statistical physics literature, and also presents new results. The review is pitched at statistical physicists who, though not professional combinatorialists, are competent and enthusiastic amateurs.
Newberry et al. (Detecting evolutionary forces in language change, 'Nature' 551, 2017) tackle an important but difficult problem in linguistics, the testing of selective theories of language change against a null model of drift. Having applied a test from population genetics (the Frequency Increment Test) to a number of relevant examples, they suggest stochasticity has a previously under-appreciated role in language evolution. We replicate their results and find that while the overall observation holds, results produced by this approach on individual time series can be sensitive to how the corpus is organized into temporal segments (binning). Furthermore, we use a large set of simulations in conjunction with binning to systematically explore the range of applicability of the Frequency Increment Test. We conclude that care should be exercised with interpreting results of tests like the Frequency Increment Test on individual series, given the researcher degrees of freedom available when applying the test to corpus data, and fundamental differences between genetic and linguistic data. Our findings have implications for selection testing and temporal binning in general, as well as demonstrating the usefulness of simulations for evaluating methods newly introduced to the field.
The availability of large diachronic corpora has provided the impetus for a growing body of quantitative research on language evolution and meaning change. The central quantities in this research are token frequencies of linguistic elements in texts, with changes in frequency taken to reflect the popularity or selective fitness of an element. However, corpus frequencies may change for a wide variety of reasons, including purely random sampling effects, or because corpora are composed of contemporary media and fiction texts within which the underlying topics ebb and flow with cultural and socio-political trends. In this work, we introduce a simple model for controlling for topical fluctuations in corpora—the topical-cultural advection model —and demonstrate how it provides a robust baseline of variability in word frequency changes over time. We validate the model on a diachronic corpus spanning two centuries, and a carefully-controlled artificial language change scenario, and then use it to correct for topical fluctuations in historical time series. Finally, we use the model to show that the emergence of new words typically corresponds with the rise of a trending topic. This suggests that some lexical innovations occur due to growing communicative need in a subspace of the lexicon, and that the topical-cultural advection model can be used to quantify this.
The mechanisms which influence coexistence of specialist and generalist species in the same environment are a key focus of ecological theory. We use an agent-based model of community assembly to show that the available resource spectrum (distribution of resources along a niche axis) can play an important role in determining the specialist-generalist balance, even in the absence of spatial structure. Our results reveal a phenomenon that we term 'resource spectrum engineering', in which opportunistic specialists occupying small niches in a mostly generalist community can change the resource spectrum that is experienced by other species, in a way that disfavours generalists and causes a community-wide shift towards specialist strategies. More generally, this suggests a mechanism by which apparently minor changes in the specialist composition of an ecological community could have knock-on effects across the entire community.
How Laggards Help Decision-MakingCollective decision-making in a social network is better when there are both early adopters and
All living languages change over time. The causes for this are many, one being the emergence and borrowing of new linguistic elements. Competition between the new elements and older ones with a similar semantic or grammatical function may lead to speakers preferring one of them, and leaving the other to go out of use. We introduce a general method for quantifying competition between linguistic elements in diachronic corpora which does not require language-specific resources other than a sufficiently large corpus. This approach is readily applicable to a wide range of languages and linguistic subsystems. Here, we apply it to lexical data in five corpora differing in language, type, genre, and time span. We find that changes in communicative need are consistently predictive of lexical competition dynamics. Near-synonymous words are more likely to directly compete if they belong to a topic of conversation whose importance to language users is constant over time, possibly leading to the extinction of one of the competing words. By contrast, in topics which are increasing in importance for language users, near-synonymous words tend not to compete directly and can coexist. This suggests that, in addition to direct competition between words, language change can be driven by competition between topics or semantic subspaces.