Abstract Sensitivity to small changes in the environment is crucial for many real-world tasks, enabling living and artificial systems to make correct behavioral decisions. It has been shown that such sensitivity is maximized when a system operates near the critical point of a phase transition. However, proximity to criticality introduces large fluctuations and diverging timescales. Hence, to leverage the maximal sensitivity, it could require impractically long integration periods. Here, we analytically and computationally demonstrate how the optimal tuning of a recurrent neural network is determined given a finite integration time. Rather than maximizing the theoretically available sensitivity, we find networks attain different sensitivities depending on the available time. Consequently, the optimal dynamic regime can shift away from criticality when integration times are finite, highlighting the necessity of incorporating finite-time considerations into studies of information processing.
Cortical neurons are complex, multi-timescale processors wired into recurrent circuits, shaped by long evolutionary pressure under stringent biological constraints. Mainstream machine learning, by contrast, predominantly builds models from extremely simple units, a default inherited from early neural-network theory. We treat this as a normative architectural question. How should one split a fixed parameter budget P between the number of units N, per-unit effective complexity k_e, and per-unit connectivity k_c? What controls the optimal allocation? This calls for a model in which per-unit complexity can be tuned independently of width and connectivity. Accordingly, we introduce the ELM Network, whose recurrent layer is built from Expressive Leaky Memory (ELM) neurons, chosen to mirror functional components of cortical neurons. The architecture allows for individually adjusting N, k_e, and k_c and trains stably across orders of magnitude in scale. We evaluate the model on two qualitatively different sequence benchmarks: the neuromorphic SHD-Adding task and Enwik8 character-level language modeling. Performance improves monotonically along each of the three axes individually. Under a fixed budget, a clear non-trivial optimum emerges in their tradeoff, and larger budgets favor both more and more complex neurons. A closed-form information-theoretic model captures these tradeoffs and attributes the diminishing returns at two ends to: per-neuron signal-to-noise saturation and across-neuron redundancy. A hyperparameter sweep spanning three orders of magnitude in trainable parameters traces a near-Pareto-frontier scaling law consistent with the framework. This suggests that the simple-unit default in ML is not obviously optimal once this tradeoff surface is probed, and offers a normative lens on cortex's reliance on complex spatio-temporal integrators.
Neural activity fluctuates over a wide range of timescales within and across brain areas. Experimental observations suggest that diverse neural timescales reflect information in dynamic environments. However, how timescales are defined and measured from brain recordings vary across the literature. Moreover, these observations do not specify the mechanisms underlying timescale variations, nor whether specific timescales are necessary for neural computation and brain function. Here, we synthesize three directions where computational approaches can distill the broad set of empirical observations into quantitative and testable theories: We review (i) how different data analysis methods quantify timescales across distinct behavioral states and recording modalities, (ii) how biophysical models provide mechanistic explanations for the emergence of diverse timescales, and (iii) how task-performing networks and machine learning models uncover the functional relevance of neural timescales. This integrative computational perspective thus complements experimental investigations, providing a holistic view on how neural timescales reflect the relationship between brain structure, dynamics, and behavior.
Biological learning achieves temporal credit assignment despite sparse and imprecise feedback, often relying on neuromodulatory signals acting over space and time. Here, we introduce a learning mechanism in which error information diffuses locally through the network, similar to volume transmission of neuromodulators. This distributed modulation allows neurons to learn even in the absence of direct feedback, using the local concentration of the diffusing credit signal. Applied to recurrent spiking neural networks with sparse feedback connectivity, diffusive credit signaling improves learning across three benchmark tasks. Using eligibility propagation as a baseline learning mechanism, we show how diffusion-based modulation can provide a plausible mechanism for credit assignment in sparsely connected neural circuits.
Developing reliable mechanisms for continuous local learning is a central challenge faced by biological and artificial systems. Yet, how the environmental factors and structural constraints on the learning network influence the optimal plasticity mechanisms remains obscure even for simple settings. To elucidate these dependencies, we study meta-learning via evolutionary optimization of simple reward-modulated plasticity rules in embodied agents solving a foraging task. We show that unconstrained meta-learning leads to the emergence of diverse plasticity rules. However, regularization and bottlenecks in the model help reduce this variability, resulting in interpretable rules. Our findings indicate that the meta-learning of plasticity rules is very sensitive to various parameters, with this sensitivity possibly reflected in the learning rules found in biological networks. When included in models, these dependencies can be used to discover potential objective functions and details of biological learning via comparisons with experimental observations.
Cultures of neurons in vitro are instrumental for studying network dynamics in normal and pathological conditions. Mature networks typically exhibit network bursting activity, which has traditionally been quantified by simplified features such as inter-burst intervals and burst durations. While these features advanced the understanding of development, disease phenotypes, and drug effects, they overlook the temporal structure of activity within bursts. Here, we developed a comprehensive framework to quantify burst shapes, the time course of network firing during bursts. Applying this approach to four datasets, including rodent- and human pluripotent stem cell-derived cultures, we show that burst shapes contain rich information about the underlying network dynamics. We quantify this information by using traditional and shape features to classify the recording conditions (types of genetic disorder, presence of pharmacological agents) and demonstrate that shapes significantly increase classification accuracy. We provide a pipeline for burst shape characterization, including simplified features that capture most of the shape information, establishing burst shape as a robust and biologically meaningful marker for functional phenotyping in disease modeling and drug screening. ### Competing Interest Statement The authors have declared no competing interest. Sofja Kovalevskaja Award from the Alexander von Humboldt Foundation Else Kröner Medical Scientist Kolleg “ClinBrAIn: Artificial Intelligence for Clinical Brain Research”, Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy – EXC number 2064/1, 390727645 Max Planck Research School for Intelligent Systems International Max Planck Research School for the Mechanisms of Mental Function and Dysfunction Joachim Herz Foundation BMBF through the Tuebingen AI Center, 01IS18039B Machine Learning Cluster of Excellence, 39072764
Variations in intrinsic neural timescales across the mammalian forebrain reflect the anatomical structure and functional specialization of brain areas and individual neurons. Yet, the organization of timescales beyond the forebrain remains unexplored. We analyzed intrinsic timescales of single neurons across the entire mouse brain. Median timescales were up to fivefold longer in the midbrain and hindbrain than in the forebrain. Spatial patterns of gene expression predicted timescale variation at a resolution finer than brain-area boundaries. Across neurons, the diversity of timescales revealed a multiscale architecture, in which fast timescales determined regional differences in medians, while slow timescales universally followed a power-law distribution with an exponent near 2, indicating a shared dynamical regime across the brain consistent with the edge of instability or chaos. These organizing principles for the dynamics of single neurons across the brain provide a foundation for linking cellular activity with regional specialization and brain-wide computation.
Schizophrenia (SCZ) is a highly heritable brain disorder marked by a wide range of changes throughout the central nervous system. These changes include alterations at the molecular and cellular levels, suggesting significant disruptions in synapse function, as well as modifications in brain structure and activity. However, it remains unclear, how changes in molecular synapse biology translate into neurophysiological and ultimately behavioral consequences across scales. Here, we narrow this translational gap in contemporary biological psychiatry by establishing a generalizable framework to bridge the scales and pinpoint biological mechanisms underlying individual psychiatric symptoms. We show that genetically driven changes in neuronal gene expression and a resulting reduction in excitatory synaptic density in vitro are linked to alterations of brain structure, electrophysiology and ultimately cognitive function in vivo. These results provide a direct connection between the molecular origins of synapse reduction in SCZ and its neurobiological and phenotypic consequences on the individual patient level, paving the way to develop new mechanism informed treatment options. ### Competing Interest Statement The authors declare that there are no conflicts of interest in relation to the subject of this study. General declaration of potential conflict of interests: SG is part-time employees by and shareholders of Systasy Bioscience GmbH, Munich, Germany. MJR is shareholder and consultant of Systasy Bioscence GmbH. AH received speaker fees from AbbVie, Advanz, Janssen, Otsuka, Lundbeck, Rovi, and Recordati and was a member of the advisory boards of these companies and Boehringer Ingelheim. BS and MZ received speaker fees from Novartis Pharma GmbH. EW was a member of the advisory boards of Boehringer Ingelheim and Recordati. OP received speaker fees from Lundbeck, Otsuka, Takeda, and Janssen and was a member of the advisory boards of Lundbeck and Janssen. PF received speaker fees from Boehringer Ingelheim, Janssen, Otsuka, Lundbeck, Recordati, and Richter and was a member of the advisory boards of these companies and Rovi. ### Funding Statement This work was supported by BMBF, eMed grant numbers 01ZX1504, 01ZX1706A (MJZ), Else-Kroener-Fresenius Stiftung grant A54 (MJZ), DFG Grants GZ: ZI 1614/5-1,ZI 1614/7-1 (MJZ). EB received funding from the Pesl-Alzheimer-Stiftung (2024-2025). DP and FJR were supported by the Else Kroener-Fresenius Foundation (Research College Translational Psychiatry) for their Residency/Ph.D. track at the International Max Planck Research School for Translational Psychiatry (IMPRS-TP), Munich, Germany. FJR and ECS were supported by the Munich Clinician Scientist Program (MCSP) of the Faculty of Medicine, LMU Munich, Munich, Germany (FoeFoLe 009/2019 and Advanced Track 01/2021, respectively). FJR received funding from the Lisa Oehler-Stiftung (2022-2024), the Pesl-Alzheimer-Stiftung (2024-2025). VY was supported by the Residency/PhD track of the International Max Planck Research School for Translational Psychiatry (IMPRS-TP) and was supported by the Faculty of Medicine at LMU Munich (FoeFoLe Reg.-Nr. 1226/2024). JM was supported by the Faculty of Medicine at LMU Munich (FoeFoLe Reg.-Nr. 1167). The study was supported by the EU HORIZON-INFRA-2024-TECH-01-04 project DTRIP4H 101188432 to PF, AS and FR. PF, AS, GH and VY received funding from the BMBF within the Era-Net Neuron project GDNF UpReg (FKZ 01EW2206). The study was endorsed by the Federal Ministry of Education and Research (Bundesministerium fuer Bildung und Forschung [BMBF]) within the initial phase of the German Center for Mental Health (DZPG) (grant: 01EE2303C to AH, and 01EE2303A, 01EE2303F to PF).The study was supported by the Supplement to BMBF funding for the German Centre for Mental Health (DZPG) by the Bavarian State Ministry for Science and the Arts with the Grant for the research project Improving Infrastructures for DZPG and NAKO Cohorts to PF, DK and BK. TS received funding through the Else Kroener Medical Scientist Kolleg ClinbrAIn: Artificial Intelligence for Clinical Brain Research and is supported by the International Max Planck Research School for Intelligent Systems (IMPRS-IS). AL is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 - project number 39072764. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethics committee of the Faculty of Medicine, LMU Munich, Project 17-880, 29.03.2018; project 18-716, 15.10.2020) and at the MPI for Psychiatry (approved by the local ethics committee of the Faculty of Medicine, LMU Munich, project numbers 350-14, 19-310, 20-314,19-678 and 18-393 I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
The critical brain hypothesis states that the brain can benefit from operating close to a second-order phase transition. While it has been shown that several computational aspects of sensory processing (e.g., sensitivity to input) can be optimal in this regime, it is still unclear whether these computational benefits of criticality can be leveraged by neural systems performing behaviorally relevant computations. To address this question, we investigate signatures of criticality in networks optimized to perform efficient coding. We consider a spike-coding network of leaky integrate-and-fire neurons with synaptic transmission delays. Previously, it was shown that the performance of such networks varies nonmonotonically with the noise amplitude. Interestingly, we find that in the vicinity of the optimal noise level for efficient coding, the network dynamics exhibit some signatures of criticality, namely, scale-free dynamics of the spiking and the presence of crackling noise relation. Our work suggests that two influential, and previously disparate theories of neural processing optimization (efficient coding and criticality) may be intimately related.
Biological cortical neurons are remarkably sophisticated computational devices, temporally integrating their vast synaptic input over an intricate dendritic tree, subject to complex, nonlinearly interacting internal biological processes. A recent study proposed to characterize this complexity by fitting accurate surrogate models to replicate the input-output relationship of a detailed biophysical cortical pyramidal neuron model and discovered it needed temporal convolutional networks (TCN) with millions of parameters. Requiring these many parameters, however, could stem from a misalignment between the inductive biases of the TCN and cortical neuron's computations. In light of this, and to explore the computational implications of leaky memory units and nonlinear dendritic processing, we introduce the Expressive Leaky Memory (ELM) neuron model, a biologically inspired phenomenological model of a cortical neuron. Remarkably, by exploiting such slowly decaying memory-like hidden states and two-layered nonlinear integration of synaptic input, our ELM neuron can accurately match the aforementioned input-output relationship with under ten thousand trainable parameters. To further assess the computational ramifications of our neuron design, we evaluate it on various tasks with demanding temporal structures, including the Long Range Arena (LRA) datasets, as well as a novel neuromorphic dataset based on the Spiking Heidelberg Digits dataset (SHD-Adding). Leveraging a larger number of memory units with sufficiently long timescales, and correspondingly sophisticated synaptic integration, the ELM neuron displays substantial long-range processing capabilities, reliably outperforming the classic Transformer or Chrono-LSTM architectures on LRA, and even solving the Pathfinder-X task with over 70% accuracy (16k context length).
Cortical neurons are versatile and efficient coding units that develop strong preferences for specific stimulus characteristics. The sharpness of tuning and coding efficiency is hypothesized to be controlled by delicately balanced excitation and inhibition. These observations suggest a need for detailed co-tuning of excitatory and inhibitory populations. Theoretical studies have demonstrated that a combination of plasticity rules can lead to the emergence of excitation/inhibition (E/I) cotuning in neurons driven by independent, low-noise signals. However, cortical signals are typically noisy and originate from highly recurrent networks, generating correlations in the inputs. This raises questions about the ability of plasticity mechanisms to self-organize co-tuned connectivity in neurons receiving noisy, correlated inputs. Here, we study the emergence of input selectivity and weight co-tuning in a neuron receiving input from a recurrent network via plastic feedforward connections. We demonstrate that while strong noise levels destroy the emergence of co-tuning in the readout neuron, introducing specific structures in the non-plastic pre-synaptic connectivity can re-establish it by generating a favourable correlation structure in the population activity. We further show that structured recurrent connectivity can impact the statistics in fully plastic recurrent networks, driving the formation of co-tuning in neurons that do not receive direct input from other areas. Our findings indicate that the network dynamics created by simple, biologically plausible structural connectivity patterns can enhance the ability of synaptic plasticity to learn input-output relationships in higher brain areas.
Neuronal cultures in vitro are a versatile system for studying the fundamental properties of individual neurons and neuronal networks. Recently, this approach has gained attention as a precision medicine tool. Mature neuronal cultures in vitro exhibit synchronized collective dynamics called network bursting. If analyzed appropriately, this activity could offer insights into the network's properties, such as its composition, topology, and developmental and pathological processes. A promising method for investigating the collective dynamics of neuronal networks is to map them onto simplified dynamical systems. This approach allows the study of dynamical regimes and the characteristics of the parameters that lead to data-consistent activity. We designed a simple biophysically inspired dynamical system and used Bayesian inference to fit it to a large number of recordings of in vitro population activity. Even with a small number of parameters, the model showed strong inter-parameter dependencies leading to invariant bursting dynamics for many parameter combinations. We further validated this observation in our analytical solution. We found that in vitro bursting can be well characterized by each of three dynamical regimes: oscillatory, bistable, and excitable. The probability of finding a data-consistent match in a particular regime changes with network composition and development. The more informative way to describe the in vitro network bursting is the effective excitability, which we analytically show to be related to the parameter-invariance of the model's dynamics. We establish that the effective excitability can be estimated directly from the experimentally recorded data. Finally, we demonstrate that effective excitability reliably detects the differences between cultures of cortical, hippocampal, and human pluripotent stem cell-derived neurons, allowing us to map their developmental trajectories. Our results open a new avenue for the model-based description of in vitro network phenotypes emerging across different experimental conditions. ### Competing Interest Statement The authors have declared no competing interest.
As more connectome data become available, the question of how to best analyse the structure of biological neural networks becomes increasingly pertinent. In brain networks, knowing that two areas are connected is often not sufficient, as the directionality and weight of the connection affect the dynamics in crucial ways. Still, the methods commonly used to estimate network properties, such as clustering and small-worldness, usually disregard features encoded in the directionality and strength of network connections. To address this issue, we propose using fully-weighted and directed clustering measures that provide higher sensitivity to non-random structural features. Using artificial networks, we demonstrate the problems with methods routinely used in the field and how fully-weighted and directed methods can alleviate them. Specifically, we highlight their robustness to noise and their ability to address thresholding issues, particularly in inferred networks. We further apply our method to the connectomes of different species and uncover regularities and correlations between neuronal structures and functions that cannot be detected with traditional clustering metrics. Finally, we extend the notion of small-worldness in brain networks to account for weights and directionality and show that some connectomes can no longer be considered ``small-world''. Overall, our study makes a case for a combined use of fully-weighted and directed measures to deal with the variability of brain networks and suggests the presence of complex patterns in neural connectivity that can only be revealed using such methods.
Structural modularity is a pervasive feature of biological neural networks, which have been linked to several functional and computational advantages. Yet, the use of modular architectures in artificial neural networks has been relatively limited despite early successes. Here, we explore the performance and functional dynamics of a modular network trained on a memory task via an iterative growth curriculum. We find that for a given classical, non-modular recurrent neural network (RNN), an equivalent modular network will perform better across multiple metrics, including training time, generalizability, and robustness to some perturbations. We further examine how different aspects of a modular network's connectivity contribute to its computational capability. We then demonstrate that the inductive bias introduced by the modular topology is strong enough for the network to perform well even when the connectivity within modules is fixed and only the connections between modules are trained. Our findings suggest that gradual modular growth of RNNs could provide advantages for learning increasingly complex tasks on evolutionary timescales, and help build more scalable and compressible artificial networks.
Multiple studies of neural avalanches across different data modalities led to the prominent hypothesis that the brain operates near a critical point. The observed exponents often indicate the mean-field directed-percolation universality class, leading to the fully-connected or random network models to study the avalanche dynamics. However, the cortical networks have distinct non-random features and spatial organization that is known to affect the critical exponents. Here we show that distinct empirical exponents arise in networks with different topology and depend on the network size. In particular, we find apparent scale-free behavior with mean-field exponents appearing as quasi-critical dynamics in structured networks. This quasi-critical dynamics cannot be easily discriminated from an actual critical point in small networks. We find that the local coalescence in activity dynamics can explain the distinct exponents. Therefore, both topology and system size should be considered when assessing criticality from empirical observables.
Recurrent neural networks (RNNs) in the brain and in silico excel at solving tasks with intricate temporal dependencies. Long timescales required for solving such tasks can arise from properties of individual neurons (single-neuron timescale, $\tau$, e.g., membrane time constant in biological neurons) or recurrent interactions among them (network-mediated timescale). However, the contribution of each mechanism for optimally solving memory-dependent tasks remains poorly understood. Here, we train RNNs to solve $N$-parity and $N$-delayed match-to-sample tasks with increasing memory requirements controlled by $N$ by simultaneously optimizing recurrent weights and $\tau$s. We find that for both tasks RNNs develop longer timescales with increasing $N$, but depending on the learning objective, they use different mechanisms. Two distinct curricula define learning objectives: sequential learning of a single-$N$ (single-head) or simultaneous learning of multiple $N$s (multi-head). Single-head networks increase their $\tau$ with $N$ and are able to solve tasks for large $N$, but they suffer from catastrophic forgetting. However, multi-head networks, which are explicitly required to hold multiple concurrent memories, keep $\tau$ constant and develop longer timescales through recurrent connectivity. Moreover, we show that the multi-head curriculum increases training speed and network stability to ablations and perturbations, and allows RNNs to generalize better to tasks beyond their training regime. This curriculum also significantly improves training GRUs and LSTMs for large-$N$ tasks. Our results suggest that adapting timescales to task requirements via recurrent interactions allows learning more complex objectives and improves the RNN's performance.
Many settings in machine learning require the selection of a rotation representation. However, choosing a suitable representation from the many available options is challenging. This paper acts as a survey and guide through rotation representations. We walk through their properties that harm or benefit deep learning with gradient-based optimization. By consolidating insights from rotation-based learning, we provide a comprehensive overview of learning functions with rotation representations. We provide guidance on selecting representations based on whether rotations are in the model's input or output and whether the data primarily comprises small angles.
Correlated fluctuations in the activity of neural populations reflect the network's dynamics and connectivity. The temporal and spatial dimensions of neural correlations are interdependent. However, prior theoretical work mainly analyzed correlations in either spatial or temporal domains, oblivious to their interplay. We show that the network dynamics and connectivity jointly define the spatiotemporal profile of neural correlations. We derive analytical expressions for pairwise correlations in networks of binary units with spatially arranged connectivity in one and two dimensions. We find that spatial interactions among units generate multiple timescales in auto-and cross-correlations. Each timescale is associated with fluctuations at a particular spatial frequency, making a hierarchical contribution to the correlations. External inputs can modulate the correlation timescales when spatial interactions are nonlinear, and the modulation effect depends on the operating regime of network dynamics. These theoretical results open new ways to relate connectivity and dynamics in cortical networks via measurements of spatiotemporal neural correlations.