Reinforcement learning (RL) offers a compelling account of how agents learn complex behaviors by trial and error, yet RL is predicated on the existence of a reward function provided by the agent's environment. By contrast, many skills are learned without external guidance, posing a challenge to RL's ability to account for self-directed learning. For instance, juvenile male zebra finches first memorize and then train themselves to reproduce the song of an adult male tutor through extensive practice. This process is believed to be guided by an internally computed assessment of performance quality, though the mechanism and development of this signal remain unknown. Here, we propose that, contrary to prevailing assumptions, tutor song memorization and performance assessment are subserved by the same neural circuit, one trained to predictively cancel tutor song. To test this hypothesis, we built models of a local forebrain circuit that learns to use contextual input from premotor regions to cancel tutor song auditory input via plasticity at different synaptic loci. We found that, after learning, excitatory projection neurons in these circuits exhibited population error codes signaling mismatches between the tutor song memory and birds' own performance, and these signals best matched experimental data when networks were trained with anti-Hebbian plasticity in the recurrent pathway through inhibitory interneurons. We also found that model learning proceeds in two stages, with an initial phase of sharpening error sensitivity followed by a fine-tuning period in which error responses to the tutor song are minimized. Finally, we showed that the error signal produced by this model can train a simple RL agent to replicate the spectrograms of adult bird songs. Together, our results suggest that purely local learning via predictive cancellation suffices for bootstrapping error signals capable of guiding self-directed learning of natural behaviors.
Socially effective vocal communication requires brain regions that encode expressive and receptive aspects of vocal communication in a social context-dependent manner. Here, we combined a novel behavioral assay with microendoscopic calcium imaging to interrogate neuronal activity (regions of interest [ROIs]) in the posterior insula (pIns) in socially interacting mice as they switched rapidly between states of vocal expression and reception. We found that largely distinct subsets of pIns ROIs were active during vocal expression and reception. Notably, pIns activity during vocal expression increased prior to vocal onset and was also detected in congenitally deaf mice, pointing to a motor signal. Furthermore, receptive pIns activity was modulated strongly by social context. Lastly, tracing experiments reveal that deep-layer neurons in the pIns directly bridge the auditory thalamus to a midbrain vocal gating region. Therefore, the pIns is a site that encodes vocal expression and reception in a manner that depends on social context.
Although learning in response to extrinsic reinforcement is theorized to be driven by dopamine signals that encode the difference between expected and experienced rewards1,2, skills that enable verbal or musical expression can be learned without extrinsic reinforcement. Instead, spontaneous execution of these skills is thought to be intrinsically reinforcing3,4. Whether dopamine signals similarly guide learning of these intrinsically reinforced behaviours is unknown. In juvenile zebra finches learning from an adult tutor, dopamine signalling in a song-specialized basal ganglia region is required for successful song copying, a spontaneous, intrinsically reinforced process5. Here we show that dopamine dynamics in the song basal ganglia faithfully track the learned quality of juvenile song performance on a rendition-by-rendition basis. Furthermore, dopamine release in the basal ganglia is driven not only by inputs from midbrain dopamine neurons classically associated with reinforcement learning but also by song premotor inputs, which act by means of local cholinergic signalling to elevate dopamine during singing. Although both cholinergic and dopaminergic signalling are necessary for juvenile song learning, only dopamine tracks the learned quality of song performance. Therefore, dopamine dynamics in the basal ganglia encode performance quality during self-directed, long-term learning of natural behaviours.
Vocal communication depends on distinguishing self-generated vocalizations from other sounds. Vocal motor corollary discharge (CD) signals are thought to support this ability by adaptively suppressing auditory cortical responses to auditory feedback. One challenge is that vocalizations, especially those produced during courtship and other social interactions, are accompanied by other movements and are emitted during a state of heightened arousal, factors that could potentially modulate auditory cortical activity. Here, we monitor auditory cortical activity, ultrasonic vocalizations (USVs), and other non-vocal courtship behaviors in a head-fixed male mouse while he interacts with a female mouse. This approach reveals a vocalization-specific signature in the auditory cortex that suppresses the activity of USV playback-excited neurons, emerges before vocal onset, and scales with USV band power. Notably, this vocal modulatory signature is also present in the auditory cortex of congenitally deaf mice, revealing an adaptive vocal CD signal that manifests independently of auditory feedback or auditory experience.
Socially effective vocal communication requires brain regions that encode expressive and receptive aspects of vocal communication in a social context-dependent manner. Here, we combined a novel behavioral assay with microendoscopy to interrogate neuronal activity in the posterior insula (pIns) in socially interacting mice as they switched rapidly between states of vocal expression and reception. We found that distinct but spatially intermingled subsets of pIns neurons were active during vocal expression and reception. Notably, pIns activity during vocal expression increased prior to vocal onset and was also detected in congenitally deaf mice, pointing to a motor signal. Furthermore, receptive pIns activity depended strongly on social cues, including female odorants. Lastly, tracing experiments reveal that deep layer neurons in the pIns directly bridge the auditory thalamus to a midbrain vocal gating region. Therefore, the pIns is a site that encodes vocal expression and reception in a manner that depends on social context.
Courtship displays often involve the concerted production of several distinct courtship behaviors. The neural circuits that enable the concerted production of the component behaviors of a courtship display are not well understood. Here, we identify a midbrain cell group (A11) that enables male zebra finches to produce their learned songs in concert with various other behaviors, including female-directed orientation, pursuit, and calling. Anatomical mapping reveals that A11 is at the center of a complex network including the song premotor nucleus HVC as well as brainstem regions crucial to calling and locomotion. Notably, lesioning A11 terminals in HVC blocked female-directed singing but did not interfere with female-directed calling, orientation, or pursuit. In contrast, lesioning A11 cell bodies strongly reduced and often abolished all female-directed courtship behaviors. However, males with either type of lesion still produced songs when in social isolation. Lastly, imaging calcium-related activity in A11 terminals in HVC showed that during courtship, A11 signals HVC about female-directed calls and during female-directed singing, about the transition from simpler introductory notes to the acoustically more complex syllables that depend intimately on HVC for their production. These results show how a brain region important to reproduction in both birds and mammals enables holistic courtship displays in male zebra finches, which include learning songs, calls, and other non-vocal behaviors.
Vocalizations facilitate mating and social affiliation but may also inadvertently alert predators and rivals. Consequently, the decision to vocalize depends on brain circuits that can weigh and compare these potential benefits and risks. Male mice produce ultrasonic vocalizations (USVs) during courtship to facilitate mating, and previously isolated female mice produce USVs during social encounters with novel females. Earlier we showed that a specialized set of neurons in the midbrain periaqueductal gray (PAG-USV neurons) are an obligatory gate for USV production in both male and female mice, and that both PAG-USV neurons and USVs can be switched on by their inputs from the preoptic area (POA) of the hypothalamus and switched off by their inputs from neurons on the border between the central and medial amygdala (AmgC/M-PAG neurons) (Michael et al., 2020). Here, we show that the USV-suppressing AmgC/M-PAG neurons are strongly activated by predator cues or during social contexts that suppress USV production in male and female mice. Further, we explored how vocal promoting and vocal suppressing drives are weighed in the brain to influence vocal production in male mice, where the drive and courtship function for USVs are better understood. We found that AmgC/M-PAG neurons receive monosynaptic inhibitory input from POA neurons that also project to the PAG, that these inhibitory inputs are active in USV-promoting social contexts, and that optogenetic activation of POA cell bodies that make divergent axonal projections to the amygdala and PAG is sufficient to elicit USV production in socially isolated male mice. Accordingly, AmgC/M-PAG neurons, along with POAPAG and PAG-USV neurons, form a nested hierarchical circuit in which environmental and social information converges to influence the decision to vocalize.
The locus coeruleus (LC) is a small noradrenergic brainstem nucleus that plays a central role in regulating arousal, attention, and performance. In the mammalian brain, individual LC neurons make divergent axonal projections to different brain regions, which are distinguished in part by which noradrenaline (NA) receptor subtypes they express. Here, we sought to determine whether similar organizational features characterize LC projections to corticobasal ganglia (CBG) circuitry in the zebra finch song system, with a focus on the basal ganglia nucleus Area X, the thalamic nucleus DLM, as well as the cortical nuclei HVC, LMAN, and RA. Single and dual retrograde tracer injections reveal that single LC–NA neurons make divergent projections to LMAN and Area X, as well as to the dopaminergic VTA/SNc complex that innervates this CBG circuit. Moreover, in situ hybridization revealed that differential expression of mRNA encoding α 2A and α 2C adrenoreceptors distinguishes LC‐recipient CBG song nuclei. Therefore, LC–NA signaling in the zebra finch CBG circuit employs a similar strategy as in mammals, which could allow a relatively small number of LC neurons to exert widespread yet distinct effects across multiple brain regions.
Learning skilled behaviors requires intensive practice over days, months, or years. Behavioral hallmarks of practice include exploratory variation and long-term improvements, both of which can be impacted by circadian processes. During weeks of vocal practice, the juvenile male zebra finch transforms highly variable and simple song into a stable and precise copy of an adult tutor's complex song. Song variability and performance in juvenile finches also exhibit circadian structure that could influence this long-term learning process. In fact, one influential study reported juvenile song regresses towards immature performance overnight, while another suggested a more complex pattern of overnight change. However, neither of these studies thoroughly examined how circadian patterns of variability may structure the production of more or less mature songs. Here we relate the circadian dynamics of song maturation to circadian patterns of song variation, leveraging a combination of data-driven approaches. In particular we analyze juvenile singing in learned feature space that supports both data-driven measures of song maturity and generative developmental models of song production. These models reveal that circadian fluctuations in variability lead to especially regressive morning variants even without overall overnight regression, and highlight the utility of data-driven generative models for untangling these contributions.
Have your ever felt as happy as a lark, feathered your nest or taken someone under your wing? As we watch birds, we cannot help but be struck by their uncannily familiar behaviors - singing, nest building, caring for their young - to name just a few. Songbirds - the oscine suborder of perching birds that constitute roughly half (∼4,000) of all known avian species - are noted for the songs that males and sometimes both sexes in this group sing to court mates and defend territory from rivals. Birdsongs contain several to many acoustically distinct syllables, typically organized into a stereotyped phrase, and span the same audio bandwidth that we exploit for speech and music, making them easy for us to hear and appreciate. Consequently, eavesdropping humans long ago detected the most striking parallel between songbirds and humans: juvenile songbirds learn to sing in a manner similar to a child learning to speak.
Increases in the scale and complexity of behavioral data pose an increasing challenge for data analysis. A common strategy involves replacing entire behaviors with small numbers of handpicked, domain-specific features, but this approach suffers from several crucial limitations. For example, handpicked features may miss important dimensions of variability, and correlations among them complicate statistical testing. Here, by contrast, we apply the variational autoencoder (VAE), an unsupervised learning method, to learn features directly from data and quantify the vocal behavior of two model species: the laboratory mouse and the zebra finch. The VAE converges on a parsimonious representation that outperforms handpicked features on a variety of common analysis tasks, enables the measurement of moment-by-moment vocal variability on the timescale of tens of milliseconds in the zebra finch, provides strong evidence that mouse ultrasonic vocalizations do not cluster as is commonly believed, and captures the similarity of tutor and pupil birdsong with qualitatively higher fidelity than previous approaches. In all, we demonstrate the utility of modern unsupervised learning approaches to the quantification of complex and high-dimensional vocal behavior.
Musical and athletic skills are learned and maintained through intensive practice to enable precise and reliable performance for an audience. Consequently, understanding such complex behaviours requires insight into how the brain functions during both practice and performance. Male zebra finches learn to produce courtship songs that are more varied when alone and more stereotyped in the presence of females 1 . These differences are thought to reflect song practice and performance, respectively 2 , 3 , providing a useful system in which to explore how neurons encode and regulate motor variability in these two states. Here we show that calcium signals in ensembles of spiny neurons (SNs) in the basal ganglia are highly variable relative to their cortical afferents during song practice. By contrast, SN calcium signals are strongly suppressed during female-directed performance, and optogenetically suppressing SNs during practice strongly reduces vocal variability. Unsupervised learning methods 4 , 5 show that specific SN activity patterns map onto distinct song practice variants. Finally, we establish that noradrenergic signalling reduces vocal variability by directly suppressing SN activity. Thus, SN ensembles encode and drive vocal exploration during practice, and the noradrenergic suppression of SN activity promotes stereotyped and precise song performance for an audience.
Holistic behaviors often require the coordination of innate and learned movements. The neural circuits that enable such coordination remain unknown. Here we identify a midbrain cell group (A11) that enables male zebra finches to coordinate their learned songs with various innate behaviors, including female-directed calling, orientation and pursuit. Anatomical mapping reveals that A11 is at the center of a complex network including the song premotor nucleus HVC as well as brainstem regions crucial to innate calling and locomotion. Notably, lesioning A11 terminals in HVC blocked female-directed singing, but did not interfere with female-directed calling, orientation or pursuit. In contrast, lesioning A11 cell bodies abolished all female-directed courtship behaviors. However, males with either type of lesion still produced songs when in social isolation. Lastly, monitoring A11 terminals in HVC showed that during courtship A11 inputs to the song premotor cortex signal the transition from innate to learned vocalizations. These results show how a brain region important to reproduction in both birds and mammals coordinates learned vocalizations with innate, ancestral courtship behaviors.
Animals vocalize only in certain behavioral contexts, but the circuits and synapses through which forebrain neurons trigger or suppress vocalization remain unknown. Here, we used transsynaptic tracing to identify two populations of inhibitory neurons that lie upstream of neurons in the periaqueductal gray (PAG) that gate the production of ultrasonic vocalizations (USVs) in mice (i.e. PAG-USV neurons). Activating PAG-projecting neurons in the preoptic area of the hypothalamus (POAPAG neurons) elicited USV production in the absence of social cues. In contrast, activating PAG-projecting neurons in the central-medial boundary zone of the amygdala (AmgC/M-PAG neurons) transiently suppressed USV production without disrupting non-vocal social behavior. Optogenetics-assisted circuit mapping in brain slices revealed that POAPAG neurons directly inhibit PAG interneurons, which in turn inhibit PAG-USV neurons, whereas AmgC/M-PAG neurons directly inhibit PAG-USV neurons. These experiments identify two major forebrain inputs to the PAG that trigger and suppress vocalization, respectively, while also establishing the synaptic mechanisms through which these neurons exert opposing behavioral effects.
•Genetic identification of brainstem neurons that gate mouse USVs.•Attractor networks may convert auditory stimuli into vocal action.•The motor cortex plays a role in timing antiphonal vocalizations.•Juvenile songbirds vocally copy optogenetically ‘implanted’ song memories.•Identification of circuits that evaluate and reinforce vocal performance.
Virtuosic motor performance requires the ability to evaluate and modify individual gestures within a complex motor sequence. Where and how the evaluative and premotor circuits operate within the brain to enable such temporally precise learning is poorly understood. Songbirds can learn to modify individual syllables within their complex vocal sequences, providing a system for elucidating the underlying evaluative and premotor circuits. We combined behavioral and optogenetic methods to identify 2 afferents to the ventral tegmental area (VTA) that serve evaluative roles in syllable-specific learning and to establish that downstream cortico-basal ganglia circuits serve a learning role that is only premotor. Furthermore, song performance-contingent optogenetic stimulation of either VTA afferent was sufficient to drive syllable-specific learning, and these learning effects were of opposite valence. Finally, functional, anatomical, and molecular studies support the idea that these evaluative afferents bidirectionally modulate VTA dopamine neurons to enable temporally precise vocal learning.
Vocalizations are fundamental to mammalian communication, but the underlying neural circuits await detailed characterization. Here, we used an intersectional genetic method to label and manipulate neurons in the midbrain periaqueductal gray (PAG) that are transiently active in male mice when they produce ultrasonic courtship vocalizations (USVs). Genetic silencing of PAG-USV neurons rendered males unable to produce USVs and impaired their ability to attract females. Conversely, activating PAG-USV neurons selectively triggered USV production, even in the absence of any female cues. Optogenetic stimulation combined with axonal tracing indicates that PAG-USV neurons gate down-stream vocal-patterning circuits. Indeed, activating PAG neurons that innervate the nucleus retroambiguus, but not those innervating the parabrachial nucleus, elicited USVs in both male and female mice. These experiments establish that a dedicated population of PAG neurons gives rise to a descending circuit necessary and sufficient for USV production while also demonstrating the communicative salience of male USVs.
Kosuke Hamaguchi (濱口航介)合作论文数Graduate School of Medicine, Kyoto University5