Manual behavior scoring is labor-intensive and subjective. Video-capable large language models (LLMs) offer a transformative, scalable solution for accelerating and standardizing neuroscience workflows. We benchmarked state-of-the-art video LLMs (Gemini 2.5 Pro, Qwen3-VL, and VideoLLaMA3) for automated behavioral segmentation and scoring of mice performing a water-reaching task. Videos of mice performing water reaching were analyzed by the LLMs. Accuracy was compared across different models and against prompt adjustments within Gemini. To assess classification determinants, video fidelity was altered through pixel interpolation and key regions blurred (paws/snout-mouth). In addition, the models were asked to describe the mouse's actions over time. Finally, an open-source rat lever-pressing dataset was utilized to validate behavioral segmentation under a few-shot learning framework, assessing the impact of visual examples on the identification of discrete action sequences. Gemini 2.5 Pro ( 0.74 ± 0.12 accuracy) and Qwen3-VL-30B ( 0.67 ± 0.13 ) exhibited the ability to classify trial outcomes. Reliable classification required a minimum pixel resolution of 0.28 mm per pixel and careful consideration of the model frame tokenization rate. Accuracy is significantly reduced upon obscuring the snout-mouth area. In 549 / 1058 of videos, Gemini 2.5 Pro also provided completely accurate frame-to-frame behavior segmentations. The inclusion of visual examples improved model detection of user-defined behaviors. Video-LLMs offer potential to accelerate neuroscience by providing scalable, objective quantification of goal-directed behaviors. By producing temporal annotations, Gemini enables fast first-pass labeling that markedly streamlines manual dataset curation.
Increasingly, experiments designed to provide practical perturbations to circuits or behavior are required for hypothesis testing in various disciplines ranging from motor learning to recovery after injury. We present the implementation and efficacy of an open-source closed-loop neurofeedback (CLNF) and closed-loop movement feedback (CLMF) system. In CLNF, we measure mm-scale cortical mesoscale activity with GCaMP6s and provide graded auditory feedback (within ∼63 ms) based on changes in dorsal-cortical activation within regions of interest (ROI) and with a specified rule. Single or dual ROIs (ROI1, ROI2) on the dorsal cortical map were selected as targets. Both motor and sensory regions supported closed-loop training in male and female mice. Mice modulated activity in rule-specific target cortical ROIs to get increasing rewards over days (RM ANOVA p=2.83e-5) and adapted to changes in ROI rules (RM ANOVA p=8.3e-10, Table 4 for different rule changes). In CLMF, feedback (within ∼67 ms) was based on tracking a specified body movement, and rewards were generated when the behavior reached a threshold. For movement training, the group that received graded auditory feedback performed significantly better (RM-ANOVA p=9.6e-7) than a control group (RM-ANOVA p=0.49) within four training days. Additionally, mice can learn a change in task rule from left forelimb to right forelimb within a day, after a brief performance drop on day 5. Offline analysis of neural data and behavioral tracking revealed changes in the overall distribution of Ca2+ fluorescence values in CLNF and body-part speed values in CLMF experiments. Increased CLMF performance was accompanied by a decrease in task latency and cortical ΔF/F0 amplitude during the task, indicating lower cortical activation as the task gets more familiar.
Huntington disease (HD) is a genetic neurodegenerative disorder characterized by progressive motor dysfunction, cognitive decline, and neuropsychiatric symptoms. Assessing early motor skill deficits in HD mouse models is challenging with traditional behavioral tasks. This study uses a home cage-based lever-pulling task, PiPaw2.0, to evaluate motor learning in 6-7 months-old zQ175 knock-in HD mice in a more naturalistic environment. In this task, mice learn to pull a lever for a water reward, with the requirement to hold the lever within a specific goal range for a required hold time. As the mice improved, the required hold time increased, thereby gradually increasing the task demands. Both wild type (WT) and zQ175 mice initially showed similar task engagement, but zQ175 mice had significant deficits in adapting to increasing hold time. The WT mice refined their strategies over time, shifting from random to more precise lever pulls, while zQ175 mice failed to make this adjustment, maintaining erratic performance. Additionally, in group-housing WT mouse lever performance benefited from peer interactions, an effect absent in zQ175 mice. Post-task neural assessments revealed that WT mice developed experience-mediated synaptic plasticity in the left striatum (contralateral to lever-pulling paw), while zQ175 mice showed no significant changes, consistent with known corticostriatal plasticity impairments in HD mouse models. In conclusion, our findings demonstrate the effectiveness of group-housed, home cage-based assessments for evaluating motor learning and adaptation in HD mouse models. This study provides insights into the motor control and adaptive learning deficits in HD, emphasizing the value of automated home cage systems in advancing neurodegenerative disease research and highlighting the importance of peer influences on performance.
Neural-network-based pose estimation models have become increasingly popular for quantitative analysis of mouse behavior, yet most recordings still use a single 2-D camera view and therefore lack the depth cues needed for accurate 3-D kinematics. Existing open-source 3-D mouse datasets for training deep-learning models cover only a narrow range of environments and do not generalize well to various laboratory settings. To overcome these limitations, we introduce PyMouse Lifter , a pipeline that automatically reconstructs 3-D mouse poses from ordinary 2-D top-view videos with minimal manual 2-D annotation. PyMouse Lifter combines (i) an anatomically realistic 3-D mouse model for automated data synthesis, (ii) a monocular depth estimation model, and (iii) a 2-D key-point estimation model, enabling accurate 3-D reconstruction (model-based 3D inference) in virtually any open-field arena without using depth or multiple camera views for reconstruction. We validate the system on multiple datasets against depth-camera ground truth and show that the lifted 3D trajectories yield improved behavior classification over 2-D data and can be implemented in real time. ### Competing Interest Statement The authors have declared no competing interest. Canadian Institute for Health Rsearch, PJT-180631, FDN-143209 Natural Sciences and Engineering Research Council, https://ror.org/01h531d29, GPIN-2022-03723
Shifts in data distribution across time can strongly affect early classification of time-series data. When decoding behavior from neural activity, early detection of behavior may help in devising corrective neural stimulation before the onset of behavior. Recurrent neural networks are common models for sequence data. However, standard recurrent neural networks are not able to handle data with temporal distributional shifts to guarantee robust classification across time. To enable the network to utilize all temporal features of the neural input data, and to enhance the memory of recurrent neural networks, this paper proposes a novel approach: recurrent neural networks with time-varying weights, here termed Time-varying recurrent neural networks. These models are able to not only predict the class of the time-sequence correctly, but also lead to accurate classification earlier in the sequence than standard recurrent neural networks, while also stabilizing gradient dynamics. This paper focuses on early sequential classification of spatially distributed neural activity across time using Time-varying recurrent neural networks applied to a variety of neural data from mice and humans, as subjects perform motor tasks. Time-varying recurrent neural networks detect self-initiated lever-pull behavior up to 6 seconds before behavior onset — 3 seconds earlier than standard recurrent neural networks. Finally, this paper explored the contribution of different brain regions on behavior classification using SHapley Additive exPlanation value, and found that the somatosensory and premotor regions play a large role in behavioral classification.
Developing high-quality training data is essential for tailoring large language models (LLMs) to specialized applications like mental health. To address privacy and legal constraints associated with real patient data, we designed a synthetic patient and interview generation framework that can be tailored to regional patient demographics. This system employs two locally run instances of Llama 3.3:70B: one as the interviewer and the other as the patient. These models produce contextually rich interview transcripts, structured by a customizable question bank, with lexical diversity similar to normal human conversation. We calculate median Distinct-1 scores of 0.44 and 0.33 for the patient and interview assistant model outputs respectively compared to 0.50 ± 0.11 as the average for 10,000 episodes of a radio program dialog. Central to this approach is the patient generation process, which begins with a locally run Llama 3.3:70B model. Given the full question bank, the model generates a detailed profile template, combining predefined variables (e.g., demographic data or specific conditions) with LLM-generated content to fill in contextual details. This hybrid method ensures that each patient profile is both diverse and realistic, providing a strong foundation for generating dynamic interactions. Demographic distributions of generated patient profiles were not significantly different from real-world population data and exhibited expected variability. Additionally, for the patient profiles we assessed LLM metrics and found an average Distinct-1 score of 0.8 (max = 1) indicating diverse word usage. By integrating detailed patient generation with dynamic interviewing, the framework produces synthetic datasets that may aid the adoption and deployment of LLMs in mental health settings.
Abstract Across cortex, a substantial component of observed neuronal activity can be explained by movement. Voluntary movements elicit cortical activity that contain elements related to both motor planning and sensory activity. However, automatic movements that do not require conscious processing, for example grooming or drinking, are to a large extent controlled by subcortical systems such as the brainstem. While the cortex may not be required to generate these automatic behaviors, it is unclear whether cortical activity represents the movements associated with them. In this work, we use a simple procedure to stimulate stereotyped grooming behaviors in the head-fixed mouse while measuring neuronal function across the dorsal cortex. We find specific cortical representations of grooming both at a mesoscale network-level and single-cell resolution. Mesoscale cortical activation was most prominent at the onset of grooming episodes, and declined to baseline levels despite continuous engagement in the behavior. Further stratification of grooming component movements revealed that more directed and unilateral grooming movements had greater cortical associated responses than more stereotyped bilateral movements. These findings help to frame the impact of forms of automatic movements on large scale cortical ensembles and suggest they engage specific and transient cellular and regional cortical ensembles.
We present the implementation and efficacy of an open-source closed-loop neurofeedback (CLNF) and closed-loop movement feedback (CLMF) system. In CLNF, we measure mm-scale cortical mesoscale activity with GCaMP6s and provide graded auditory feedback (within ∼50 ms) based on changes in dorsal-cortical activation within regions of interest (ROI) and with a specified rule. Single or dual ROIs (ROI1, ROI2) on the dorsal cortical map were selected as targets. Both motor and sensory regions supported closed-loop training in male and female mice. Mice modulated activity in rule-specific target cortical ROIs to get increasing rewards over days (RM ANOVA p=2.83e-5) and adapted to changes in ROI rules (RM ANOVA p=8.3e-10, Table 4 for different rule changes). In CLMF, feedback was based on tracking a specified body movement, and rewards were generated when the behavior reached a threshold. For movement training, the group that received graded auditory feedback performed significantly better (RM-ANOVA p=9.6e-7) than a control group (RM-ANOVA p=0.49) within four training days. Additionally, mice can learn a change in task rule from left forelimb to right forelimb within a day, after a brief performance drop on day 5. Offline analysis of neural data and behavioral tracking revealed changes in the overall distribution of ΔF/F 0 values in CLNF and body-part speed values in CLMF experiments. Increased CLMF performance was accompanied by a decrease in task latency and cortical ΔF/F 0 amplitude during the task, indicating lower cortical activation as the task gets more familiar.
The Allen Brain Institute (ABI) and the International Brain Laboratory (IBL) have produced large high quality open behavioural and electrophysiological datasets collected from behaving mice using neuropixels probes. Shared data from these projects are in the form of spike times, raw video footage, scored behaviour, but also local field potentials (LFP). These probes often pass through hippocampus while simultaneously recording from a number of other regions, providing the opportunity to evaluate how hippocampal LFP features during synchronized high-frequency bursts known as sharp-wave ripples (SWRs) impact behavioural task variables or spiking activity on other recorded contacts. Currently, there are no data standards or file formats for sharing SWRs. Here we present the SWR data from the ABI and IBL datasets in a sharable format, which integrates with their APIs. We have extracted, curated using field standards, and shared over 967,431 SWR events from 210 mice from these datasets as well as the code used to process them.
Significance: Behavior scoring is labor-intensive and subjective, introducing variability in results. Large Language Models (LLMs) capable of video understanding offer a transformative solution to manual scoring, crucial for accelerating and standardizing neuroscience workflows. Aim: We sought to benchmark state-of-the-art video LLMs (Gemini 2.5 Pro, Qwen3-VL, and VideoLLaMA3) for automated behavioural segmentation and scoring of mice performing a water-reaching task. Approach: Videos of mice performing water reaching from the front view were analysed by the LLMs. Accuracy was compared across different models and against prompt adjustments within Gemini. To assess classification determinants, video fidelity was altered through pixel interpolation and key regions blurred (paws/snout-mouth). In addition, the models were asked to describe the mouse's actions over time. Results: Gemini 2.5 Pro (Mean: 0.74, SD: 0.12) and Qwen3-VL-30B (Mean: 0.67, SD: 0.13) exhibited ability to classify trial outcomes. Reliable classification required a minimum pixel resolution of 0.28 mm per pixel. Accuracy is significantly reduced upon obscuring the snout-mouth area. In 549/1058 of videos, Gemini 2.5 Pro also provided completely accurate frame-to- frame behaviour segmentations. Conclusions: Video-LLMs offer potential to accelerate neuroscience by providing scalable, objective quantification of goal-directed behaviors. By producing temporal annotations, Gemini enables fast first-pass labelling that markedly streamlines manual dataset curation. ### Competing Interest Statement The authors have declared no competing interest. Canadian Institutes of Health Research, PJT-180631 Natural Sciences and Engineering Research Council, GPIN-2022-03723
BACKGROUND:Huntington disease (HD) is a neurodegenerative disorder with complex motor and behavioural manifestations. The Q175 knock-in mouse model of HD has gained recent popularity as a genetically accurate model of the human disease. However, behavioural phenotypes are often subtle and progress slowly in this model. Here, we have implemented machine-learning algorithms to investigate behaviour in the Q175 model and compare differences between sexes and disease stages. We explore distinct behavioural patterns and motor functions in open field, rotarod, water T-maze, and home cage lever-pulling tasks.RESULTS:In the open field, we observed habituation deficits in two versions of the Q175 model (zQ175dn and Q175FDN, on two different background strains), and using B-SOiD, an advanced machine learning approach, we found altered performance of rearing in male manifest zQ175dn mice. Notably, we found that weight had a considerable effect on performance of accelerating rotarod and water T-maze tasks and controlled for this by normalizing for weight. Manifest zQ175dn mice displayed a deficit in accelerating rotarod (after weight normalization), as well as changes to paw kinematics specific to males. Our water T-maze experiments revealed response learning deficits in manifest zQ175dn mice and reversal learning deficits in premanifest male zQ175dn mice; further analysis using PyMouseTracks software allowed us to characterize new behavioural features in this task, including time at decision point and number of accelerations. In a home cage-based lever-pulling assessment, we found significant learning deficits in male manifest zQ175dn mice. A subset of mice also underwent electrophysiology slice experiments, revealing a reduced spontaneous excitatory event frequency in male manifest zQ175dn mice.CONCLUSIONS:Our study uncovered several behavioural changes in Q175 mice that differed by sex, age, and strain. Our results highlight the impact of weight and experimental protocol on behavioural results, and the utility of machine learning tools to examine behaviour in more detailed ways than was previously possible. Specifically, this work provides the field with an updated overview of behavioural impairments in this model of HD, as well as novel techniques for dissecting behaviour in the open field, accelerating rotarod, and T-maze tasks.
Traumatic brain injury (TBI) is the leading cause of death in young people and can cause cognitive and motor dysfunction and disruptions in functional connectivity between brain regions. In human TBI patients and rodent models of TBI, functional connectivity is decreased after injury. Recovery of connectivity after TBI is associated with improved cognition and memory, suggesting an important link between connectivity and functional outcome. We examined widespread alterations in functional connectivity following TBI using simultaneous widefield mesoscale GCaMP7c calcium imaging and electrocorticography (ECoG) in mice injured using the controlled cortical impact (CCI) model of TBI. Combining CCI with widefield cortical imaging provides us with unprecedented access to characterize network connectivity changes throughout the entire injured cortex over time. Our data demonstrate that CCI profoundly disrupts functional connectivity immediately after injury, followed by partial recovery over 3 weeks. Examining discrete periods of locomotion and stillness reveals that CCI alters functional connectivity and reduces theta power only during periods of behavioral stillness. Together, these findings demonstrate that TBI causes dynamic, behavioral state-dependent changes in functional connectivity and ECoG activity across the cortex.
The cortex and cerebellum form multi-synaptic reciprocal connections. We investigate the functional connectivity between single spiking cerebellar neurons and the population activity of the mouse dorsal cortex using mesoscale imaging. Cortical representations of individual cerebellar neurons vary significantly across different brain states but are drawn from a common set of cortical networks. These cortical-cerebellar connectivity features are observed in mossy fibers and Purkinje cells as well as neurons in different cerebellar lobules, albeit with variations across cell types and regions. Complex spikes of Purkinje cells preferably associate with the sensorimotor cortex, whereas simple spikes display more diverse cortical connectivity patterns. The spontaneous functional connectivity patterns align with cerebellar neurons’ functional responses to external stimuli in a modality-specific manner. The tuning properties of subsets of cerebellar neurons differ between anesthesia and awake states, mirrored by state-dependent changes in their long-range functional connectivity patterns with mesoscale cortical activity.
The availability of large-scale neuronal population datasets necessitates new methods to model population dynamics and extract interpretable, scientifically translatable insights. Existing deep learning methods often overlook the biological mechanisms underlying population activity and thus exhibit suboptimal performance with neuronal data and provide little to no interpretable information about neurons and their interactions. In response, we introduce SynapsNet, a novel deep-learning framework that effectively models population dynamics and functional interactions between neurons. Within this biologically realistic framework, each neuron, characterized by a latent embedding, sends and receives currents through directed connections. A shared decoder uses the input current, previous neuronal activity, neuron embedding, and behavioral data to predict the population activity in the next time step. Unlike common sequential models that treat population activity as a multichannel time series, SynapsNet applies its decoder to each neuron (channel) individually, with the learnable functional connectivity serving as the sole pathway for information flow between neurons. Our experiments, conducted on mouse cortical activity from publicly available datasets and recorded using the two most common population recording modalities (Ca imaging and Neuropixels) across three distinct tasks, demonstrate that SynapsNet consistently outperforms existing models in forecasting population activity. Additionally, our experiments on both real and synthetic data showed that SynapsNet accurately learns functional connectivity that reveals predictive interactions between neurons.
Academic departments, research clusters and evaluators analyze author and citation data to measure research impact and to support strategic planning. We created Scholar Metrics Scraper (SMS) to automate the retrieval of bibliometric data for a group of researchers. The project contains Jupyter notebooks that take a list of researchers as an input and exports a CSV file of citation metrics from Google Scholar (GS) to visualize the group's impact and collaboration. A series of graph outputs are also available. SMS is an open solution for automating the retrieval and visualization of citation data.
We present a cost-effective, compact foot-print, and open-source Raspberry Pi-based widefield imaging system. The compact nature allows the system to be used for close-proximity dual-brain cortical mesoscale functional-imaging to simultaneously observe activity in two head-fixed animals in a staged social touch-like interaction. We provide all schematics, code, and protocols for a rail system where head-fixed mice are brought together to a distance where the macrovibrissae of each mouse make contact. Cortical neuronal functional signals (GCaMP6s; genetically encoded Ca2+ sensor) were recorded from both mice simultaneously before, during, and after the social contact period. When the mice were together, we observed bouts of mutual whisking and cross-mouse correlated cortical activity across the cortex. Correlations were not observed in trial-shuffled mouse pairs, suggesting that correlated activity was specific to individual interactions. Whisking-related cortical signals were observed during the period where mice were together (closest contact). The effects of social stimulus presentation extend outside of regions associated with mutual touch and have global synchronizing effects on cortical activity.
PyMouseTracks (PMT) is a scalable and customizable computer vision and radio frequency identification (RFID)-based system for multiple rodent tracking and behavior assessment that can be set up within minutes in any user-defined arena at minimal cost. PMT is composed of the online Raspberry Pi (RPi)-based video and RFID acquisition with subsequent offline analysis tools. The system is capable of tracking up to six mice in experiments ranging from minutes to days. PMT maintained a minimum of 88% detections tracked with an overall accuracy.85% when compared with manual validation of videos containing one to four mice in a modified home-cage. As expected, chronic recording in home-cage revealed diurnal activity patterns. In open- field, it was observed that novel noncagemate mouse pairs exhibit more similarity in travel trajectory patterns than cagemate pairs over a 10-min period. Therefore, shared features within travel trajectories between ani- mals may be a measure of sociability that has not been previously reported. Moreover, PMT can interface with open-source packages such as DeepLabCut and Traja for pose estimation and travel trajectory analysis, respectively. In combination with Traja, PMT resolved motor deficits exhibited in stroke animals. Overall, we present an affordable, open-sourced, and customizable/scalable mouse behavior recording and analysis system.
Accurate capture of animal behavior and posture requires the use of multiple cameras to reconstruct three-dimensional (3D) representations. Typically, a paper ChArUco (or checker) board works well for correcting distortion and calibrating for 3D reconstruction in stereo vision. However, measuring the error in two-dimensional (2D) is also prone to bias related to the placement of the 2D board in 3D. We proposed a procedure as a visual way of validating camera placement, and it also can provide some guidance about the positioning of cameras and potential advantages of using multiple cameras. We propose the use of a 3D printable test object for validating multi-camera surround-view calibration in small animal video capture arenas. The proposed 3D printed object has no bias to a particular dimension and is designed to minimize occlusions. The use of the calibrated test object provided an estimate of 3D reconstruction accuracy. The approach reveals that for complex specimens such as mice, some view angles will be more important for accurate capture of keypoints. Our method ensures accurate 3D camera calibration for surround image capture of laboratory mice and other specimens.