Neural circuit function is shaped both by the cell types that comprise the circuit and the connections between them1. Neural cell types have previously been defined by morphology2,3, electrophysiology4, transcriptomic expression5,6, connectivity7-9 or a combination of such modalities10-12. The Patch-seq technique enables the characterization of morphology, electrophysiology and transcriptomic properties from individual cells13-15. These properties were integrated to define 28 inhibitory, morpho-electric-transcriptomic (MET) cell types in mouse visual cortex16, which do not include synaptic connectivity. Conversely, large-scale electron microscopy (EM) enables morphological reconstruction and a near-complete description of a neuron's local synaptic connectivity, but does not include transcriptomic or electrophysiological information. Here, we leveraged morphological information from Patch-seq to predict the transcriptomically defined cell subclass and/or MET-type of inhibitory neurons within a large-scale EM dataset. We further analysed Martinotti cells-a somatostatin (Sst)-positive17 morphological cell type18,19-which were classified successfully into Sst MET-types with distinct axon myelination and synaptic output connectivity patterns. We demonstrate that morphological features can be used to link cell types across experimental modalities, enabling further comparison of connectivity to gene expression and electrophysiology. We observe unique connectivity rules for predicted Sst cell types.
Despite significant progress in characterizing neocortical cell types, a complete understanding of the synaptic connections of individual excitatory cells remains elusive. This study investigates the connectivity of mouse visual cortex thick tufted layer 5 pyramidal cells, also known as extratelencephalic neurons (L5-ETns), using a 1 mm3 publicly available electron microscopy dataset. The analysis reveals that, in their immediate vicinity, L5-ETns primarily establish connections with a group of inhibitory cell types, which, in turn, specifically target the L5-ETns back. The most common excitatory targets of L5-ETns are layer 5 intertelencephalic neurons (L5-ITns) and layer 6 (L6) pyramidal cells, whereas synapses with other L5-ETns are less common. When L5-ETns extend their axons to other cortical regions, they tend to connect more with excitatory cells. Our results highlight a circuit motif where a subclass of excitatory cells forms a subcircuit with specific inhibitory cell types. This is achieved using a publicly available, automated approach for synapse recognition and automated cell typing, offering a framework for exploring the connectivity of other neuron types.
Effective personalization of LLMs is critical for a broad range of user-interfacing applications such as virtual assistants and content curation. Inspired by the strong in-context capabilities of LLMs, we propose few-shot preference optimization (FSPO), an algorithm for LLM personalization that reframes reward modeling as a meta-learning problem. Under FSPO, an LLM learns to quickly infer a personalized reward function for a user via a few labeled preferences. FSPO also utilizes user description rationalization (RAT) to encourage better reward modeling and instruction following, recovering performance with the oracle user description. Since real-world preference data is challenging to collect at scale, we propose careful design choices to construct synthetic preference datasets for personalization, generating over 1M synthetic personalized preferences using publicly available LLMs. To successfully transfer from synthetic data to real users, we find it crucial for the data to exhibit both high diversity and coherent, self-consistent structure. We evaluate FSPO on personalized open-ended generation for up to 1,500 synthetic users across three domains: movie reviews, education, and open-ended question answering. We also run a controlled human study. Overall, FSPO achieves an 87
Advances in Electron Microscopy, image segmentation and computational infrastructure have given rise to large-scale and richly annotated connectomic datasets which are increasingly shared across communities. To enable collaboration, users need to be able to concurrently create new annotations and correct errors in the automated segmentation by proofreading. In large datasets, every proofreading edit relabels cell identities of millions of voxels and thousands of annotations like synapses. For analysis, users require immediate and reproducible access to this constantly changing and expanding data landscape. Here, we present the Connectome Annotation Versioning Engine (CAVE), a computational infrastructure for immediate and reproducible connectome analysis in up-to petascale datasets (~1mm3) while proofreading and annotating is ongoing. For segmentation, CAVE provides a distributed proofreading infrastructure for continuous versioning of large reconstructions. Annotations in CAVE are defined by locations such that they can be quickly assigned to the underlying segment which enables fast analysis queries of CAVE's data for arbitrary time points. CAVE supports schematized, extensible annotations, so that researchers can readily design novel annotation types. CAVE is already used for many connectomics datasets, including the largest datasets available to date.
Mammalian cortex features a vast diversity of neuronal cell types, each with characteristic anatomical, molecular and functional properties1. Synaptic connectivity shapes how each cell type participates in the cortical circuit, but mapping connectivity rules at the resolution of distinct cell types remains difficult. Here we used millimetre-scale volumetric electron microscopy2 to investigate the connectivity of all inhibitory neurons across a densely segmented neuronal population of 1,352 cells spanning all layers of mouse visual cortex, producing a wiring diagram of inhibition with more than 70,000 synapses. Inspired by classical neuroanatomy, we classified inhibitory neurons based on targeting of dendritic compartments and developed an excitatory neuron classification based on dendritic reconstructions with whole-cell maps of synaptic input. Single-cell connectivity showed a class of disinhibitory specialist that targets basket cells. Analysis of inhibitory connectivity onto excitatory neurons found widespread specificity, with many interneurons exhibiting differential targeting of spatially intermingled subpopulations. Inhibitory targeting was organized into 'motif groups', diverse sets of cells that collectively target both perisomatic and dendritic compartments of the same excitatory targets. Collectively, our analysis identified new organizing principles for cortical inhibition and will serve as a foundation for linking contemporary multimodal neuronal atlases with the cortical wiring diagram.
We are in the era of millimetre-scale electron microscopy volumes collected at nanometre resolution1,2. Dense reconstruction of cellular compartments in these electron microscopy volumes has been enabled by recent advances in machine learning3-6. Automated segmentation methods produce exceptionally accurate reconstructions of cells, but post hoc proofreading is still required to generate large connectomes that are free of merge and split errors. The elaborate 3D meshes of neurons in these volumes contain detailed morphological information at multiple scales, from the diameter, shape and branching patterns of axons and dendrites, down to the fine-scale structure of dendritic spines. However, extracting these features can require substantial effort to piece together existing tools into custom workflows. Here, building on existing open source software for mesh manipulation, we present Neural Decomposition (NEURD), a software package that decomposes meshed neurons into compact and extensively annotated graph representations. With these feature-rich graphs, we automate a variety of tasks such as state-of-the-art automated proofreading of merge errors, cell classification, spine detection, axonal-dendritic proximities and other annotations. These features enable many downstream analyses of neural morphology and connectivity, making these massive and complex datasets more accessible to neuroscience researchers.
Mammalian neocortex contains a highly diverse set of cell types. These cell types have been mapped systematically using a variety of molecular, electrophysiological and morphological approaches 1–4 . Each modality offers new perspectives on the variation of biological processes underlying cell-type specialization. Cellular-scale electron microscopy provides dense ultrastructural examination and an unbiased perspective on the subcellular organization of brain cells, including their synaptic connectivity and nanometre-scale morphology. In data that contain tens of thousands of neurons, most of which have incomplete reconstructions, identifying cell types becomes a clear challenge for analysis 5 . Here, to address this challenge, we present a systematic survey of the somatic region of all cells in a cubic millimetre of cortex using quantitative features obtained from electron microscopy. This analysis demonstrates that the perisomatic region is sufficient to identify cell types, including types defined primarily on the basis of their connectivity patterns. We then describe how this classification facilitates cell-type-specific connectivity characterization and locating cells with rare connectivity patterns in the dataset.
Neurons in the neocortex exhibit astonishing morphological diversity, which is critical for properly wiring neural circuits and giving neurons their functional properties. However, the organizational principles underlying this morphological diversity remain an open question. Here, we took a data-driven approach using graph-based machine learning methods to obtain a low-dimensional morphological "bar code" describing more than 30,000 excitatory neurons in mouse visual areas V1, AL, and RL that were reconstructed from the millimeter scale MICrONS serial-section electron microscopy volume. Contrary to previous classifications into discrete morphological types (m-types), our data-driven approach suggests that the morphological landscape of cortical excitatory neurons is better described as a continuum, with a few notable exceptions in layers 5 and 6. Dendritic morphologies in layers 2-3 exhibited a trend towards a decreasing width of the dendritic arbor and a smaller tuft with increasing cortical depth. Inter-area differences were most evident in layer 4, where V1 contained more atufted neurons than higher visual areas. Moreover, we discovered neurons in V1 on the border to layer 5, which avoided deeper layers with their dendrites. In summary, we suggest that excitatory neurons' morphological diversity is better understood by considering axes of variation than using distinct m-types.
The diversity of contexts in which large language models (LLMs) are deployed requires the ability to modify or customize default model behaviors to incorporate nuanced requirements and preferences. A convenient interface to specify such model adjustments is high-level verbal feedback, such as "Don't use emojis when drafting emails to my boss." However, while writing high-level feedback is far simpler than collecting annotations for reinforcement learning from human feedback (RLHF), we find that simply prompting a model with such feedback leads to overgeneralization of the feedback to contexts where it is not relevant. We study the problem of incorporating verbal feedback without such overgeneralization, inspiring a new method Contextualized Critiques with Constrained Preference Optimization (C3PO). C3PO uses a piece of high-level feedback to generate a small synthetic preference dataset specifying how the feedback should (and should not) be applied. It then fine-tunes the model in accordance with the synthetic preference data while minimizing the divergence from the original model for prompts where the feedback does not apply. Our experimental results indicate that our approach effectively applies verbal feedback to relevant scenarios while preserving existing behaviors for other contexts. For both human- and GPT-4-generated high-level feedback, C3PO effectively adheres to the given feedback comparably to in-context baselines while reducing overgeneralization by 30%.
Widely used language models (LMs) are typically built by scaling up a two-stage training pipeline: a pre-training stage that uses a very large, diverse dataset of text and a fine-tuning (sometimes, 'alignment') stage that uses targeted examples or other specifications of desired behaviors. While it has been hypothesized that knowledge and skills come from pre-training, and fine-tuning mostly filters this knowledge and skillset, this intuition has not been extensively tested. To aid in doing so, we introduce a novel technique for decoupling the knowledge and skills gained in these two stages, enabling a direct answer to the question, "What would happen if we combined the knowledge learned by a large model during pre-training with the knowledge learned by a small model during fine-tuning (or vice versa)?" Using an RL-based framework derived from recent developments in learning from human preferences, we introduce emulated fine-tuning (EFT), a principled and practical method for sampling from a distribution that approximates (or 'emulates') the result of pre-training and fine-tuning at different scales. Our experiments with EFT show that scaling up fine-tuning tends to improve helpfulness, while scaling up pre-training tends to improve factuality. Beyond decoupling scale, we show that EFT enables test-time adjustment of competing behavioral traits like helpfulness and harmlessness without additional training. Finally, a special case of emulated fine-tuning, which we call LM up-scaling, avoids resource-intensive fine-tuning of large pre-trained models by ensembling them with small fine-tuned models, essentially emulating the result of fine-tuning the large pre-trained model. Up-scaling consistently improves helpfulness and factuality of instruction-following models in the Llama, Llama-2, and Falcon families, without additional hyperparameters or training.
The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In particular, our models can reason about our safety policies in context when responding to potentially unsafe prompts, through deliberative alignment. This leads to state-of-the-art performance on certain benchmarks for risks such as generating illicit advice, choosing stereotyped responses, and succumbing to known jailbreaks. Training models to incorporate a chain of thought before answering has the potential to unlock substantial benefits, while also increasing potential risks that stem from heightened intelligence. Our results underscore the need for building robust alignment methods, extensively stress-testing their efficacy, and maintaining meticulous risk management protocols. This report outlines the safety work carried out for the OpenAI o1 and OpenAI o1-mini models, including safety evaluations, external red teaming, and Preparedness Framework evaluations.
The fluency and creativity of large pre-trained language models (LLMs) have led to their widespread use, sometimes even as a replacement for traditional search engines. Yet language models are prone to making convincing but factually inaccurate claims, often referred to as 'hallucinations.' These errors can inadvertently spread misinformation or harmfully perpetuate misconceptions. Further, manual fact-checking of model responses is a time-consuming process, making human factuality labels expensive to acquire. In this work, we fine-tune language models to be more factual, without human labeling and targeting more open-ended generation settings than past work. We leverage two key recent innovations in NLP to do so. First, several recent works have proposed methods for judging the factuality of open-ended text by measuring consistency with an external knowledge base or simply a large model's confidence scores. Second, the direct preference optimization algorithm enables straightforward fine-tuning of language models on objectives other than supervised imitation, using a preference ranking over possible model responses. We show that learning from automatically generated factuality preference rankings, generated either through existing retrieval systems or our novel retrieval-free approach, significantly improves the factuality (percent of generated claims that are correct) of Llama-2 on held-out topics compared with RLHF or decoding strategies targeted at factuality. At 7B scale, compared to Llama-2-chat, we observe 58% and 40% reduction in factual error rate when generating biographies and answering medical questions, respectively.
Reward models trained on aggregate preferences often fail to capture individual users' values, but existing adaptation methods such as fine-tuning or long-context conditioning are too costly for real-time personalization. We propose Hypothesis Reweighting (HyRe), which enables real-time personalization by reweighting ensemble members using just 1-5 labeled examples from the target user or domain. Our method builds on the empirical observation that when different heads capture different valid interpretations of preference data, reweighting them can substantially outperform uniform averaging. HyRe trains a single network with multiple prediction heads that capture different valid interpretations of preference data, then uses a Bayesian update to upweight the heads that best match the target user's preferences. This requires only a single forward pass with negligible (<1
The effectiveness of large language models (LLMs) is not only measured by their ability to generate accurate outputs but also by their calibration—how well their confidence scores reflect the probability of their outputs being correct. While unsupervised pre-training has been shown to yield LLMs with well-calibrated conditional probabilities, recent studies have shown that after fine-tuning with reinforcement learning from human feedback (RLHF), the calibration of these models degrades significantly. In this work, we introduce Adaptive Temperature Scaling (ATS), a post-hoc calibration method that predicts a temperature scaling parameter for each token prediction. The predicted temperature values adapt based on token-level features and are fit over a standard supervised fine-tuning (SFT) dataset. The adaptive nature of ATS addresses the varying degrees of calibration shift that can occur after RLHF fine-tuning. ATS improves calibration by over 10-50% across three downstream natural language evaluation benchmarks compared to prior calibration methods and does not impede performance improvements from RLHF.
The reconstruction of neural circuits from serial section electron microscopy (ssEM) images is being accelerated by automatic image segmentation methods. Segmentation accuracy is often limited by the preceding step of aligning 2D section images to create a 3D image stack. Precise and robust alignment in the presence of image artifacts is challenging, especially as datasets are attaining the petascale. We present a computational pipeline for aligning ssEM images with several key elements. Self-supervised convolutional nets are trained via metric learning to encode and align image pairs, and they are used to initialize iterative fine-tuning of alignment. A procedure called vector voting increases robustness to image artifacts or missing image data. For speedup the series is divided into blocks that are distributed to computational workers for alignment. The blocks are aligned to each other by composing transformations with decay, which achieves a global alignment without resorting to a time-consuming global optimization. We apply our pipeline to a whole fly brain dataset, and show improved accuracy relative to prior state of the art. We also demonstrate that our pipeline scales to a cubic millimeter of mouse visual cortex. Our pipeline is publicly available through two open source Python packages.
Reinforcement learning with AI feedback (RLAIF) is a popular paradigm for improving the instruction-following abilities of powerful pre-trained language models. RLAIF first performs supervised fine-tuning (SFT) using demonstrations from a teacher model and then further fine-tunes the model with reinforcement learning (RL), using feedback from a critic model. While recent popular open-source models have demonstrated substantial improvements in performance from the RL step, in this paper we question whether the complexity of this RL step is truly warranted for AI feedback. We show that the improvements of the RL step are virtually entirely due to the widespread practice of using a weaker teacher model (e.g. GPT-3.5) for SFT data collection than the critic (e.g., GPT-4) used for AI feedback generation. Specifically, we show that simple supervised fine-tuning with GPT-4 as the teacher outperforms existing RLAIF pipelines. More generally, we find that the gains from RLAIF vary substantially across base model families, test-time evaluation protocols, and critic models. Finally, we provide a mechanistic explanation for when SFT may outperform the full two-step RLAIF pipeline as well as suggestions for making RLAIF maximally useful in practice.
Mammalian cortex features a vast diversity of neuronal cell types, each with characteristic anatomical, molecular and functional properties. Synaptic connectivity powerfully shapes how each cell type participates in the cortical circuit, but mapping connectivity rules at the resolution of distinct cell types remains difficult. Here, we used millimeter-scale volumetric electron microscopy1to investigate the connectivity of all inhibitory neurons across a densely-segmented neuronal population of 1352 cells spanning all layers of mouse visual cortex, producing a wiring diagram of inhibitory connections with more than 70,000 synapses. Taking a data-driven approach inspired by classical neuroanatomy, we classified inhibitory neurons based on the relative targeting of dendritic compartments and other inhibitory cells and developed a novel classification of excitatory neurons based on the morphological and synaptic input properties. The synaptic connectivity between inhibitory cells revealed a novel class of disinhibitory specialist targeting basket cells, in addition to familiar subclasses. Analysis of the inhibitory connectivity onto excitatory neurons found widespread specificity, with many interneurons exhibiting differential targeting of certain subpopulations spatially intermingled with other potential targets. Inhibitory targeting was organized into “motif groups,” diverse sets of cells that collectively target both perisomatic and dendritic compartments of the same excitatory targets. Collectively, our analysis identified new organizing principles for cortical inhibition and will serve as a foundation for linking modern multimodal neuronal atlases with the cortical wiring diagram.
Understanding the relationship between circuit connectivity and function is crucial for uncovering how the brain implements computation. In the mouse primary visual cortex (V1), excitatory neurons with similar response properties are more likely to be synaptically connected, but previous studies have been limited to within V1, leaving much unknown about broader connectivity rules. In this study, we leverage the millimeter-scale MICrONS dataset to analyze synaptic connectivity and functional properties of individual neurons across cortical layers and areas. Our results reveal that neurons with similar responses are preferentially connected both within and across layers and areas - including feedback connections - suggesting the universality of the 'like-to-like' connectivity across the visual hierarchy. Using a validated digital twin model, we separated neuronal tuning into feature (what neurons respond to) and spatial (receptive field location) components. We found that only the feature component predicts fine-scale synaptic connections, beyond what could be explained by the physical proximity of axons and dendrites. We also found a higher-order rule where postsynaptic neuron cohorts downstream of individual presynaptic cells show greater functional similarity than predicted by a pairwise like-to-like rule. Notably, recurrent neural networks (RNNs) trained on a simple classification task develop connectivity patterns mirroring both pairwise and higher-order rules, with magnitude similar to those in the MICrONS data. Lesion studies in these RNNs reveal that disrupting 'like-to-like' connections has a significantly greater impact on performance compared to lesions of random connections. These findings suggest that these connectivity principles may play a functional role in sensory processing and learning, highlighting shared principles between biological and artificial systems.
The fluency and general applicability of large language models (LLMs) has motivated significant interest in detecting whether a piece of text was written by a language model. While both academic and commercial detectors have been deployed in some settings, particularly education, other research has highlighted the fragility of these systems. In this paper, we demonstrate a data-efficient attack that fine-tunes language models to confuse existing detectors, leveraging recent developments in reinforcement learning of language models. We use the 'human-ness' score (often just a log probability) of various open-source and commercial detectors as a reward function for reinforcement learning, subject to a KL-divergence constraint that the resulting model does not differ significantly from the original. For a 7B parameter Llama-2 model, fine-tuning for under a day reduces the AUROC of the OpenAI RoBERTa-Large detector from 0.84 to 0.62, while perplexity on OpenWebText increases from 8.7 to only 9.0; with a larger perplexity budget, we reduce AUROC to 0.30 (worse than random), with a perplexity increase to 9.9. Similar to traditional adversarial attacks, we find that this increase in 'detector evasion' generalizes to other detectors not used during training. In light of our empirical results, we advise against continued reliance on LLM-generated text detectors.
Connections between neurons can be mapped by acquiring and analyzing electron microscopic (EM) brain images. In recent years, this approach has been applied to chunks of brains to reconstruct local connectivity maps that are highly informative, yet inadequate for understanding brain function more globally. Here, we present the first neuronal wiring diagram of a whole adult brain, containing 5×10 7 chemical synapses between ∼130,000 neurons reconstructed from a female Drosophila melanogaster . The resource also incorporates annotations of cell classes and types, nerves, hemilineages, and predictions of neurotransmitter identities. Data products are available by download, programmatic access, and interactive browsing and made interoperable with other fly data resources. We show how to derive a projectome, a map of projections between regions, from the connectome. We demonstrate the tracing of synaptic pathways and the analysis of information flow from inputs (sensory and ascending neurons) to outputs (motor, endocrine, and descending neurons), across both hemispheres, and between the central brain and the optic lobes. Tracing from a subset of photoreceptors all the way to descending motor pathways illustrates how structure can uncover putative circuit mechanisms underlying sensorimotor behaviors. The technologies and open ecosystem of the FlyWire Consortium set the stage for future large-scale connectome projects in other species.