A Last-Layer Ensemble (LLE), $K$ linear units on one shared frozen feature map, is an efficient single-pass approach to the disagreement-based epistemic uncertainty for out-of-distribution (OOD) detection. Its weakness is that members share the backbone gradient and can converge toward the same function, collapsing the inter-member diversity the signal depends on. Whether last-layer diversity can be restored, and what mitigates the collapse, is an open question. The weight-orthonormality defining Orthonormal Certificates (OC), the weight-orthonormal special case of the LLE, is only an indirect correction; it decorrelates the weights of the members, not their predictions. Here, we instead target the collapse directly in function space, with a Covariance Last-Layer Ensemble (cov-LLE) that places a direct covariance penalty on member activations. Cov-LLE restores the function-space diversity that weight-orthonormality cannot, and at matched $K$ recovers much of the diversity and calibration of a deep ensemble at $1\times$ backbone cost (in-distribution prediction variance $0.05\!\to\!9.3$ vs.\ $22.1$ ($\times10^{-3}$), and ECE $0.135\!\to\!0.090$ vs.\ $0.035$, for a $K\times$-cost deep ensemble), at no cost to accuracy. Viewing OC as a last-layer ensemble also organizes detectors into a two-axis taxonomy (by how their units are trained and how their outputs are scored) and exposes the OC score as a magnitude, motivating a scale-invariant, label-free direction score that repairs its near-OOD failure, adding $+0.16$ to $+0.18$ ROC AUC on every backbone.
Automating the annotation of benthic imagery (i.e., images of the seafloor and its associated organisms, habitats, and geological features) is critical for monitoring rapidly changing ocean ecosystems. Deep learning approaches have succeeded in this purpose; however, consistent annotation remains challenging due to ambiguous seafloor images, potential inter-user annotation disagreements, and out-of-distribution samples. Marine scientists implementing deep learning models often obtain predictions based on one-hot representations trained using a cross-entropy loss objective with softmax normalization, resulting with a single set of model parameters. While efficient, this approach may lead to overconfident predictions for context-challenging datasets, raising reliability concerns that present risks for downstream tasks such as benthic habitat mapping and marine spatial planning. In this study, we investigated classification uncertainty as a tool to improve the labeling of benthic habitat imagery. We developed a framework for two challenging sub-datasets of the recently publicly available BenthicNet dataset using Bayesian neural networks, Monte Carlo dropout inference sampling, and a proposed single last-layer committee machine. This approach resulted with a > 95% reduction of network parameters to obtain per-sample uncertainties while obtaining near-identical performance compared to computationally more expensive strategies such as Bayesian neural networks, Monte Carlo dropout, and deep ensembles. The method proposed in this research provides a strategy for obtaining prioritized lists of uncertain samples for human-in-the-loop interventions to identify ambiguous, mislabeled, out-of-distribution, and/or difficult images for enhancing existing annotation tools for benthic mapping and other applications.
In hierarchical multi-label classification, a persistent challenge is enabling model predictions to reach deeper levels of the hierarchy for more detailed or fine-grained classifications. This difficulty partly arises from the natural rarity of certain classes (or hierarchical nodes) and the hierarchical constraint that ensures child nodes are almost always less frequent than their parents. To address this, we propose a weighted loss objective for neural networks that combines node-wise imbalance weighting with focal weighting components, the latter leveraging modern quantification of ensemble uncertainties. By emphasizing rare nodes rather than rare observations (data points), and focusing on uncertain nodes for each model output distribution during training, we observe improvements in recall by up to a factor of five on benchmark datasets, along with statistically significant gains in F_1 score. We also show our approach aids convolutional networks on challenging tasks, as in situations with suboptimal encoders or limited data.
The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.
Machine learning applications require fast and reliable per-sample uncertainty estimation. A common approach is to use predictive distributions from Bayesian or approximation methods and additively decompose uncertainty into aleatoric (i.e., data-related) and epistemic (i.e., model-related) components. However, additive decomposition has recently been questioned, with evidence that it breaks down when using finite-ensemble sampling and/or mismatched predictive distributions. This paper introduces Variance-Gated Ensembles (VGE), an intuitive, differentiable framework that injects epistemic sensitivity via a signal-to-noise gate computed from ensemble statistics. VGE provides: (i) a Variance-Gated Margin Uncertainty (VGMU) score that couples decision margins with ensemble predictive variance; and (ii) a Variance-Gated Normalization (VGN) layer that generalizes the variance-gated uncertainty mechanism to training via per-class, learnable normalization of ensemble member probabilities. We derive closed-form vector-Jacobian products enabling end-to-end training through ensemble sample mean and variance. VGE matches or exceeds state-of-the-art information-theoretic baselines while remaining computationally efficient. As a result, VGE provides a practical and scalable approach to epistemic-aware uncertainty estimation in ensemble models. An open-source implementation is available at: https://github.com/nextdevai/vge.
The ongoing discussion of AGI is often mixed with claims about the ability of models to mimic brain functions. This position paper argues that mimicking human behavior at some level is not enough to understand brain functions. We point specifically to system-level organization of brain functions facilitated by complementary processes such as dual decision pathways, called System 1 and System 2 by Daniel Kahneman. System 1 resembles many of the abilities captured by deep learning, such as large language models (LLMs). System 2 is mostly associated with structural causal models (SCMs). We outline some important areas where more research is needed. This includes the interplay between multiple process pathways and how System 2 learns and evolves to represent the causal relations in our world during the lifetime of agents that are interacting with the world. We also argue about the danger of how systems beyond the capabilities of the human brain carry the risk of considerable harm to our society. Addressing them requires serious discussion.
Background: Mnemonic discrimination (MD) involves distinguishing new stimuli from memories of highly similar “lure” items or events, and is a putative indirect probe of dentate gyrus functioning. MD is impaired in the elderly and in individuals with hippocampal lesions, schizophrenia, major depressive disorder, and Alzheimer’s disease. The gold-standard MD test, called the mnemonic similarity task (MST), is rarely used in clinical research. We therefore aimed to validate a novel analysis method that extracts information about MD and recognition memory in widely clinically used recognition memory tests which do not have categorical distinctions between “lures” and “foils.”Methods: By fitting a logistic function to the relationship between stimulus interference and the probability of classifying a stimulus as novel, at the single participant level, we derived participant-level indices of MD (λ) and overall recognition memory performance (Δ). We applied the novel measures to MST data from two independent datasets (N=18; N=67). Using linear mixed-effects modelling, we sought to confirm that λ predicts the MST’s lure discrimination index (LDI), while Δ predicts the MST’s overall recognition memory index (REC). Results: Across both datasets, λ predicted LDI (β=0.76, 95% CI [0.62-0.91], p<0.001), but not REC (β=-0.06, 95% CI [-0.20-0.09], p=0.438), while Δ predicted REC (β=0.93, 95% CI [0.83-1.02], p<0.001), but not LDI (β=0.06, 95% CI [-0.03-0.15], p=0197). The λ and Δ indices were not correlated.Conclusion: Our novel measure accurately indexes MD, without correlating with overall recognition memory performance. Future studies should apply it to large clinical datasets with widely used recognition memory tests, such as the California Verbal Learning Test.
Rehabilitation robots and assistive devices that detect the motor intent of their users can provide more intuitive and effective control. Pupil dilation occurs when people perform motor activities, but its utility for detecting motor intent has not been explored previously. In this work, a human participant research study is conducted to determine if pupillometric data can be used to differentiate between a person's intent to pick up or observe an object. Thirty participants were recruited to perform 120 trials of picking up and observing objects while their pupil dilation was recorded by an eye tracking headset. Features were extracted from the time series data and used to train a neural network classifier. The classifier was tested using leave-one-out cross-validation. The classifier achieved an average accuracy of 59.4% and F1 score of 0.578 across the thirty test datasets. The performance varied significantly depending on the participant used for testing, suggesting that the pupillometric approach to intent detection may be better suited to some participants than others. Future work should determine whether intent detection can be improved with more advanced machine learning methods, such as convolutional neural networks (CNN), and whether intent detection can be performed in real time.
ABSTRACTIntroductionPatients with bipolar disorder (BD) demonstrate episodic memory deficits, which may be hippocampal‐dependent and may be attenuated in lithium responders. Induced pluripotent stem cell–derived CA3 pyramidal cell–like neurons show significant hyperexcitability in lithium‐responsive BD patients, while lithium nonresponders show marked variance in hyperexcitability. We hypothesize that this variable excitability will impair episodic memory recall, as assessed by cued retrieval (pattern completion) within a computational model of the hippocampal CA3.MethodsWe simulated pattern completion tasks using a computational model of the CA3 with different degrees of pyramidal cell excitability variance. Since pyramidal cell excitability variance naturally leads to a mix of hyperexcitability and hypoexcitability, we also examined what fraction (hyper‐ vs. hypoexcitable) was predominantly responsible for pattern completion errors in our model.ResultsPyramidal cell excitability variance impaired pattern completion (linear model β = −2.00, SE = 0.03, p < 0.001). The effect was invariant to all other parameter settings in the model. Excitability variance, specifically hyperexcitability, increased the number of spuriously active neurons, increasing false alarm rates and producing pattern completion deficits. Excessive inhibition also induces pattern completion deficits by limiting the number of correctly active neurons during pattern retrieval.ConclusionsExcitability variance in CA3 pyramidal cell–like neurons observed in lithium nonresponders may predict pattern completion deficits in these patients. These cognitive deficits may not be fully corrected by medications that minimize excitability. Future studies should test our predictions by examining behavioral correlates of pattern completion in lithium‐responsive and ‐nonresponsive BD patients.
Advances in underwater imaging enable collection of extensive seafloor image datasets necessary for monitoring important benthic ecosystems. The ability to collect seafloor imagery has outpaced our capacity to analyze it, hindering mobilization of this crucial environmental information. Machine learning approaches provide opportunities to increase the efficiency with which seafloor imagery is analyzed, yet large and consistent datasets to support development of such approaches are scarce. Here we present BenthicNet: a global compilation of seafloor imagery designed to support the training and evaluation of large-scale image recognition models. An initial set of over 11.4 million images was collected and curated to represent a diversity of seafloor environments using a representative subset of 1.3 million images. These are accompanied by 3.1 million annotations translated to the CATAMI scheme, which span 190,000 of the images. A large deep learning model was trained on this compilation and preliminary results suggest it has utility for automating large and small-scale image analysis tasks. The compilation and model are made openly available for reuse.
Evaluation of per-sample uncertainty quantification from neural networks is essential for decision-making involving high-risk applications. A common approach is to use the predictive distribution from Bayesian or approximation models and decompose the corresponding predictive uncertainty into epistemic (model-related) and aleatoric (data-related) components. However, additive decomposition has recently been questioned. In this work, we propose an intuitive framework for uncertainty estimation and decomposition based on the signal-to-noise ratio of class probability distributions across different model predictions. We introduce a variance-gated measure that scales predictions by a confidence factor derived from ensembles. We use this measure to discuss the existence of a collapse in the diversity of committee machines.
The hematology analytics used for detection and classification of small blood components is a significant challenge. In particular, when objects exists as small pixel-sized entities in a large context of similar objects. Deep learning approaches using supervised models with pre-trained weights (e.g., ImageNet), such as residual networks (ResNets) and vision transformers (ViTs) have demonstrated success for many applications. Unfortunately, when applied to images outside the domain of learned representations, these methods often result with less than acceptable performance. A strategy to overcome this can be achieved by using self-supervised models, where representations are learned from in-domain images and weights are then applied for downstream applications, such as semantic segmentation. Recently, masked autoencoders (MAEs) have proven to be effective in pre-training ViTs to obtain representations that captures global context information. By masking regions of an image and having the model learn to reconstruct both the masked and non-masked regions, the resulting weights can be used for various applications. However, if the sizes of the objects in images are less than the size of the mask (i.e., patch size), the global context information is lost, making it almost impossible to reconstruct the image. In this study, we investigated the effect of mask ratios and patch sizes for blood components using a "small-scale" MAE to obtain learned ViT encoder representations. We then applied the encoder weights to train a U-Net Transformer (UNETR) for semantic segmentation to obtain both local and global contextual information. Our experimental results demonstrates that both smaller mask ratios and patch sizes improve the reconstruction of images using a MAE. We also show the results of semantic segmentation with and without pre-trained weights, where smaller-sized blood components benefited with pre-training. Overall, our proposed method offers an efficient and effective strategy for the segmentation and classification of small objects.
Induced pluripotent stem cell (iPSC) derived hippocampal dentate granule cell-like neurons from individuals with bipolar disorder (BD) are hyperexcitable and more spontaneously active relative to healthy control (HC) neurons. Furthermore, these abnormalities are normalised after the application of lithium in neurons derived from clinical lithium responders (LR) only. How these abnormalities impact hippocampal microcircuit computation is not understood. We aimed to investigate the impacts of BD-associated abnormal granule cell (GC) activity on pattern separation (PS) using a computational model of the dentate gyrus. We used parameter optimization to fit the parameters of biophysically realistic granule cell (GC) models to electrophysiological data from iPSC GCs from patients with BD. These cellular models were incorporated into dentate gyrus networks to assess impacts on PS using an adapted spatiotemporal task. Relationships between BD, lithium and spontaneous activity were analysed using a linear mixed-effects model. Lithium and BD negatively impacted PS, consistent with clinical reports of cognitive slowing and memory impairment during lithium therapy. By normalising spontaneous activity levels, lithium improved PS performance in LRs only. Improvements in PS after lithium therapy in LRs may therefore be attributable to the normalisation of spontaneous activity levels, rather than reductions in GC intrinsic excitability as we hypothesised. Our results mirror previous research demonstrating that mnemonic discrimination improves after lithium therapy in lithium responders only, supporting a hypothesised link between behavioural mnemonic discrimination and dentate gyrus PS. Our work can be expanded to also consider the effects of lithium-induced neurogenesis on PS.
Pattern separation is a computational process by which dissimilar neural patterns are generated from similar input patterns. We present an information-geometric formulation of pattern separation, where a pattern separator is modeled as a family of statistical distributions on a manifold. Such a manifold maps an input (i.e., coordinates) to a probability distribution that generates firing patterns. Pattern separation occurs when small coordinate changes result in large distances between samples from the corresponding distributions. Under this formulation, we implement a two-neuron system whose probability law forms a three-dimensional manifold with mutually orthogonal coordinates representing the neurons’ marginal and correlational firing rates. We use this highly controlled system to examine the behavior of spike train similarity indices commonly used in pattern separation research. We find that all indices (except scaling factor) are sensitive to relative differences in marginal firing rates, but no index adequately captures differences in spike trains that result from altering the correlation in activity between the two neurons. That is, existing pattern separation metrics appear (A) sensitive to patterns that are encoded by different neurons but (B) insensitive to patterns that differ only in relative spike timing (e.g., synchrony between neurons in the ensemble).
In this work, we apply state-of-the-art selfsupervised learning techniques on a large dataset of seafloor imagery, BenthicNet, and study their performance for a complex hierarchical multi-label (HML) classification downstream task. In particular, we demonstrate the capacity to conduct HML training in scenarios where there exist multiple levels of missing annotation information, an important scenario for handling heterogeneous real-world data collected by multiple research groups with differing data collection protocols. We find that, when using smaller one-hot image label datasets typical of local or regional scale benthic science projects, models pre-trained with self-supervision on a larger collection of in-domain benthic data outperform models pre-trained on ImageNet. In the HML setting, we find the model can attain a deeper and more precise classification if it is pre-trained with self-supervision on indomain data. We hope this work can establish a benchmark for future models in the field of automated underwater image annotation tasks and can guide work in other domains with hierarchical annotations of mixed resolution.
Introduction: Mnemonic discrimination (MD), the ability to discriminate new stimuli from similar memories, putatively involves dentate gyrus pattern separation. Since lithium may normalize dentate gyrus functioning in lithium -responsive bipolar disorder (BD), we hypothesized that lithium treatment would be associated with better MD in lithium -responsive BD patients. Methods: BD patients (N = 69; NResponders = 16 [23 %]) performed the Continuous Visual Memory Test (CVMT), which requires discriminating between novel and previously seen images. Before testing, all patients had prophylactic lithium responsiveness assessed over >= 1 year of therapy (with the Alda Score), although only thirtyeight patients were actively prescribed lithium at time of testing (55 %; 12/16 responders, 26/53 nonresponders). We then used computational modelling to extract patient -specific MD indices. Linear models were used to test how (A) lithium treatment, (B) lithium responsiveness via the continuous Alda score, and (C) their interaction, affected MD. Results: Superior MD performance was associated with lithium treatment exclusively in lithium -responsive patients (Lithium x AldaScore beta = 0.257 [SE 0.078], p = 0.002). Consistent with prior literature, increased age was associated with worse MD (beta = -0.03 [SE 0.01], p = 0.005). Limitations: Secondary pilot analysis of retrospectively collected data in a cross-sectional design limits generalizability. Conclusion: Our study is the first to examine MD performance in BD. Lithium is associated with better MD performance only in lithium responders, potentially due to lithium's effects on dentate gyrus granule cell excitability. Our results may influence the development of behavioural probes for dentate gyrus neuronal hyperexcitability in BD.
High-resolution seabed sediment information is critical for a range of marine spatial planning applications in multi-use shelf environments. To establish this information for the Bay of Fundy, Canada, legacy seabed sediment measurements were obtained from regional data compilations, and eight parameters describing the grain size were modelled across the extent of the bay using high resolution acoustic seafloor mapping and oceanographic datasets. This was achieved using a purpose-made convolutional neural network configured for geospatial modelling of multivariate grain size parameters. Shared information between the response parameters enabled model training with partially complete observations from the varied legacy data sources, and an explicit multiscale model architecture ensured that environmental predictors were implemented at appropriate scales for modelling each parameter. This avoids typical exhaustive exploration and selection of scale-specific predictor sets that often precede model building. Compositional grain size parameters were additionally accommodated using appropriate output activation functions, providing an efficient alternative to compositional data transformation and imputation. Results agreed well with our current understanding of the surficial geology of the bay, and cross-validation was used to quantitatively evaluate map predictions. Of the eight predicted parameters, the mean grain size and mud (clay and silt) fractions were predicted with high accuracy (> 50% variance explained); the accuracy of grain size skewness was comparatively low (24% variance explained). Exploration of variable importance suggested that compiled acoustic backscatter was the most important environmental variable for predicting the grain size, but that geographic information describing the latitude and longitude within the bay was also highly useful. We hypothesize an interaction between these variables that enables location-specific prediction. Data layers of predicted grain size parameter values are made available for further sedimentological and ecological exploration, and for marine spatial planning activities within the bay.
Mnemonic discrimination (MD) may be dependent on oscillatory perforant path input frequencies to the hippocampus in a "U"-shaped fashion, where some studies show that slow and fast input frequencies support MD, while other studies show that intermediate frequencies disrupt MD. We hypothesize that pattern separation (PS) underlies frequency-dependent MD performance. We aim to study, in a computational model of the hippocampal dentate gyrus (DG), the network and cellular mechanisms governing this putative "U"-shaped PS relationship. We implemented a biophysical model of the DG that produces the hypothesized "U"-shaped input frequency-PS relationship, and its associated oscillatory electrophysiological signatures. We subsequently evaluated the network's PS ability using an adapted spatiotemporal task. We undertook systematic lesion studies to identify the network-level mechanisms driving the "U"-shaped input frequency-PS relationship. A minimal circuit of a single granule cell (GC) stimulated with oscillatory inputs was also used to study potential cellular-level mechanisms. Lesioning synapses onto GCs did not impact the "U"-shaped input frequency-PS relationship. Furthermore, GC inhibition limits PS performance for fast frequency inputs, while enhancing PS for slow frequency inputs. GC interspike interval was found to be input frequency dependent in a "U"-shaped fashion, paralleling frequency-dependent PS observed at the network level. Additionally, GCs showed an attenuated firing response for fast frequency inputs. We conclude that independent of network-level inhibition, GCs may intrinsically be capable of producing a "U"-shaped input frequency-PS relationship. GCs may preferentially decorrelate slow and fast inputs via spike timing reorganization and high frequency filtering.