Insects are of crucial importance for terrestrial ecosystems, but many populations decline rapidly. Conventional collecting methods are usually time-consuming, resulting in a low temporal, spatial and taxonomic resolution of data. Automated camera light traps (CLTs) allow non-lethal monitoring of species-rich moths (Lepidoptera) and other nocturnal insects, but so far little is known about their performance compared to conventional collecting methods. By observing the behaviour of moths in previous field work, we hypothesised that CLTs perform well in moth groups in which species tend to sit down quietly after approaching the lamp (such as Geometridae) but worse in moth groups in which species are persistently active (such as Sphingidae). We tested the performance of two CLTs, equipped with Sony alpha 7II (24 megapixel sensor) cameras that resulted in images with approx. Four hundred twenty dpi resolution. The study was carried out in a forested area near Bielefeld in NW Germany for 196 nights in a row from March to October 2023, and photos were taken every 2 min during the night. All macromoths recognisable in the photographs were identified and counted individually. We directly compared the data from the CLTs with moth samples obtained from conventional funnel light traps (FLTs) during 12 nights which were spread across the flight season. The resulting images from the CLTs allowed reliable species identification of all observed macromoths with only a few exceptions due to technical problems. In direct comparison during 12 nights, CLTs recorded 39 species exclusively, FLTs recorded 48 species exclusively, and 53 species were recorded by both methods equally. During the whole sampling period of 196 nights in a row, a total of 225 moth species were recorded by CLTs. We found six indicator species for CLTs (all Geometridae) and one species for FLT, a hawkmoth species. Families differed in the length to which they remained on the screen of the CLTs. Our study was the first to systematically compare the methods and it shows that CLTs perform overall very well. Results from CLTs differ to a certain extent from conventional trapping methods because they seem to perform worse in groups with highly active species and perform better in calmer groups like geometrid moths. CLTs are promising devices for insect monitoring since they deliver data with high resolution in time, space and taxonomy. The use of artificial intelligence (AI) for the analysis of images is intended as the next logic step
Monitoring insect populations has become an urgent priority given the ongoing biodiversity crisis, and the resulting changes to ecosystem services. However, the available data on insect population trends are severely limited in terms of taxonomic, spatial, and temporal resolution, due to the time-consuming nature of insect collection and identification. The goal of LEPMON (LEPidoptera MONitoring) is to develop a powerful, stable and scalable automated nocturnal insect recording system for long-term monitoring and answering ecological questions. The project runs from December 2024 to November 2027. We provide an overview of the entire project, summarizing the original research proposal and current developments as of August 2026. LEPMON uses time-lapse digital photography of nocturnal insects attracted to a white screen using UV light (LepiLED) with a high resolution of 16 px/mm. Artificial intelligence (AI) is applied for automated processing of the collected images. The recording system comprises two different models of Automated Recorders for Nocturnal Insects (ARNIs). The ARNI-Pro is the high-end model with the highest image quality and durable components, designed for professional users. ARNI-CS is the more affordable and portable alternative with slightly reduced image quality, built largely using 3D-printed components. As of August 2026, 67 ARNI-Pros and 34 ARNI-CS's have been installed in the field. The ARNI-Pro models were set up (1) along eight urbanization gradients to test the system’s ability to detect community changes, and (2) in a variety of natural habitats across Germany to capture as many species as possible, ranging from raised bogs in the north to alpine habitats in the south. We also determine technical limits at extreme locations such as forest canopies and tropical environments and shortly assess the project’s risks and exploitation perspectives. The ARNI-CS models further support the recording of the community composition of nocturnal insects across various habitats in a citizen science context. The images and data generated by the ARNIs are uploaded to a scalable data management platform (LAUP = LEPMON Annotation and Upload Portal). LAUP processes the images and uses AI to enable large-scale object detection and species identification. As accurate AI models require extensive species-labelled training data, large numbers of manually identified images are needed. To obtain these identifications, we involve both taxonomic experts and citizen scientists. We aim to build international collaborations and share knowledge between countries, extending moth monitoring beyond Germany to strengthen LEPMON as a long-term biodiversity monitoring network.
We present LEPY, a free and openly available Python-based pipeline for the automated extraction and analysis of morphological and colour traits, from mounted specimens of Lepidoptera (butterflies and moths). The pipeline uses an automatically detected scale bar for accurate morphological measurements, together with image segmentation that separates the specimen from the background, with users able to pre-select from a set of segmentation models. We designed LEPY to be user-friendly and reproducible, ensuring efficient and consistent analysis of large image datasets. The pipeline also supports the integration of ultraviolet (UV) photographs for improved colour analysis, an innovative feature rarely available in existing trait-analysis tools.LEPY computes morphological traits such as body length, forewing length, and specimen area. It also extracts colour traits including hue, saturation, intensity from the red, green, and blue (RGB) channels, as well as brightness, contrast, chromaticity, and luminance from both RBG and UV channels. The pipeline uses the data to calculate colour diversity with the Shannon index, exports results in a structured, machine-readable format, and it also generates visual summaries of each image pair.We tested LEPY on different moth groups spanning a wide range of body sizes and colouration patterns. As an ecological case study, we applied the pipeline to complete datasets of Sphingidae and Saturniidae collected along an elevational gradient in the Peruvian Andes. The resulting trait data revealed taxon-dependent morphological and colour responses to elevation, thereby demonstrating LEPY's utility for analysing large-scale trait datasets.LEPY provides a robust and fully automated approach for the analysis of morphological and colour traits in Lepidoptera, supporting ecological and evolutionary research. Its scalability and ability to generate standardised, high-resolution trait datasets make it a valuable tool for biodiversity monitoring, macroecological research, and the development of global trait databases.
Vision-Language Models (VLMs) learn joint representations by mapping images and text into a shared latent space. However, recent research highlights that deterministic embeddings from standard VLMs often struggle to capture the uncertainties arising from the ambiguities in visual and textual descriptions and the multiple possible correspondences between images and texts. Existing approaches tackle this by learning probabilistic embeddings during VLM training, which demands large datasets and does not leverage the powerful representations already learned by large-scale VLMs like CLIP. In this paper, we propose GroVE, a post-hoc approach to obtaining probabilistic embeddings from frozen VLMs. GroVE builds on Gaussian Process Latent Variable Model (GPLVM) to learn a shared low-dimensional latent space where image and text inputs are mapped to a unified representation, optimized through single-modal embedding reconstruction and cross-modal alignment objectives. Once trained, the Gaussian Process model generates uncertainty-aware probabilistic embeddings. Evaluation shows that GroVE achieves state-of-the-art uncertainty calibration across multiple downstream tasks, including cross-modal retrieval, visual question answering, and active learning.
Fine-grained visual classification (FGVC) requires distinguishing between visually similar categories through subtle, localized features - a task that remains challenging due to high intra-class variability and limited inter-class differences. Existing part-based methods often rely on complex localization networks that learn mappings from pixel to sample space, requiring a deep understanding of image content while limiting feature utility for downstream tasks. In addition, sampled points frequently suffer from high spatial redundancy, making it difficult to quantify the optimal number of required parts. Inspired by human saccadic vision, we propose a two-stage process that first extracts peripheral features (coarse view) and generates a sample map, from which fixation patches are sampled and encoded in parallel using a weight-shared encoder. We employ contextualized selective attention to weigh the impact of each fixation patch before fusing peripheral and focus representations. To prevent spatial collapse - a common issue in part-based methods - we utilize non-maximum suppression during fixation sampling to eliminate redundancy. Comprehensive evaluation on standard FGVC benchmarks (CUB-200-2011, NABirds, Food-101 and Stanford-Dogs) and challenging insect datasets (EU-Moths, Ecuador-Moths and AMI-Moths) demonstrates that our method achieves comparable performance to state-of-the-art approaches while consistently outperforming our baseline encoder.
Deep clustering has proven successful in analyzing complex, high-dimensional real-world data. Typically, features are extracted from a deep neural network and then clustered. However, training the network to extract features that can be clustered efficiently in a semantically meaningful way is particularly challenging when data is sparse. In this paper, we present a semi-supervised method to fine-tune a deep learning network using Model-Agnostic Meta-Learning, commonly employed in Few-Shot Learning. We apply episodic training with a novel multivariate scatter loss, designed to enhance inter-class feature separation while minimizing intra-class variance, thereby improving overall clustering performance. Our approach works with state-of-the-art deep learning models, spanning convolutional neural networks and vision transformers, as well as different clustering algorithms like K-means and Spectral clustering. The effectiveness of our method is tested on several commonly used Few-Shot Learning datasets, where episodic fine-tuning with our multivariate scatter loss and a ConvNeXt backbone outperforms other models, achieving adjusted rand index scores of 89.7% on the EU moths dataset and 86.9% on the Caltech birds dataset, respectively. Hence, our proposed method can be applied across various practical domains, such as clustering images of animal species in biology.
Plant community data, like the species composition of the community and the phenology of the occurring species, are paramount for environmental research. Such data can be used to detect species responses to environmental changes, but the collection is very laborious, slow, and prone to human error. These detriments can be counteracted with automatic camera systems in combination with machine learning approaches that are able to extract the vegetation data from collected images in a consistent and fast manner. We introduce PlantCAPNet, an application to automate the analysis of herbaceous plant communities from images by extracting plant cover and phenology, addressing the tedious and biased nature of manual field collection. The system has an easy-to-use web interface with a single image prediction tool, a batch prediction function for image series, and a training interface for users to build novel models. We offer PlantCAPNet with two operational modes: a ’cover-trained’ mode for predicting cover and phenology using user-provided labeled data, and a ’zero-shot’ mode capable of predicting cover using only web-sourced data, thus lowering the barrier for entry. Our evaluations show that PlantCAPNet performs comparably or better than independent human experts in estimating plant cover. The zero-shot method reflects the reference estimates with a correlation of 0.625, and the cover-trained method with one of 0.790 compared to a correlation of 0.620 from independent experts. Moreover, we show that our system performs reliably for dataset with few species, and the cover prediction is also reliable for the most abundant species in datasets with many species, while the phenology prediction is dependent on the amount of training data. In total, our system offers higher consistency than human experts, and enables the extraction of high-temporal-resolution ecological data, facilitating novel environmental research.
Deep neural networks are prone to shortcut bias, where models rely on features that are statistically associated with the target label but lack causal relevance, leading to poor generalization under distribution shifts. To address this, debiasing methods aim to improve robustness by reducing reliance on these spurious features. Unfortunately, existing approaches typically assume unbiased test distributions, an idealized scenario that rarely holds in practice. As a result, they often underperform on the original biased distribution when compared with standard empirical risk minimization (ERM) models. We propose a novel Adaptive Model SELection approach for expanding post hoc debiasing called AMSEL, which maintains strong performance across test distributions with varying strength of spurious correlation. Using the fixed feature extractor of the biased model, AMSEL trains a family of lightweight classifier heads on simulated distributions ranging from the original biased data to a fully balanced version. At test time, it estimates the degree of spurious correlation in the test data and selects the most suitable classifier. We validate AMSEL on CelebA and ChestX-ray14, demonstrating that it matches the performance of debiased models under unbiased conditions while preserving the accuracy of the original biased model when spurious correlations are prevalent. AMSEL thus offers an adaptive solution to mitigate the impact of spurious correlations when their strength is either unknown or varies across application environments. Code and models are publicly available at https://github.com/debiasing/AMSEL.
The InsectAI COST action will support insect monitoring and conservation at the national and continental scale in order to understand and counteract widespread insect declines. The Action will bring together a critical mass of researchers and stakeholders in image-based insect AI technologies to direct and drive the research agenda, build research capacity across Europe and support innovation and application.There is mounting evidence that populations of insects around the world are in sharp decline. Understanding trends in species and their drivers is key to knowing the size of the challenge, its causes and how to address it. To identify solutions that lead to sustainable biodiversity alongside economic prosperity, insect monitoring should be efficient and provide standardised and frequently updated status indicators to guide conservation actions.The EU Biodiversity Strategy 2030 identifies the critical challenge of delivering standardised information about the state of nature and image-based insect AI can contribute to this. Specifically, the EU Nature Restoration Law will likely set binding targets for the high resolution data that cameras can provide. Thus, outputs of the Action will contribute directly to EU policies implementation, where biodiversity monitoring is considered a key component.The InsectAI COST Action will organise workshops, conferences, short-term scientific missions, hackathons, design-sprints and much more, across four Working Groups. These groups will address how image-based insect AI technologies can best address Societal Needs, support innovation in Image Collection hardware, create standardised approaches for Image Processing and develop novel Data Analysis and Integration methods for turning data into actionable insights.
Global change has a detrimental impact on the environment and changes biodiversity patterns, which can be observed, among others, via analyzing changes in the composition of plant communities. Typically, vegetation relevées are done manually, which is time-consuming, laborious, and subjective. Applying an automatic system for such an analysis that can also identify co-occurring species would be beneficial as it is fast, effortless to use, and consistent. Here, we introduce such a system based on Convolutional Neural Networks for automatically predicting the species-wise plant cover. The system is trained on freely available image data of herbaceous plant species from web sources and plant cover estimates done by experts. With a novel extension of our original approach, the system can even be applied directly to vegetation images without requiring such cover estimates. Our extended approach, not utilizing dedicated training data, performs similarly to humans concerning the relative species abundances in the vegetation relevées. When trained on dedicated training annotations, it reflects the original estimates more closely than (independent) human experts, who manually analyzed the same sites. Our method is, with little adaptation, usable in novel domains and could be used to analyze plant community dynamics and responses of different plant species to environmental changes.
Machine learning has achieved considerable success in data-intensive applications, yet encounters challenges when confronted with small datasets. Recently, few-shot learning (FSL) has emerged as a promising solution to address this limitation. By leveraging prior knowledge, FSL exhibits the ability to swiftly generalize to new tasks, even when presented with only a handful of samples in an accompanied support set. This paper extends the scope of few-shot learning by incorporating novelty detection for samples of categories not present in the support set of FSL. This extension holds substantial promise for real-life applications where the availability of samples for each class is either sparse or absent. Our approach involves adapting existing FSL methods with a cosine similarity function, complemented by the learning of a probabilistic threshold to distinguish between known and outlier classes. During episodic training with domain generalization, we introduce a scatter loss function designed to disentangle the distribution of similarities between known and outlier classes, thereby enhancing the separation of novel and known classes. The efficacy of the proposed method is evaluated on commonly used FSL datasets and the EU Moths dataset characterized by few samples. Our experimental results showcase accuracy, ranging from 95.4
It is common for domain experts like physicians in medical studies to examine features for their reliability with respect to a specific domain task. When introducing machine learning, a common expectation is that machine learning models use the same features as these human experts to solve a task, but that is not always the case. Moreover, datasets often contain features that are known from domain knowledge to generalize badly to the real world, referred to as biases. Current debiasing methods only remove such influences. To additionally integrate the domain knowledge about well-established features into the training of a model, their relevance should be increased. We present a method that allows the manipulation of the relevance of features by actively steering the model's feature selection during the training process. That is, it allows both the discouragement of biases and encouragement of well-established features to incorporate domain knowledge about the feature reliability. We model our objectives for actively steering the feature selection process as a constrained optimization problem, which we implement via a loss regularization that is based on batch-wise feature attributions. We evaluate our approach on a novel synthetic regression dataset and a dataset from the computer vision domain. We observe that it successfully steers the features a model selects during the training process. This is a strong indicator that our method can be used to integrate domain knowledge about well-established features into a model.
There are still open questions about how the learned representations of deep models change during the training process. Understanding this process could aid in validating the training. Towards this goal, previous works analyze the training in the mutual information plane. We use a different approach and base our analysis on a method built on Reichenbach’s common cause principle. Using this method, we test whether the model utilizes information contained in human-defined features. Given such a set of features, we investigate how the relative feature usage changes throughout the training process. We analyze multiple networks training on different tasks, including melanoma classification as a real-world application. We find that over the training, models concentrate on features containing information relevant to the task. This concentration is a form of representation compression. Crucially, we also find that the selected features can differ between training from-scratch and finetuning a pre-trained network.
Automatic camera-assisted monitoring of insects for abundance estimations is crucial to understand and counteract ongoing insect decline. In this paper, we present two datasets of nocturnal insects, especially moths as a subset of Lepidoptera, photographed in Central Europe. One of the datasets, the EU-Moths dataset, was captured manually by citizen scientists and contains species annotations for 200 different species and bounding box annotations for those. We used this dataset to develop and evaluate a two-stage pipeline for insect detection and moth species classification in previous work. We further introduce a prototype for an automated visual monitoring system. This prototype produced the second dataset consisting of more than 27,000 images captured on 95 nights. For evaluation and bootstrapping purposes, we annotated a subset of the images with bounding boxes enframing nocturnal insects. Finally, we present first detection and classification baselines for these datasets and encourage other scientists to use this publicly available data.
Biodiversity monitoring is crucial for tracking and counteracting adverse trends in population fluctuations. However, automatic recognition systems are rarely applied so far, and experts evaluate the generated data masses manually. Especially the support of deep learning methods for visual monitoring is not yet established in biodiversity research, compared to other areas like advertising or entertainment. In this paper, we present a deep learning pipeline for analyzing images captured by a moth scanner, an automated visual monitoring system of moth species developed within the AMMOD project. We first localize individuals with a moth detector and afterward determine the species of detected insects with a classifier. Our detector achieves up to 99.01% mean average precision and our classifier distinguishes 200 moth species with an accuracy of 93.13% on image cutouts depicting single insects. Combining both in our pipeline improves the accuracy for species identification in images of the moth scanner from 79.62% to 88.05%.
Part-based approaches for fine-grained recognition do not show the expected performance gain over global methods, although explicitly focusing on small details that are relevant for distinguishing highly similar classes. We assume that part-based methods suffer from a missing representation of local features, which is invariant to the order of parts and can handle a varying number of visible parts appropriately. The order of parts is artificial and often only given by ground-truth annotations, whereas viewpoint variations and occlusions result in not observable parts. Therefore, we propose integrating a Fisher vector encoding of part features into convolutional neural networks. The parameters for this encoding are estimated by an online EM algorithm jointly with those of the neural network and are more precise than the estimates of previous works. Our approach improves state-of-the-art accuracies for three bird species classification datasets.
Rapid changes of the biosphere observed in recent years are caused by both small and large scale drivers, like shifts in temperature, transformations in land-use, or changes in the energy budget of systems. While the latter processes are easily quantifiable, documentation of the loss of biodiversity and community structure is more difficult. Changes in organismal abundance and diversity are barely documented. Censuses of species are usually fragmentary and inferred by often spatially, temporally and ecologically unsatisfactory simple species lists for individual study sites. Thus, detrimental global processes and their drivers often remain unrevealed. A major impediment to monitoring species diversity is the lack of human taxonomic expertise that is implicitly required for large-scale and fine-grained assessments. Another is the large amount of personnel and associated costs needed to cover large scales, or the inaccessibility of remote but nonetheless affected areas. To overcome these limitations we propose a network of Automated Multisensor stations for Monitoring of species Diversity (AMMODs) to pave the way for a new generation of biodiversity assessment centers. This network combines cutting-edge technologies with biodiversity informatics and expert systems that conserve expert knowledge. Each AMMOD station combines autonomous samplers for insects, pollen and spores, audio recorders for vocalizing animals, sensors for volatile organic compounds emitted by plants (pVOCs) and camera traps for mammals and small invertebrates. AMMODs are largely self-containing and have the ability to pre-process data (e.g. for noise filtering) prior to transmission to receiver stations for storage, integration and analyses. Installation on sites that are difficult to access require a sophisticated and challenging system design with optimum balance between power requirements, bandwidth for data transmission, required service, and operation under all environmental conditions for years. An important prerequisite for automated species identification are databases of DNA barcodes, animal sounds, for pVOCs, and images used as training data for automated species identification. AMMOD stations thus become a key component to advance the field of biodiversity monitoring for research and policy by delivering biodiversity data at an unprecedented spatial and temporal resolution. (C) 2022 Published by Elsevier GmbH on behalf of Gesellschaft fur Okologie.
Bias in classifiers is a severe issue of modern deep learning methods, especially for their application in safetyand security-critical areas. Often, the bias of a classifier is a direct consequence of a bias in the training dataset, frequently caused by the co-occurrence of relevant features and irrelevant ones. To mitigate this issue, we require learning algorithms that prevent the propagation of bias from the dataset into the classifier. We present a novel adversarial debiasing method, which addresses a feature that is spuriously connected to the labels of training images but statistically independent of the labels for test images. Thus, the automatic identification of relevant features during training is perturbed by irrelevant features. This is the case in a wide range of bias-related problems for many computer vision tasks, such as automatic skin cancer detection or driver assistance. We argue by a mathematical proof that our approach is superior to existing techniques for the abovementioned bias. Our experiments show that our approach performs better than state-of-the-art techniques on a well-known benchmark dataset with real-world images of cats and dogs.
Weakly supervised object localization (WSOL) enables the detection and segmentation of objects in applications where localization annotations are hard or too expensive to obtain. Nowadays, most relevant WSOL approaches are based on class activation mapping (CAM), where a classification network utilizing global average pooling is trained for object classification. The classification layer that follows the pooling layer is then repurposed to generate segmentations using the unpooled features. The resulting localizations are usually imprecise and primarily focused around the most discriminative areas of the object, making a correct indication of the object location difficult. We argue that this problem is inherent in training with global average pooling due to its averaging operation. Therefore, we investigate two alternative pooling strategies: global max pooling and global log-sum-exp pooling. Furthermore, to increase the crispness and resolution of localization maps, we also investigate the application of Feature Pyramid Networks, which are commonplace in object detection. We confirm the usefulness of both alternative pooling methods a, well as the Feature Pyramid Network on the CUB-200-2011 and OpenImages datasets.
Animal re-identification based on image data, either recorded manually by photographers or automatically with camera traps, is an important task for ecological studies about biodiversity and conservation that can be highly automatized with algorithms from computer vision and machine learning. However, fixed identification models only trained with standard datasets before their application will quickly reach their limits, especially for long-term monitoring with changing environmental conditions, varying visual appearances of individuals over time that differ a lot from those in the training data, and new occurring individuals that have not been observed before. Hence, we believe that active learning with human-in-the-loop and continuous lifelong learning is important to tackle these challenges and to obtain high-performance recognition systems when dealing with huge amounts of additional data that become available during the application. Our general approach with image features from deep neural networks and decoupled decision models can be applied to many different mammalian species and is perfectly suited for continuous improvements of the recognition systems via lifelong learning. In our identification experiments, we consider four different taxa, namely two elephant species: African forest elephants and Asian elephants, as well as two species of great apes: gorillas and chimpanzees. Going beyond classical re-identification, our decoupled approach can also be used for predicting attributes of individuals such as gender or age using classification or regression methods. Although applicable for small datasets of individuals as well, we argue that even better recognition performance will be achieved by improving decision models gradually via lifelong learning to exploit huge datasets and continuous recordings from long-term applications. We highlight that algorithms for deploying lifelong learning in real observational studies exist and are ready for use. Hence, lifelong learning might become a valuable concept that supports practitioners when analyzing large-scale image data during long-term monitoring of mammals.