Passive acoustic monitoring generates datasets that increasingly rely on globallytrained birdsong models for species detection. Standard workflows apply a confidencethreshold to filter classifier outputs, but this practice has the potential to biasresults, and the position of the threshold is often largely arbitrary. Here we present ahybrid workflow that combines supervised classifier outputs with human-in-the-loopclustering of deep-learning embeddings to validate detections in a threshold-agnosticmanner. As clustering an entire acoustic dataset dataset is computationally expensive,the supervised model acts as a filter: detections are retained at low modelconfidence, audio snippets are generated for the full detection pool, and embeddingsof those snippets are then clustered, reviewed and labelled by an expert. To supportthis workflow we introduce pardalote, a graphical tool that allows users to clusterembeddings by similarity, audition representative audio and spectrograms for eachcluster, label and discard clusters that do not contain the target signal, and re-clusterthe remainder. This iterative clustering approach allows for flexible, expert-led validationusing Uniform Manifold Approximation and Projection (UMAP) and HierarchicalDensity-Based Spatial Clustering of Applications with Noise (HDBSCAN).We applied the workflow to 64,596 BirdNET detections of four bird species, obtainedfrom 60 acoustic monitoring points across 15 vineyard sites in southeast Queensland,Australia, using embeddings from both BirdNET and Perch v2. Clustering retained11,726 confirmed vocalisations at a precision of 0.927 and a recall of 0.942 relative tothe validated pool, and removed 51,229 of 52,154 false positives (98.2%). This processoutperformed the supervised models for all species at any threshold. Confirmedvocalisations occurred across all thresholds for every species, with false positives alsopersisting at high confidence for the worse-performing supervised models. Hybridpipelines that merge supervised classifiers with human-in-the-loop clustering offer ascalable framework for validating automated acoustic detections. This process allowsfor the examination of model performance across all thresholds, the rapid removalof false-positive clusters, and the collection of diverse training data. In practice,it decouples validation from the confidence threshold, and recovers detections thatotherwise would have been missed.
Monitoring threatened species is crucial for conservation, providing the necessary information on species distribution and abundance required to detect declines and evaluate the impact of conservation actions. Passive acoustic monitoring networks, such as the Australian Acoustic Observatory (A2O), have significant potential to aid the monitoring of vocal threatened species by collecting continuous data across many locations in a cost-effective manner. We analysed over 2 million hours of acoustic recordings collected as part of the A2O from 63 sites between 2019 and 2022 to determine the site presence of 74 threatened species. Using the existing deep-learning model BirdNET to detect species belonging to classes within the model, and embeddings search for species outside the model, we successfully detected 41 out of the 74 threatened target species at a minimum of one A2O location, with some broadly distributed threatened species detected at more than ten sites. Threatened species belonging to all three threatened categories (i.e., Vulnerable, Endangered, Critically Endangered), and all three broad taxonomic groupings (i.e., birds, frogs, mammals) were found within A2O recordings, and the majority of sites searched (83%) detected at least one threatened species. We provide information on the distribution of detections for all target species, an evaluation of classification performance for all species within the existing BirdNET model, as well as easy-to-use linear classifiers trained on BirdNET embeddings for all threatened species detected, which have improved performance over the standard BirdNET model. We conclude that acoustic monitoring networks such as the A2O, powered by deep-learning models, can serve as an important tool for monitoring threatened species.
Passive acoustic monitoring (PAM) has shown great promise in helping ecologists understand the health of animal populations and ecosystems. However, extracting insights from millions of hours of audio recordings requires the development of specialized recognizers. This is typically a challenging task, necessitating large amounts of training data and machine learning expertise. In this work, we introduce a general, scalable and data-efficient system for developing recognizers for novel bioacoustic problems in under an hour. Our system consists of several key components that tackle problems in previous bioacoustic workflows: 1) highly generalizable acoustic embeddings pre-trained for birdsong classification minimize data hunger; 2) indexed audio search allows the efficient creation of classifier training datasets, and 3) precomputation of embeddings enables an efficient active learning loop, improving classifier quality iteratively with minimal wait time. Ecologists employed our system in three novel case studies: analyzing coral reef health through unidentified sounds; identifying juvenile Hawaiian bird calls to quantify breeding success and improve endangered species monitoring; and Christmas Island bird occupancy modeling. We augment the case studies with simulated experiments which explore the range of design decisions in a structured way and help establish best practices. Altogether these experiments showcase our system's scalability, efficiency, and generalizability, enabling scientists to quickly address new bioacoustic challenges.
Passive acoustic monitoring and machine learning are increasingly being used to survey threatened species. When automated detection models are applied to large novel datasets, false-positive detections are likely even for high-performing models, and arbitrary thresholds may result in missed detections. Manual validation of outputs is time consuming, and additional fine-scale annotation of individual notes is impractical for large datasets and difficult to automate when using passive field recordings. This research presents an acoustic monitoring pipeline which employs a multi-stage hybrid approach: initial detection using a convolutional neural network classifier, followed by segmentation and iterative unsupervised clustering of extracted acoustic features using UMAP and HDBSCAN to remove label noise. We applied the pipeline to a large acoustic dataset comprised of 2764 h of environmental recordings and test the utility of the approach on territorial calls of Australia's largest owl: the threatened Powerful Owl (Ninox strenua). The pipeline reduced the large acoustic dataset into 10,116 annotations, of which 9399 (93 %) were correctly annotated individual notes of the target species. The clustering process also eliminated 88 % of false positive detections while retaining 95 % true positives (F1 = 0.94). The approach is highly scalable, can be applied to very large acoustic datasets, and can rapidly collect note-level annotations from noisy field recordings. The acoustic features derived from this methodology identified population differences in our test dataset and enable further exploration of song structure, geographic variation, and vocal individuality. The clustering process also facilitates a semi-supervised learning approach, allowing rapid selection of uncertain examples for model improvement. The pipeline helps to address two key challenges in bioacoustic monitoring: the need for manual validation of automated detections and the difficulty of obtaining accurate note-level annotations in noisy field recordings. Adaptation of these methods to other species and vocalisations may facilitate improved detection and investigation of vocal characteristics across different populations or regions.
Aim: The urgency for remote, reliable and scalable biodiversity monitoring amidst mounting human pressures on ecosystems has sparked worldwide interest in Passive Acoustic Monitoring (PAM), which can track life underwater and on land. However, we lack a unified methodology to report this sampling effort and a comprehensive overview of PAM coverage to gauge its potential as a global research and monitoring tool. To address this gap, we created the Worldwide Soundscapes project, a collaborative network and growing database comprising metadata from 416 datasets across all realms (terrestrial, marine, freshwater and subterranean). Location: Worldwide, 12,343 sites, all ecosystem types. Time Period: 1991 to present. Major Taxa Studied: All soniferous taxa. Methods: We synthesise sampling coverage across spatial, temporal and ecological scales using metadata describing sampling locations, deployment schedules, focal taxa and audio recording parameters. We explore global trends in biological, anthropogenic and geophysical sounds based on 168 selected recordings from 12 ecosystems across all realms. Results: Terrestrial sampling is spatially denser (46 sites per million square kilometre-Mkm(2)) than aquatic sampling (0.3 and 1.8 sites/Mkm(2) in oceans and fresh water) with only two subterranean datasets. Although diel and lunar cycles are well sampled across realms, only marine datasets (55%) comprehensively sample all seasons. Across the 12 ecosystems selected for exploring global acoustic trends, biological sounds showed contrasting diel patterns across ecosystems, declined with distance from the Equator, and were negatively correlated with anthropogenic sounds. Main Conclusions: PAM can inform macroecological studies as well as global conservation and phenology syntheses, but representation can be improved by expanding terrestrial taxonomic scope, sampling coverage in the high seas and subterranean ecosystems, and spatio-temporal replication in freshwater habitats. Overall, this worldwide PAM network holds promise to support cross-realm biodiversity research and monitoring efforts.
AbstractPassive acoustic recorders have emerged as powerful tools for ecological monitoring. However, effective monitoring is not simply an act of recording sounds. To have meaning for conservation and management, acoustic monitoring needs to be properly planned and analyzed to yield high quality information. Here, we provide a set of considerations for the design of an effective acoustic monitoring program. We argue that such a program, has the following attributes: (1) has established appropriate partnerships with landowners, Traditional Owners, researchers, or other relevant stakeholders, (2) is based on clear objectives and questions, (3) is explicit in its target sound signals, (4) has considered in‐field sensor placement for a range of factors, including experimental design, statistical power, background noise, and potential impacts on human privacy and animal disturbance, (5) has a justified recording schedule and periodicity, (6) has methods to process sound data in line with objectives, and (7) has protocols for permanent data storage and access. Acoustic monitoring is increasingly used in large‐scale programs and will be important in addressing global biodiversity targets and new biodiversity markets. It is critical that new monitoring programs are designed to effectively and efficiently capture data that address pertinent and emerging issues in conservation.
Climate change and biodiversity loss are significant global environmental issues. However, to understand their impacts we need to know how fauna respond to environmental and climatic variation over time. In this study, remote sensing techniques (satellite imagery and passive acoustic recorders) were used to investigate the variation in biophony over different timescales, ranging from one day to one year, in a sub-tropical woodland in eastern Australia. The prominent sources of biophony were birds at dawn and during the day, nocturnal insects at dusk and during the night, and diurnal birds and insects (mainly cicadas) over the summer period of December, January, and February. While different envi-ronmental factors were found to be key drivers of phenological response in different faunal groups, temperature, hu-midity and the interactions between temperature, humidity, moon illumination and vegetation greenness were most important factors overall. Using observed temperatures relative to the historical mean for each day of the year, we eval-uated the impact of higher-than-average temperatures on calling activity. We found that nocturnal insects call less fre-quently on days when the temperature was hotter than average in winter months (June, July, and August), and birds call less frequently in hot spring days (September, October, and November) meaning these groups can be susceptible to temperature increase as consequence, for example, of climate change. This study demonstrates how animal calling be-haviour is affected by different environmental variables over different temporal scales. This study also demonstrates the utility of remote sensing techniques for assessing the impacts of climate change on biodiversity. It is highly recom-mended that monitoring schemes and impact assessments account for phenological changes and environmental vari-ability, as these are complex and important processes shaping animal communities.
Effective monitoring tools are key for tracking biodiversity loss and informing management intervention strategies. Passive acoustic monitoring promises to provide a cheap and effective way to monitor biodiversity across large spatial and temporal scales, however, extracting useful information from long-duration audio recordings still proves challenging. Recently, a range of acoustic indices have been developed, which capture different aspects of the soundscape, and may provide a way to estimate traditional biodiversity measures. Here we investigated the relationship between 13 acoustic indices obtained from passive acoustic monitoring and biodiversity estimates of various vertebrate taxonomic groupings obtained from manual surveys at six sites spanning over 20 degrees of latitude along the Australian east coast. We found a number of individual acoustic indices that correlated well with species richness, Shannon’s diversity index, and total individual count estimates obtained from traditional survey methods. Correlations were typically greater for avian and total vertebrate biodiversity than for anuran and non-avian vertebrate biodiversity. Acoustic indices also correlated better with species richness and total individual count than with Shannon’s diversity index. Random forest models incorporating multiple acoustic indices provided more accurate predictions than single indices alone. Out of the acoustic indices tested, cluster count, mid-frequency cover and spectral density contributed the greatest predictive ability to models. Our results suggest that models incorporating multiple acoustic indices could be a useful tool for monitoring certain vertebrate groups. Further work is required to understand how site-specific variables can be incorporated into models to improve predictive capabilities and how to improve the monitoring of taxa besides avians, particularly anurans.
Observatories are designed to collect data for a range of uses. The Australian Acoustic Observatory (A2O) was established to collect environmental sound, including audible species calls, from 344 recorders at 86 sites around Australia. We examine the potential of the A2O to monitor near threatened, threatened, endangered and critically endangered species, based on their vocal behaviour, geographic distributions in relation to the sites of the A2O and on some knowledge of habitat use. Using IUCN and EPBC lists of threatened and endangered species, we extracted species that vocalized in the audible range, and using conservative estimates of their geographic ranges, determined whether there was a possibility of hearing them at these sites. We found that it may be possible to detect up to 171 threatened species at sites established for the A2O, and that individual sites have the potential to detect up to 40 threatened species. All 86 sites occurred in locations where threatened species could possibly be detected, and the list of detectable species included birds, amphibians, and mammals. We have incidentally detected one mammal and four bird species in the data during other work. Threatening processes to which potentially detectable species were exposed included all but two IUCN threat categories. We concluded that with applications of technology to search the audio data from the A2O, it could serve as an important tool for monitoring threatened species.
Traditional ecological survey methods are expensive and time-consuming as they typically involve a domain expert in the field. As a result of this, continuous audio recordings are playing an ever more critical role in conservation and biodiversity monitoring. However, listening to these recordings is often infeasible, as they can be thousands of hours long. The knowledge of domain experts can be efficiently leveraged using visualization. Traditionally, spectrograms are used to visualize audio data. However, there is a limit to the duration of audio that can fit on a computer screen as a spectrogram. Several techniques have been adapted to overcome this. Each of these techniques has its own set of advantages and disadvantages that make it ideal for some situations but unsuitable for others. In this paper, we propose a novel visualization based on embeddings produced by Frequency Preserving Autoencoders. We evaluate by visually comparing with traditional spectrogram and spectral index false-color spectrogram, using human-generated annotations as a baseline. We find that calls from some species are more consistently visible in our autoencoder-based visualization, while others are more clearly visible in the alternative visualizations. This novel visualization presents a new opportunity to assist ecologists in identifying bird calls in long-duration audio, if calls are difficult to identify using other methods.
Conservation and sustainable management efforts in tropical forests often lack reliable, effective, and easily-communicated ways to measure the biodiversity status of a protected or managed landscape. The sounds that many tropical species make can be recorded by pre-programmed devices and analysed to yield measures of biodiversity. Interpreting the resulting soundscapes has developed along two paths: analysing the whole soundscape using acoustic indices, used as a proxy of biodiversity, or focusing on individual species that can be either manually or automatically recognized from the soundscape. Here we develop an intermediate approach to divide the soundscape into frequency categories belonging to broad taxonomic groups of vocalizing animals. While the method was unable to distinguish between amphibian and mammal communities, it was successful in assigning parts of the soundscape as likely produced by birds and insects. Applying the approach in Borneo revealed that, with increasing land use intensity, i) the spectral saturation of the soundscape, a proxy of species richness, loses dawn and dusk peaks, ii) bird acoustic communities lose recurrent diurnal patterns, becoming less synchronized across sites, and that iii) insect Soundscape Saturation increases at night. If soundscapes are partitioned similarly in different regions, our method could be used to bridge soundscape-level and individual-species level analyses. Regaining dawn and dusk peaks, the synchrony of bird acoustic communities, and losing nocturnal dominance of insect could be used as a set of simple indicators of tropical forest retaining high levels of biodiversity.
Context Semi-arid landscapes are naturally heterogeneous with several factors influencing this variation. Fauna responses and adaptations vary in xeric environments, and the scale of observation is important. Biodiversity monitoring at several scales can be challenging, and acoustics are an alternative to this issue. Objectives We investigated how audible biodiversity is influenced by environmental factors (e.g.: vegetation metrics, climatic variables, etc.) across a fine spatial scale, aiming to provide a better understanding of the variation in audible species across recording locations placed close together. These results will improve the current knowledge on ecoacoustics as a tool for measuring ecological processes in this biome, and better inform conservation plans. Methods We collected data in the semi-arid region in Queensland, Australia placing 24 recorders 200 m apart for 48 h. We also sampled environmental attributes (e.g.: temperature and vegetation structure metrics) and used acoustic indices in a time-series algorithm to categorise sound into classes. Bird species and feeding guilds were also identified. Results We found significant differences between proximate sensors, demonstrating that soundscape differences occur across fine spatial scales. Birds and insects were the predominant biophonic sound observed and both groups were associated with shrub cover and subcanopy height. Environments with higher shrub and subcanopy cover had a higher percentage of all birds’ feeding guilds and insects. Sixty-three bird species were identified, including a threatened bird species in Queensland. Conclusion We show biodiversity is influenced by vegetation heterogeneity across fine spatial scales in semi-arid regions, identifying which attributes sustain higher levels of biodiversity activity. Our study reveals the practicality of acoustic surveys for this biodiversity monitoring by covering a large area in 48 h. However, we caution that scale is an important consideration when designing surveys.
1. Forests on private land have a wide range of uses that span activities such as recreation, primary production and nature conservation. Traditionally, it has been difficult for researchers to access private land to undertake systematic surveys. We used mini-acoustic sensors (Audiomoth) mailed via the postal service to overcome landholder concerns about researchers accessing private property, with a focus on properties used for private native forestry. 2. We surveyed koalas, an iconic threatened marsupial, in north-east New South Wales, Australia using passive acoustics, with repeat surveys over consecutive nights to account for imperfect detection in an occupancy modelling framework. 3. Over 3 years, we surveyed 128 sites and recorded 2,560 male bellows. Detection probability over seven nights was high (>0.79), but varied substantially between years, due to use of different sensors, housings and weather conditions. After accounting for detection probability, modelling revealed that koalas commonly occupied private native forests of the study region (probability of occupancy = 0.58 +/- 0.08). 4. Occupancy was modelled against several covariates and it varied with the landscape extent of sealed roads (.ve), NDVI (.ve) and a habitat suitability model (+ve, but minor). There was no support for occupancy in private forests to be related to a range of other factors including extent of surrounding cleared land, timber harvesting history, fire and other measured habitat features. 5. Synthesis and applications. We conclude that mini--acoustic recorders mailed to landholders were a highly effective method for assessing koala occupancy, after accounting for variable detection, and the approach could be deployed more widely for a range of species. Private native forests in partly cleared landscapes are commonly occupied by koalas, highlighting that this tenure is crucial for koala conservation and that practices seeking to balance conservation and production should be encouraged. In addition to sensitive habitat management in private forests, sealed road density is a major threat needing to be addressed.
Context It is notoriously difficult to estimate the size of animal populations, especially for cryptic or threatened species that occur in low numbers. Recent advances with acoustic sensors make the detection of animal populations cost effective when coupled with software that can recognise species-specific calls. Aims We assess the potential for acoustic sensors to estimate koala, Phascolarctos cinereus, density, when individuals are not identified, using spatial count models. Sites were selected where previous independent estimates of density were available. Methods We established acoustic arrays at each of five sites representing different environments and densities of koalas in New South Wales. To assess reliability, we compared male koala density estimates derived from spatial count modelling to independently derived estimates for each site. Key results A total 11 312 koala bellows were verified across our five arrays. Koalas were detected at most of our sample locations (96–100% of sensors; n = 130), compared with low detection rates from rapid scat searches at trees near each sensor (scats at <2% of trees searched, n = 889, except one site where scats were present at 69% of trees, n = 129). Independent estimates of koala density at our study areas varied from a minimum of 0.02 male koalas ha−1 to 0.32 ha−1. Acoustic arrays and the spatial count method yielded plausible estimates of male koala density, which, when converted to total koalas (assuming 1:1 sex ratio), were mostly equivalent to independent estimates previously derived for each site. The greatest discrepancy occurred where the acoustic estimate was larger (although within the bounds of uncertainty) than the independent mark–recapture estimate at a fragmented, high koala-density site. Conclusions Spatial count modelling of acoustic data from arrays provides plausible and reliable estimates of koala density and, importantly, associated measures of uncertainty as well as an ability to model spatial variations in density across an array. Caution is needed when applying models to higher-density populations where home ranges overlap extensively and calls are evenly spread across the array. Implications The results add to the opportunities of acoustic methods for wildlife, especially where monitoring of density requires cost-effective repeat surveys.
The compatibility of forestry and koala conservation is a controversial issue. We used a BACIPS design to assess change in koala density after selective harvesting with regulations to protect environmental values. We also assessed additional sites heavily harvested 5–10 years previously, now dominated by young regeneration. We used replicate arrays of acoustic sensors and spatial count modelling of male bellowing to estimate male koala density over 3600 ha. Paired sites in nearby National Parks served as controls. Naïve occupancy was close to 100% before and after harvesting, indicating koalas were widespread across all arrays. Average density was higher than expected for forests in NSW, varying between arrays from 0.03–0.08 males ha −1 . There was no significant effect of selective harvesting on density and little change evident between years. Density 5–10 years after previous heavy harvesting was equivalent to controls, with one harvested array supporting the second highest density in the study. Within arrays, density was similar between areas mapped as selectively harvested or excluded from harvest. Density was also high in young regeneration 5–10 years after heavy harvesting. We conclude that native forestry regulations provided sufficient habitat for koalas to maintain their density, both immediately after selective harvesting and 5–10 years after heavy harvesting.
Continuous recording of environmental sounds could allow long-term monitoring of vocal wildlife, and scaling of ecological studies to large temporal and spatial scales. However, such opportunities are currently limited by constraints in the analysis of large acoustic data sets. Computational methods and automation of call detection require specialist expertise and are time consuming to develop, therefore most biological researchers continue to use manual listening and inspection of spectrograms to analyze their sound recordings. False-color spectrograms were recently developed as a tool to allow visualization of long-duration sound recordings, intending to aid ecologists in navigating their audio data and detecting species of interest. This paper explores the efficacy of using this visualization method to identify multiple frog species in a large set of continuous sound recordings and gather data on the chorusing activity of the frog community. We found that, after a phase of training of the observer, frog choruses could be visually identified to species with high accuracy. We present a method to analyze such data, including a simple R routine to interactively select short segments on the false-color spectrogram for rapid manual checking of visually identified sounds. We propose these methods could fruitfully be applied to large acoustic data sets to analyze calling patterns in other chorusing species.
Many organizations are attempting to scale ecoacoustic monitoring for conservation but are hampered at the stages of data management and analysis. We reviewed current ecoacoustic hardware, software, and standards, and conducted workshops with 23 participants across 10 organizations in Australia to learn about their current practices, and to identify key trends and challenges in their use of ecoacoustics data. We found no existing metadata schemas that contain enough ecoacoustics terms for current practice, and no standard approaches to annotation. There was a strong need for free acoustics data storage, discoverable learning resources, and interoperability with other ecological modeling tools. In parallel, there were tensions regarding intellectual property management, and siloed approaches to studying species within organizations across different regions and between organizations doing similar work. This research contributes directly to the development of an open ecoacoustics platform to enable the sharing of data, analyses, and tools for environmental conservation.
Automatically detecting the calls of species of interest in audio recordings is a common but often challenging exercise in ecoacoustics. This challenge is increasingly being tackled with deep neural networks that generally require a rich set of training data. Often, the available training data might not be from the same geographical region as the study area and so may contain important differences. This mismatch in training and deployment datasets can impact the accuracy at deployment, mainly due to confusing sounds absent from the training data generating false positives, as well as some variation in call types. We have developed a multiclass convolutional neural network classifier for seven target bird species to track presence absence of these species over time in cotton growing regions. We started with no training data from cotton regions but we did have an unbalanced library of calls from other locations. Due to the relative scarcity of calls in recordings from cotton regions, manually scanning and labeling the recordings was prohibitively time consuming. In this paper we describe our process of overcoming this data mismatch to develop a recognizer that performs well on the cotton recordings for most classes. The recognizer was trained on recordings from outside the cotton regions and then applied to unlabeled cotton recordings. Based on the resulting outputs a verification set was chosen to be manually tagged and incorporated in the training set. By iterating this process, we were gradually able to build the training set of cotton audio examples. Through this process, we were able to increase the average class F1 score (the harmonic mean of precision and recall) of the recognizer on target recordings from 0.45 in the first iteration to 0.74.
One quarter of all terrestrial native bird species have become extinct since human arrival in New Zealand, leading to a pervasive silence in many natural environments due to the decrease in native bird song. Passive acoustic techniques are a potential tool for environmental monitoring, especially for testing whether the control of mammals can reverse the 'silent forest' effect. Here we compare soundscapes from two nearby sites within the Waitakere Ranges Regional Park, New Zealand, that have contrasting predator control levels: one with high-level pest mammal control, and the other with low-level pest control. Measurements of twelve acoustic indices extracted from two seasons of passive acoustic recordings are split into 20 acoustic regions to identify which regions best discriminate between the two management regimes. We define the acoustic regions as units of analysis bounded by a specific time period and frequency range chosen to capture the main groups of biologically relevant acoustic events within a soundscape. Analysis of variance and pairwise comparisons indicated the acoustic region bounded from 9 pm to 11:59 pm and a range of 0.988-3.609 kHz in autumn presented the greatest differences between sites. The sounds responsible for these acoustic differences were generated by invasive mammals in the site with no pest control. Results also supports spring season as the most important for bird monitoring in New Zealand. Acoustic indices analysis did not detect a reversal of the "silence forest" effect in the site with high-level predator control.
Wildlife calls are the best witnesses to the health of ecosystems, if only we know how to listen to them. Efforts to understand and inform restoration of healthy ecosystems with environmental audio recordings languish from insufficient tools to learn and identify sounds in recordings. To address this problem, we designed and playtested the Bristle Whistle Challenge prototype with ten players. We explored how to design delightful interactions with audio for gaining awareness of nature sounds and supporting wildlife conservation through citizen science. We found that rather than presenting audio alone, it was necessary to connect sounds to other senses and experiences in creative ways to impart meaning and enhance engagement. We offer recommendations to design creative and contextual interactions with media to build awareness of nature's wonders. We call for greater efforts in interaction design to engage people with nature, which is the key to turning around our environmental crisis.