Abstract Ecoacoustics mainly aims at monitoring soundscapes by means of non‐invasive protocols. Despite the widespread adoption of machine and deep learning techniques, existing ecoacoustic models predominantly rely on supervised learning and, consequently, face two primary limitations: (1) the necessity of annotated data; and (2) the restriction to fixed pre‐defined classes. In this work, we leverage recent advances in contrastive language‐audio pretraining (CLAP) and envision its first application to large‐scale soundscape analysis. As trained on extensive datasets of paired audio and global text descriptions using contrastive learning, CLAP allows computing similarity scores between audio and text prompts without the constraints of pre‐defined categorical lists. This flexibility enables a comprehensive investigation of various elements within the recordings from coarse‐grained (e.g. mammals, weather, humans, vehicles) to fine‐grained (e.g. dog, rain, speech, airplane) descriptions. Here, we first conducted a preliminary experiment on a calibration dataset, featuring audio events likely to occur in soundscapes, which is shared with the community and constituted from an online, free sound library. Then, we developed a methodology to define reproducible, bounded, independent and interpretable Contrastive Ecoacoustic Indices (CEI), which can characterize the prevalence of four primary sound categories in soundscapes—biophony, geophony, anthropophony and technophony. We finally computed these new CEI on 9‐month field recordings (189,137 1‐min excerpts) monitoring both tropical (Ecuador) and temperate (France) soundscapes, portraying an anthropic gradient from protected forests to urban city centres. This experiment reveals clear soundscape patterns associated with human population density, suggesting that the CEI could be used in other ecological contexts.
更多