Uploaded August 2025 | Updated September 2026, 58 minutes ago
Vanessa Süßle
Abstract Title: Automated Bioacoustic Monitoring on Unlabeled Data Using Artificial Intelligence
Co-Authors: Michael Doell; Dominik Kuehn; Matthew J. Burnett; Colleen T. Downs, Andreas Weinmann, Elke Hergenroether
Author affiliations: Department of Computer Science, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany and, P/Bag X01, Scottsville, Pietermaritzburg, 3209, South Africa;, Schoefferstrasse 3, Darmstadt, Germany; Department of Computer Science, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany; Centre for Functional Biodiversity, School of Life Sciences, University of KwaZulu-Natal, P/Bag X01, Scottsville, Pietermaritzburg, 3209, South Africa; Centre for Functional Biodiversity, School of Life Sciences, University of KwaZulu-Natal, P/Bag X01, Scottsville, Pietermaritzburg, 3209, South Africa; Algorithms for Computer Vision, Imaging and Data Analysis Group, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany; Department of Computer Science, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany;
Abstract: Analyses for biodiversity monitoring based on passive acoustic monitoring (PAM) recordings are time-consuming and challenged by the presence of background noise in recordings.
The analysis process can be supported by methods based on Artificial Intelligence (AI) to save time. By converting an audio file into a spectrogram creating a 2D-signal (image) that can be processed with computer vision architectures.In general, the training of AI models requires large amounts of annotated data and for example can only classify into the species it was trained on. Therefore, existing models work only on certain geo-specific species they were trained on and the development of models for new habitats required annotated data. We therefore sought to address this. We developed a data-processing-pipeline that automatically extracted labeled data from available platforms for the species of interest. The extracted annotated data were embedded into background recordings, including environmental sounds and noise, and were used to train and fine-tune Convolutional Recurrent Neural Network (CRNN) models, an architecture which considers temporal dependencies in the data. The pipeline was evaluated on unprocessed real world data recorded in urban KwaZulu-Natal habitats to classify six different pre-selected bird species of interest. The approach to automatically extract labeled data for chosen avian species enables an easy adaption of PAM to other species and habitats for future conservation projects. The prediction for new unlabelled data can be performed in a developed graphical user interface for an easy application.
Vanessa Süßle
Abstract Title: Automated Bioacoustic Monitoring on Unlabeled Data Using Artificial Intelligence
Co-Authors: Michael Doell; Dominik Kuehn; Matthew J. Burnett; Colleen T. Downs, Andreas Weinmann, Elke Hergenroether
Author affiliations: Department of Computer Science, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany and, P/Bag X01, Scottsville, Pietermaritzburg, 3209, South Africa;, Schoefferstrasse 3, Darmstadt, Germany; Department of Computer Science, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany; Centre for Functional Biodiversity, School of Life Sciences, University of KwaZulu-Natal, P/Bag X01, Scottsville, Pietermaritzburg, 3209, South Africa; Centre for Functional Biodiversity, School of Life Sciences, University of KwaZulu-Natal, P/Bag X01, Scottsville, Pietermaritzburg, 3209, South Africa; Algorithms for Computer Vision, Imaging and Data Analysis Group, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany; Department of Computer Science, University of Applied Sciences Darmstadt, Schoefferstrasse 3, Darmstadt, Germany;
Abstract: Analyses for biodiversity monitoring based on passive acoustic monitoring (PAM) recordings are time-consuming and challenged by the presence of background noise in recordings.
The analysis process can be supported by methods based on Artificial Intelligence (AI) to save time. By converting an audio file into a spectrogram creating a 2D-signal (image) that can be processed with computer vision architectures.In general, the training of AI models requires large amounts of annotated data and for example can only classify into the species it was trained on. Therefore, existing models work only on certain geo-specific species they were trained on and the development of models for new habitats required annotated data. We therefore sought to address this. We developed a data-processing-pipeline that automatically extracted labeled data from available platforms for the species of interest. The extracted annotated data were embedded into background recordings, including environmental sounds and noise, and were used to train and fine-tune Convolutional Recurrent Neural Network (CRNN) models, an architecture which considers temporal dependencies in the data. The pipeline was evaluated on unprocessed real world data recorded in urban KwaZulu-Natal habitats to classify six different pre-selected bird species of interest. The approach to automatically extract labeled data for chosen avian species enables an easy adaption of PAM to other species and habitats for future conservation projects. The prediction for new unlabelled data can be performed in a developed graphical user interface for an easy application.










