生物多样性监测中的主动学习:从标签效率到可靠的生态推断
摘要
本文综合了针对生物多样性监测的主动学习研究,关注标签效率,并强调了需要支持验证和可靠生态推断的方法。
arXiv:2609.27409v1 Announce Type: new
Abstract: Limited expert annotation capacity is a pervasive constraint in biodiversity monitoring. Passive acoustic recorders and camera traps generate data faster than experts can analyse them. Machine learning (ML) models can process these data at scale, but their reliability depends on the quality, quantity, and coverage of labelled samples, so expert time remains a constraint. Active learning (AL) eases this bottleneck by selecting, under a fixed annotation budget, the samples expected to improve a model most, and published evidence shows it can reduce the labels needed to reach a target performance. Monitoring programmes, however, face a broader question: how should a limited expert budget be divided so that model training, validation, and the ecological estimates built on model outputs all remain reliable? Because AL selects samples non-randomly, its labels are unsuitable for validation, calibration, or threshold selection, a tension rarely acknowledged. We synthesise AL research across acoustic and image modalities and identify gaps and opportunities. Most studies evaluate query strategies on pre-labelled benchmarks with simulated annotators; deployments in real monitoring workflows are rare and concentrate on birds and cetaceans. Bats, insects, amphibians, and fish are underrepresented, and multimodal applications remain largely unexplored. Evaluation centres on headline reductions in annotation effort, often without random-sampling baselines, per-class results, or calibration analysis, and rarely accounts for the labels required for validation. We provide a tutorial treatment of the AL loop that makes these budget decisions explicit, and a roadmap towards AL methods that support label-efficient training, validation, and trustworthy downstream ecological inference.
查看缓存全文
缓存时间: 2026/09/24 09:41
# Active Learning for Biodiversity Monitoring: From Label Efficiency to Reliable Ecological Inference
Source: [https://arxiv.org/html/2609.27409](https://arxiv.org/html/2609.27409)
###### Abstract
Limited expert annotation capacity is a pervasive constraint in biodiversity monitoring\. Passive acoustic recorders and camera traps generate raw data faster than experts can analyse them\. Machine learning \(ML\) models can process these data at scale, but their reliability depends on the quality, quantity, and coverage of labelled samples, thus expert time continues to be a constraint\. Active learning \(AL\) eases this bottleneck by selecting, under a fixed annotation budget, the samples expected to improve a model most\. The published evidence shows that AL can reduce the number of labels needed to reach a target level of predictive performance\. Monitoring programmes, however, face a broader question: how should a limited expert budget be divided so that model training, model validation, and the ecological estimates built on model outputs all remain reliable? Because AL selects samples non\-randomly, the labels it produces are unsuitable for validation, calibration, or threshold selection, and this tension is rarely acknowledged\. We synthesise AL research across acoustic, image modalities, and identify several gaps and opportunities\. Most published works evaluate query strategies on pre\-labelled benchmark datasets with simulated annotators; deployments embedded in real monitoring workflows are rare and concentrate on birds and cetaceans\. Bats, insects, amphibians, and fish are underrepresented, and multimodal applications remain largely unexplored\. Evaluation practice centres on headline reductions in annotation effort, often without random\-sampling baselines, per\-class results, or calibration analysis, and almost never accounts for the labels required for validation\. We provide a tutorial treatment of the AL loop that makes these budget decisions explicit, and a roadmap towards AL methods that support not only label\-efficient training but also validation and trustworthy downstream ecological inference\.
###### keywords
Active Learning; Bioacoustics; Biodiversity; Camera trap; Machine Learning
††articletype:REVIEW††affiliation:aInstitute for Biodiversity and Ecosystem Dynamics \(IBED\), University of Amsterdam, Amsterdam, the Netherlands††affiliation:bFaculty of Information Technology and Communication Sciences, Tampere University, Finland††affiliation:cLeiden Institute of Advanced Computer Science, Leiden University, The Netherlands††affiliation:dNaturalis Biodiversity Centre, Leiden, The Netherlands## 1Introduction
Passive acoustic recorders and camera traps are expanding the spatial coverage and temporal extent of biodiversity monitoring\. These approaches produce large volumes of recordings, images, and molecular observations, yet converting raw observations into species detections and ecological measures still depend heavily on expert annotation\. Machine learning \(ML\) can process these data at scale, but model reliability depends on the quality, quantity, and coverage of labelled samples\([Stowell, 2022](https://arxiv.org/html/2609.27409#bib.bib44);[Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66);[Kitzes et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib43)\)\. Expert time therefore becomes a scarce resource in the monitoring workflow, and many projects can label only a small fraction of the data they collect\([Kholghi et al\., 2018](https://arxiv.org/html/2609.27409#bib.bib23);[Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66);[van Osta et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib18)\)\(Figure[1](https://arxiv.org/html/2609.27409#S1.F1)\)\. This constraint creates a crucial allocation problem:when the pool of candidate data exceeds the annotation budget, which data is most valuable to label?
Figure 1:In\-situ monitoring devices significantly increase the spatial and temporal resolution of monitoring initiatives yet they also increase data volumes\. Ecological inference is constrained by both the information captured due to study design and the efficiency with which that information is extracted and validated\.Active learning \(AL\) addresses this problem by using the current model and its representation of the data to guide subsequent sample selection\([Settles, 2009](https://arxiv.org/html/2609.27409#bib.bib13)\)\. It aims to concentrate expert effort on samples expected to improve a specified learning objective, such as reducing classification error or improving recognition of rare classes\([McEwen et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib32);[Kurinchi\-Vendhan and Beery, 2026](https://arxiv.org/html/2609.27409#bib.bib12)\)\. Annotation, model updating, and further querying form an iterative cycle \(Figure[2](https://arxiv.org/html/2609.27409#S1.F2)\) in which the allocation of a limited annotation budget adapts to the model’s needs\. Studies in acoustic and image\-based monitoring have used this approach to reduce the number of labels required to train species classifiers and detectors\([Kholghi et al\., 2018](https://arxiv.org/html/2609.27409#bib.bib23);[Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66);[van Osta et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib18)\)\. AL thus converts annotation from a one\-off data preparation step into a iterative decision process within model development\.
Figure 2:Overview of active learning pipeline\.Label efficiency, however, is not the only requirement that biodiversity monitoring places on annotation\. Ecological data have long\-tailed class distributions, strong spatial and temporal variation, heterogeneous observation conditions, and per\-label annotation costs that vary widely\. More fundamentally, the expert time that funds model training comes from the same budget as the expert time required for model validation and the verification of detections used for ecological estimates\. A query strategy can improve classification performance while shifting the labelled sample away from the population distribution\([Farquhar et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib21)\); labels acquired this way cannot be reused for validation, calibration or threshold selection without correction\. Evaluating predictive performance at a given label count therefore says little about whether the occupancy, abundance, or phenology estimates built on the resulting model can be trusted\. In short, existing biodiversity AL research largely demonstrates that fewer labels can suffice for predictive performance, whereas monitoring programmes need to know how a fixed expert budget should support reliable training, validation, and ecological inference together \(Figure[3](https://arxiv.org/html/2609.27409#S1.F3)\)\. This review is organised around that gap\.
Figure 3:Model training and validation both draw from the same budget and are generally not interchangeable resulting in a dual allocation problem\.𝒰\\mathcal\{U\}denotes the pool of unlabelled samples, andBtrainB\_\{\\mathrm\{train\}\}andBvalB\_\{\\mathrm\{val\}\}the shares of the expert budget spent on training and validation labels \(Section[2\.2](https://arxiv.org/html/2609.27409#S2.SS2)\)\.Relevant studies are dispersed across acoustic and image modalities and across methodological communities, and in most cases evaluate query strategies on pre\-labelled benchmark datasets with simulated oracle annotation\. Ecological practitioners therefore find it difficult to judge whether a method fits a particular monitoring objective or to compare reported gains across studies\. We synthesise 164 studies of active learning and machine learning for biodiversity monitoring\.
### 1\.1Contributions
This review makes three contributions\.
- •A tutorial treatment of the active learning loop written for ecologists and field biologists, making explicit where each stage of the loop consumes expert effort and where design choices affect the reliability of later validation and ecological inference\.
- •A systematic map of 164 studies across acoustic and image modalities, showing that the literature concentrates on label\-efficient training evaluated on pre\-labelled benchmarks, and that validation practice, budget accounting, and downstream ecological objectives receive little attention\.
- •A roadmap for closing this gap, covering evaluation standards, explicit budget allocation between training and validation, cost\-aware annotation, and the treatment of query\-induced sampling bias in ecological inference\.
The remainder of this paper is structured as follows\. Section[2](https://arxiv.org/html/2609.27409#S2)introduces the AL loop in tutorial form and identifies where each stage draws on the expert budget\. Section[3](https://arxiv.org/html/2609.27409#S3)describes the survey method\. Section[4](https://arxiv.org/html/2609.27409#S4)maps the state of current literature\. Section[5](https://arxiv.org/html/2609.27409#S5)analyses why label\-efficient training alone does not deliver reliable monitoring and sets out a roadmap\.
## 2A Budget\-Aware Introduction to Active Learning
This section provides a working introduction for readers who plan to use AL within a monitoring workflow\. Unlike general treatments, we track a single quantity throughout: the expert budget\. Every stage of the loop spends from it, and design choices at some stages determine whether the rest of the budget can be spent effectively\. Readers seeking the general theory and method families of AL are referred to[Settles \(2009\)](https://arxiv.org/html/2609.27409#bib.bib13)and[Ren et al\. \(2021\)](https://arxiv.org/html/2609.27409#bib.bib81)\.
### 2\.1Definitions
Amodelin this review is a function, generally a deep neural network, that maps an input \(an audio segment or an image\) to an output \(predicted classes with confidence scores\)\. Internally, the network converts the raw input into a set of numerical descriptors calledfeatures\. The vector of features produced by the final layers of the network is called anembedding\. Samples that lie close together in embedding space are treated as similar by the model, so embeddings provide a geometry over the data that many AL methods exploit\. Apretrained modelhas been trained beforehand on a large dataset, as with BirdNET\([Kahl et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib91)\)\. Itsencoder, the part that produces embeddings, can befrozen\(its weights held fixed\) and used only to embed the data, with a lightweight classification head trained on top\.
This review concerns deep learning: most of the surveyed studies use deep neural networks, and concepts such as embeddings and frozen encoders presuppose them\. The general AL framework is also applicable to most ML methods\.
Active learning\(AL\) is an iterative process of model\-guided sample selection and label acquisition, which aims to reach a specified learning objective with fewer labelled samples than passive \(random\) sampling\([Settles, 2009](https://arxiv.org/html/2609.27409#bib.bib13)\)\. The most common objective is the discriminative performance of the model, but broader objectives such as annotation cost or calibration fit the same framework\. The rule that scores and selects samples is thequery strategy\(Section[2\.4](https://arxiv.org/html/2609.27409#S2.SS4)\)\.
Label efficiencyis the number of labelled samples required to reach a specified level of performance, or equivalently the performance reached for a given number of labels\. AL improves label efficiency when it reaches the target with fewer labels than random sampling\. The learning curve \(Section[4\.4](https://arxiv.org/html/2609.27409#S4.SS4)\) is its standard visualisation\.
Theexpert budgetis the total amount of expert annotation effort a monitoring programme can afford, counted in labels or in expert hours\. Every stage of the AL loop draws on it\.
Pool\-basedAL assumes that a large collection of unlabelled samples, the pool, is available before annotation starts, as with an archive of recordings or camera\-trap images\. The query strategy scores the whole pool and selects from it\. Instream\-basedAL, samples arrive one at a time, for example on a recording device, and the decision whether to query each one is made on arrival\. Almost all biodiversity AL is pool\-based, matching the offline archives that passive sensors produce\. A separate distinction concerns how many samples are queried per cycle:sequentialAL queries one sample and updates the model before the next query, whereasbatch\-modeAL queries a batch of samples per cycle \(Section[2\.3](https://arxiv.org/html/2609.27409#S2.SS3)\)\.
AL is frequently paired withhuman\-in\-the\-loop\(HITL\) workflows\([Monarch, 2021](https://arxiv.org/html/2609.27409#bib.bib68)\)\. The two are related but distinct\. HITL refers to any closed\-loop process in which people participate in model development or verification, whereas AL refers specifically to model\-guided sample selection\.
AL as defined above selects labels for modeltraining\.Active testingapplies the same idea to modelvalidation: it selects which samples an expert should verify so that model performance can be estimated to a target precision with fewer labels\([Kossen et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib14)\)\. Both draw on the same expert budget, an allocation challenge discussed in Section[2\.2](https://arxiv.org/html/2609.27409#S2.SS2)\.
### 2\.2One Expert Budget, Two Demands
Biodiversity monitoring consumes expert effort at several stages of a pipeline: deploying sensors across ecological gradients, annotating detected signals, and validating and interpreting model outputs\. This review concerns the latter two stages, which share one allocation problem: reduce uncertainty about a quantity of interest with as few expert observations as possible\. Active learning addresses model training: it selects which samples to label so that the model improves most efficiently\([Settles, 2009](https://arxiv.org/html/2609.27409#bib.bib13)\)\. Active testing addresses model validation: it selects which model outputs an expert should verify to characterise model performance most precisely under a fixed budget\([Kossen et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib14)\), or stops verification once the confidence interval is acceptably narrow\([Perez et al\., 2024a](https://arxiv.org/html/2609.27409#bib.bib84);[Perez et al\., 2024b](https://arxiv.org/html/2609.27409#bib.bib80)\)\. The two demands share one objective, maximising information per unit of expert effort, and they compete for the same budget\. WritingBBfor the total number of expert labels \(or expert hours\) a programme can afford,B=Btrain\+BvalB=B\_\{\\mathrm\{train\}\}\+B\_\{\\mathrm\{val\}\}, whereBtrainB\_\{\\mathrm\{train\}\}is spent on training labels andBvalB\_\{\\mathrm\{val\}\}on validation labels\. How the budget is divided is itself a design decision, and it is the decision this review keeps returning to\.
A related idea,adaptive sampling, applies the same principle to field deployment itself, reallocating monitoring effort towards locations or periods where ecological utility is highest\([Balantic and Donovan, 2019](https://arxiv.org/html/2609.27409#bib.bib16);[Alawad, 2022](https://arxiv.org/html/2609.27409#bib.bib15)\)\. Adaptive sampling is absent from the reports and deployments we survey, and requires specific circumstances such as mobile monitoring tools\. We therefore leave it outside the scope of this review, noting only that adaptive sampling would place additional effects on the same expert budget\.
The budget question extends beyond which samples to label for training\. A query strategy selects a sample precisely because it is unusual under the current model, so the labelled set it produces is not a random sample of the monitored population\([Farquhar et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib21)\)\. Uncertainty sampling, for example, concentrates on decision\-boundary cases: error rates estimated on such a set are systematically pessimistic, and calibration curves or detection thresholds fitted to it do not transfer to the full data stream\. Labels bought through active learning therefore serve training well but validation poorly, and the two purposes require separately budgeted data\([Lowell et al\., 2019](https://arxiv.org/html/2609.27409#bib.bib17);[Kitzes et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib43)\)\.
The distinction matters most when model outputs feed ecological analyses\. Occupancy, abundance, and phenology estimates inherit the error characteristics of the detections they are built on, and therefore require validated error rates across the relevant sites, seasons, and classes rather than a single global accuracy figure\([Kitzes et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib43)\)\. Throughout this review we accordingly apply a simple test:an AL study isbudget\-completeonly if it accounts forBvalB\_\{\\mathrm\{val\}\}, the labels needed to validate the resulting model, alongsideBtrainB\_\{\\mathrm\{train\}\}\. Section[5](https://arxiv.org/html/2609.27409#S5)returns to the mechanics of query\-induced bias and to methods for validating under a shared budget\.
### 2\.3The Active Learning Loop
A full loop consists of initialisation, inference, querying, annotation, training, and a stopping criterion \(Figure[4](https://arxiv.org/html/2609.27409#S2.F4)\)\. We describe each stage in turn and note where it draws on the expert budget\. In our notation,ftf\_\{t\}denotes the model at cyclett,𝒰\\mathcal\{U\}the pool of unlabelled samples,ℒt\\mathcal\{L\}\_\{t\}the accumulated labelled set on whichftf\_\{t\}is trained,𝒬t⊂𝒰\\mathcal\{Q\}\_\{t\}\\subset\\mathcal\{U\}the batch queried in cyclett, andb=\|𝒬t\|b=\|\\mathcal\{Q\}\_\{t\}\|the batch size\.
Figure 4:Active learning cycle showing active fine\-tuning of model, inference, querying and expert labelling\. At cyclett, the model trained on the labelled setℒt\\mathcal\{L\}\_\{t\}is run over the unlabelled pool𝒰\\mathcal\{U\}, and the query method selects a batch𝒬t\\mathcal\{Q\}\_\{t\}ofbbsamples for expert labelling\.#### Initialisation\.
Sample selection is guided by model outputs, so the quality of selection depends on the discriminative capability of the initial modelf0f\_\{0\}\. For an uninitialised model, early selections can be poor \(the “cold\-start problem”\)\. The initialisation \(warm\-up\) phase typically is the first phase of training on a small labelled setℒ0\\mathcal\{L\}\_\{0\}, which is generally randomly sampled\([Lindholm et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib33);[Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66)\)or may already exist\.
It has become increasingly common to use a frozen pretrained encoder and apply AL to the training of a lightweight classification head \(Section[4\.2](https://arxiv.org/html/2609.27409#S4.SS2)\)\. In this case, other initialisation methods can diversify sampling across the embedding space, such as Core\-Set \(farthest\-first traversal\)\([Sener and Savarese, 2018](https://arxiv.org/html/2609.27409#bib.bib99)\), sampling of representative high\-density regions such as TypiClust\([Hacohen et al\., 2022](https://arxiv.org/html/2609.27409#bib.bib98)\), compared on bioacoustic datasets by[Rauch et al\. \(2024\)](https://arxiv.org/html/2609.27409#bib.bib30), or k\-means\+\+\([Arthur and Vassilvitskii, 2007](https://arxiv.org/html/2609.27409#bib.bib19)\), applied to aquatic invasive species recognition by[Chowdhury et al\. \(2024\)](https://arxiv.org/html/2609.27409#bib.bib65)\. When labelled samples are present but very scarce, distance\-based sampling can be applied to prototypical class embeddings \(the mean of a class support set\)\([McEwen et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib32)\)\. With increasingly capable pretrained models it is also common to skip initialisation and apply informativeness\-based or hybrid methods directly\. Clearer guidance on when initialisation utility and method are needed\. A reasonable default is that a pretrained encoder covering the target classes removes the need for a separate warm\-up\([Tamkin et al\., 2022](https://arxiv.org/html/2609.27409#bib.bib78)\)\.
#### Inference\.
The current modelftf\_\{t\}is run as a single forward pass over the unlabelled pool \(e\.g\. audio segments or images\), producing for each sample a predicted class, a confidence score, and an embedding\. The model’s confidence scores and embeddings are then used to inform the query strategy\.
#### Querying\.
The query strategy selects a subset𝒬t\\mathcal\{Q\}\_\{t\}of the unlabelled pool to present to the annotator, ranked by estimated utility\. Utility is typically decomposed into two complementary properties:informativeness, the expected reduction in model uncertainty if the sample were labelled, andnon\-redundancy, the degree to which a sample adds information not already represented by the other samples selected in the same cycle\.
In sequential AL, one sample is queried per cycle \(b=1b=1\) and the model is updated before the next query\. This maximises informativeness but is computationally prohibitive at the scales typical of passive biodiversity monitoring, both because of the volume of data collected and because of the infrastructural difficulty of closing the loop between querying, model updates, and human annotators\. Sequential AL is therefore uncommon in the biodiversity literature and viable only in simulation, through oracle sampling111Oracle labelling is the process of revealling pre\-labelled samples, simulating the querying and expert annotation process\.of pre\-labelled data\.
Batch\-mode AL, which selectsb\>1b\>1samples per cycle, is more commonly applied\. The model is updated once after the batch is annotated, trading sample efficiency for computational efficiency\. Note the batch sizebbrefers to the number of samples selected per cycle,notthe batch size used for model training\. Although performance generally degrades at larger batch sizes\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22);[Zhang et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib82)\), the batch size is more often set by feasibility constraints or simple heuristics\. Common batch sizes range from 20\([McEwen et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib32);[Kath et al\., 2024c](https://arxiv.org/html/2609.27409#bib.bib37)\)to 100\([Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66);[Qian et al\., 2017](https://arxiv.org/html/2609.27409#bib.bib46)\)\. Some studies use larger batches, such as the 1,712 images per cycle of[Nguyen and Nguyen \(2025\)](https://arxiv.org/html/2609.27409#bib.bib73), and work outside biodiversity extends to extreme batches of 100k to 1M images\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22)\)\. Ablation studies exploring batch size are limited\([Zhang et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib82)\)\.[Mononen et al\. \(2025\)](https://arxiv.org/html/2609.27409#bib.bib70)set the batch size from the number of predicted classes \(50 samples per class\)\. Evidence from the general AL literature suggests that dynamic batch sizes and batch scheduling, starting small and growing, could improve performance\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22)\), but this remains largely unexplored\. Given the resource and computational constraints of biodiversity monitoring, batch size and large\-batch AL merit further investigation\.
Batch mode matches offline, pool\-based annotation workflows but introduces a redundancy problem: samples selected independently by an uncertainty criterion may be near\-duplicates, wasting annotation budget\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22)\)\. Diversity\-aware strategies \(Section[2\.4](https://arxiv.org/html/2609.27409#S2.SS4)\) address this by enforcing coverage of the unlabelled pool within each batch\.
#### Annotation\.
The selected samples𝒬t\\mathcal\{Q\}\_\{t\}are presented to a domain expert, typically a trained ecologist or taxon specialist, who provides labels, so each cycle spendsbblabels ofBtrainB\_\{\\mathrm\{train\}\}\. These labels are generally treated as ground truth, although inter\-annotator disagreement is sometimes considered\([van Osta et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib18)\)\. The per\-sample annotation burden in biodiversity monitoring varies substantially: a clearly vocalising target species takes seconds to confirm, while an ambiguous vocalisation, a rare species, or an occluded image may require extensive analysis\. Ambiguous samples sometimes need to be skipped\([Mononen et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib70)\)or labelled as uncertain\([van Osta et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib18)\)\. Samples acquired to optimise model sample efficiency can therefore come at the cost of annotation efficiency, since query strategies often select ambiguous cases\.
The annotation task itself also varies in form and duration: confirming a species\-level prediction is simpler than drawing a bounding box on an image or spectrogram, and providing strong multi\-label annotations is more cumbersome than weaker single\-label annotations\([Martinsson et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib62)\)\. This variability motivates methods that reduce the cost of each label independently of sample selection, such as model\-guided weak labels\([Martinsson et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib29)\); we return to these in Section[5\.3](https://arxiv.org/html/2609.27409#S5.SS3)\.
In AL benchmarking and method development, the annotation process is simulated throughoracle labelling: AL runs on a pre\-labelled dataset and labels are revealed after querying\. As a consequence, under oracle labelling, annotation cost is largely ignored\.
#### Training\.
Following annotation, the queried batch joins the labelled set,ℒt\+1=ℒt∪𝒬t\\mathcal\{L\}\_\{t\+1\}=\\mathcal\{L\}\_\{t\}\\cup\\mathcal\{Q\}\_\{t\}, and the model is fine\-tuned on the full accumulated labelled set to giveft\+1f\_\{t\+1\}\. When only a classification head is updated on a frozen backbone, the embedding space remains fixed across cycles\. When the full network is fine\-tuned\([Mohaimenuzzaman et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib47);[Chowdhury et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib65)\), training and querying become directly coupled: as model weights update, the embedding space shifts, reshaping what the query strategy selects in the next cycle\.[Mohaimenuzzaman et al\. \(2023\)](https://arxiv.org/html/2609.27409#bib.bib47)demonstrate this in bioacoustic AL, showing that updating the feature extractor within the loop outperforms static\-feature approaches, and attribute the gain to the representation adapting to the newly labelled distribution\.[Tamkin et al\. \(2022\)](https://arxiv.org/html/2609.27409#bib.bib78)show that AL effectiveness is an emergent property of representation quality: pretrained models with linearly separable feature spaces require up to five times fewer labels, and hence a lowerBtrainB\_\{\\mathrm\{train\}\}, than their non\-pretrained counterparts\.
#### Stopping criterion\.
The annotation budget defines the maximum annotation effort available across all AL cycles; in biodiversity monitoring it is typically constrained by expert availability\. Reported budgets vary considerably, from a few hundred samples to tens of thousands for large camera\-trap datasets\([Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66)\), reflecting the diversity of task difficulty, pool size, and deployment context\. Cycles terminate when the budget is exhausted, when a fixed number of iterations is reached\([Chowdhury et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib65)\), or when a performance threshold is met\([Williams et al\., 2025a](https://arxiv.org/html/2609.27409#bib.bib58)\)\. The last approach is preferable in principle because it adapts to the difficulty of the problem, but it requires a held\-out validation set to monitor performance\. When labels are scarce, partitioning them for validation is costly;[Tamkin et al\. \(2022\)](https://arxiv.org/html/2609.27409#bib.bib78)propose stopping when the training loss falls to a fixed fraction of its initial value, avoiding a separate validation set\. In practice, stopping criteria are rarely reported explicitly in the biodiversity AL literature: most studies run until a fixed budget is exhausted or report results at a predetermined number of cycles, which makes cross\-study comparison of label efficiency difficult\. Notably, a performance\-based stopping criterion is the first point inside the loop where validation labels must be spent\.
### 2\.4Query Strategy Families
Query strategies can be grouped into three families by the signal used to rank samples\.Uncertainty\-basedstrategies use the current model’s predictions: least\-confidence, margin, and entropy criteria rely on class probabilities, while ensemble or Bayesian approximations additionally estimate model \(epistemic\) uncertainty at a higher computational cost\([Settles, 2009](https://arxiv.org/html/2609.27409#bib.bib13);[Ren et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib81)\)\. Uncertainty sampling concentrates queries near decision boundaries, which is efficient for refining a classifier but is also the source of the sampling bias discussed in Section[5](https://arxiv.org/html/2609.27409#S5)\.
Geometry\-basedstrategies use only the structure of the embedding space: diversity methods such as Core\-Set maximise the distance between samples within a batch\([Sener and Savarese, 2018](https://arxiv.org/html/2609.27409#bib.bib99)\); coverage methods ensure that all regions of the pool are represented; and typicality methods such as TypiClust prioritise high\-density regions\([Hacohen et al\., 2022](https://arxiv.org/html/2609.27409#bib.bib98)\)\. Because these strategies do not depend on reliable class probabilities, they are particularly useful in cold\-start and low\-budget settings\.
Hybridstrategies combine informativeness with diversity and are the default choice in batch\-mode settings because they suppress within\-batch redundancy\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22);[Zhang and Virtanen, 2025](https://arxiv.org/html/2609.27409#bib.bib42);[Kath et al\., 2024c](https://arxiv.org/html/2609.27409#bib.bib37)\)\. Beyond these three families, the ranking objective itself can change: cost\-aware strategies include the expected annotation cost of each sample in its utility, and ecological\-utility strategies select samples by the expected reduction in uncertainty of a downstream ecological quantity rather than model accuracy, for example sampling in service of species distribution models\([Lange et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib24)\)or stratification across spatial, temporal, and ecological gradients\([McEwen et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib34)\)\. This last class of objectives is the most directly relevant to biodiversity monitoring and the least studied\.
## 3Review Methodology
### 3\.1Survey design
An initial set of searches was conducted on the 27th of November 2025 using Google Scholar, covering the passive monitoring modalities relevant to biodiversity monitoring \(acoustics and camera traps\)\. The search queries for each modality are listed below\. No date range was specified\. The number of relevant papers and pre\-prints increases over time \(Figure[5](https://arxiv.org/html/2609.27409#S3.F5)\)\.
Figure 5:Number of published papers related to active learning and biodiversity monitoring over time\.#### Bioacoustics:
553 search results with192 selectedprovided by the following query: \("active learning" OR "query learning" OR "selective sampling" OR "human\-in\-the\-loop"\) AND \(bioacoust\* OR ecoacoust\* OR vocali\* OR “animal calls” OR "passive acoustic monitoring" OR "soundscape"\) AND \(animal OR bird\* OR cetacean\* OR insect\* OR mammal\*\)
#### Camera trap:
571 search results with171 selectedprovided by the following query: \("active learning" OR "query learning" OR "selective sampling" OR "human\-in\-the\-loop"\) AND \("camera trap\*"\) AND \(animal\* OR wildlife OR mammal\* OR bird\* OR "biodiversity monitoring"\)
The searches returned1128 resultsin total\. After a preliminary review of abstracts,367 paperswere selected for further review and exported from Google Scholar\. After deduplication and removal of irrelevant literature \(books, theses and inaccessible papers\),156 papersentered comprehensive review\. A supplementary search was conducted on the 21st of April 2026 using Claude Opus 4\.6 \(in research mode\) with the Consensus AI and BioRxiv connectors enabled, prompted with the existing queries and paper list; this identified8 additional relevant papers\. The final corpus comprises164 papers\.
#### Schema:
Following the identification of relevant literature, a schema of 24 questions across seven categories was developed: paper metadata \(e\.g\. preprint status\), AL configuration \(e\.g\. query strategy, budget\), performance and evaluation \(e\.g\. evaluation metrics\), application domain \(e\.g\. modality, taxonomic coverage\), technical details \(e\.g\. base model, datasets\), context \(e\.g\. motivation and use case\), and summary\. The full schema is provided \(Appendices[A](https://arxiv.org/html/2609.27409#A1)\)\. The schema gave a consistent framework for evaluating the reviewed papers\. Generative AI \(Claude\) was used to aid the extraction of schema information; all information reported in this review has been manually reviewed and verified by the authors\.
## 4The State of the Field
We organise the literature review around four questions: which data and taxa are studied \(Section[4\.1](https://arxiv.org/html/2609.27409#S4.SS1)\), which models the loop is built on \(Section[4\.2](https://arxiv.org/html/2609.27409#S4.SS2)\), which query strategies are used in practice \(Section[4\.3](https://arxiv.org/html/2609.27409#S4.SS3)\), and how success is evaluated and supported by tools \(Sections[4\.4](https://arxiv.org/html/2609.27409#S4.SS4)and[4\.5](https://arxiv.org/html/2609.27409#S4.SS5)\)\. The answers to all four point to the same conclusion: the experimental designs of the current literature demonstrate label\-efficient training, rather than the budget\-complete validation and inference that monitoring and ecological inference requires\.
### 4\.1Modalities, Taxa, and Datasets
The current literature is heavily weighted toward methods papers that reuse existing ecological datasets \(BirdSet[Rauch et al\. \(2025b\)](https://arxiv.org/html/2609.27409#bib.bib11), AudioSet[Gemmeke et al\. \(2017\)](https://arxiv.org/html/2609.27409#bib.bib8), Snapshot Serengeti[Swanson et al\. \(2015\)](https://arxiv.org/html/2609.27409#bib.bib9), AnuraSet[Cañas et al\. \(2023\)](https://arxiv.org/html/2609.27409#bib.bib10)\)\. A small number of corpora recur, with different query strategies evaluated on the same underlying recordings\. This reuse, while valuable for the development and benchmarking of AL methods, does inflate taxonomic coverage\. We therefore distinguish throughout between benchmark coverage, where a taxon is present only because it is contained in a reused dataset, and deployment coverage, where active learning is embedded in a workflow that answers a question about a real population\. Deployment is the rarer case: 20 of the 66 taxon\-related studies \(30%\) report an applied deployment rather than a method or tool demonstration\. These deployments are concentrated in terrestrial mammals \(9 studies, mostly conservation camera\-trap programmes for endangered species\) and birds \(7\), followed by marine mammals \(4\) and fish \(3\)\.
Figure 6:a\) Data modality split across active learning literature and b\) taxonomic split across 66 taxon\-specific active learning papers\.#### Taxonomic Coverage
Assigning each of the 66 taxon\-related studies, the largest group is multi\-taxon \(35%\), which are evaluated on multi\-species benchmarks\. Terrestrial mammals follow at 26% and birds at 20%, then marine mammals \(8%\), fish \(5%\), freshwater species \(3%\), and the remaining studies focus on marine invertebrates and plants\. Bats are absent from AL literature\. Insects and anurans are almost entirely confined to the multi\-taxon studies: insects are the dominant taxon in no study despite appearing in seven, and work on anurans is entirely constrained to AnuraSet\([Cañas et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib10)\)\.
Birds are the most\-investigated taxon and the one with the clearest record of field deployment\. On the methods side,[Rauch et al\. \(2024\)](https://arxiv.org/html/2609.27409#bib.bib30)benchmark entropy\-based and hybrid query strategies on BirdSet, and[Qian et al\. \(2017\)](https://arxiv.org/html/2609.27409#bib.bib46)compare uncertainty\- and diversity\-based selection across sixty species\. Deployments pair strategies with monitoring targets:[McEwen et al\. \(2025\)](https://arxiv.org/html/2609.27409#bib.bib34)apply stratified uncertainty sampling over BirdNET embeddings in the transnational TABMON network[Cretois et al\. \(2026\)](https://arxiv.org/html/2609.27409#bib.bib20),[Ayers et al\. \(2021\)](https://arxiv.org/html/2609.27409#bib.bib28)prioritise recordings for more than 150 threatened coastal species,[Bellafkir et al\. \(2023\)](https://arxiv.org/html/2609.27409#bib.bib54)drive edge\-deployed recognisers of 214 species using ensemble reliability scores,[van Osta et al\. \(2023\)](https://arxiv.org/html/2609.27409#bib.bib18)develop a recogniser for the endangered southern black\-throated finch\.
Terrestrial mammal studies occupy a similar share of the corpus but with a different emphasis\. Much of the volume is benchmark\-driven, with[Norouzzadeh et al\. \(2021\)](https://arxiv.org/html/2609.27409#bib.bib66)comparing query strategies on Snapshot Serengeti and NACTI and[Mononen et al\. \(2025\)](https://arxiv.org/html/2609.27409#bib.bib70)applying core\-selection to global camera\-trap data, while genuine deployments pair strategies with sites, including[Miao et al\. \(2021\)](https://arxiv.org/html/2609.27409#bib.bib25)at Gorongosa National Park and[Bothmann et al\. \(2023\)](https://arxiv.org/html/2609.27409#bib.bib67)and[Auer et al\. \(2021\)](https://arxiv.org/html/2609.27409#bib.bib71)in the Bavarian Forest\. A distinct sub\-strand applies active learning to individual re\-identification[Kulits et al\. \(2021\)](https://arxiv.org/html/2609.27409#bib.bib75);[Brust et al\. \(2020\)](https://arxiv.org/html/2609.27409#bib.bib76);[Sani et al\. \(2025\)](https://arxiv.org/html/2609.27409#bib.bib72)\. The apparent richness of mammalian coverage is concentrated in a handful of reused megafauna and savanna\-community datasets rather than distributed across mammalian monitoring contexts\.
Coverage of the remaining groups is thinner and, in several cases, an artefact of dataset reuse\. Amphibian coverage is entirely attributable to AnuraSet, which supplies the frog and toad classes evaluated by[Kath et al\. \(2024c\)](https://arxiv.org/html/2609.27409#bib.bib37)and[Kath et al\. \(2024d\)](https://arxiv.org/html/2609.27409#bib.bib31)and appears as a testbed in generalisation studies\([McEwen et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib34)\); we found no study in which active learning was deployed to monitor a real anuran population\. Marine mammals are the second group \(after birds\) with a substantive deployment record, through humpback whale song\([Allen et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib53)\), seasonal Bryde’s whale calls\([Allen et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib63)\), and blue and fin whale call\-density estimation\([Alksne et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib26)\), with benchmark work on the BioSED pilot\-whale task\([Zhang and Virtanen, 2025](https://arxiv.org/html/2609.27409#bib.bib42)\)\. Insect coverage is incidental, arising through the BioSED mosquito task\([Zhang and Virtanen, 2025](https://arxiv.org/html/2609.27409#bib.bib42)\), cicada and cricket sounds in long\-duration soundscapes\([Kholghi et al\., 2018](https://arxiv.org/html/2609.27409#bib.bib23)\), and image\-based work on bees\([Boiński and Szymański, 2020](https://arxiv.org/html/2609.27409#bib.bib86)\), even though reviews of emerging monitoring technology repeatedly flag insects as a priority\([Van Klink et al\., 2022](https://arxiv.org/html/2609.27409#bib.bib61);[Sheard et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib64)\)\. Fish coverage is also low\([Bordoux et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib36);[Williams et al\., 2025a](https://arxiv.org/html/2609.27409#bib.bib58);[Dumoulin et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib40)\)\. Freshwater systems, which we separate from marine ones because their monitoring conditions differ, amount to two studies: semi\-automated selection for endangered river dolphins in the Brazilian Amazon\([Erbs et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib60)\), the only freshwater deployment in the corpus, and contrastive representation learning with k\-means selection for invasive dreissenid mussel larvae\([Chowdhury et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib65)\)\. Reptiles appear only as minor classes within multi\-taxa camera\-trap datasets, and no study takes them as a target\. Bats are absent, which is notable given that they are among the most intensively acoustically monitored taxa\.
#### Modalities
Acoustic data accounts for 53% of the studies with an identifiable sensing modality and camera\-trap or other image data for 42%, leaving 5% across everything else \(Figure[6](https://arxiv.org/html/2609.27409#S4.F6)\)\. Those remaining studies show active learning reaching sensing contexts the field has otherwise left alone\.[Kang et al\. \(2021\)](https://arxiv.org/html/2609.27409#bib.bib87)exploit temporal and spatial proximity structure to surface rare hummingbird events in a 9,000 GB field video archive\.[Kozlova et al\. \(2025\)](https://arxiv.org/html/2609.27409#bib.bib59)apply Bayesian active learning by disagreement to behavioural segmentation from pose and video features, shifting the annotation target from species identity to behaviour\.[Lange et al\. \(2023\)](https://arxiv.org/html/2609.27409#bib.bib24)estimate species ranges from geographic coordinates and environmental covariates, using a weighted committee of pre\-trained per\-species models to select where to sample next across a thousand mixed taxa\.[Tamkin et al\. \(2022\)](https://arxiv.org/html/2609.27409#bib.bib78), counted here as image work, applies one selection method to both camera\-trap imagery and NLP tasks\.[van Ommen Kloeke et al\. \(2025\)](https://arxiv.org/html/2609.27409#bib.bib85)is the only multi\-sensor entry, combining wildlife cameras, acoustic recorders, insect cameras and radar in a national monitoring infrastructure, though it describes interactive training cycles as an intended capability rather than reporting a query strategy in use\.
### 4\.2Models
Deep neural networks are the base model in over 94% of the surveyed studies\. In much of the recent bioacoustic and camera\-trap literature, the starting point has shifted from training task\-specific classifiers from scratch towards adapting pretrained encoders, detectors, or classifiers\. The governing question is increasingly how to adapt a pretrained model or its classification head to a new domain \(for example, a new location or recording condition\) or a new application\([Dumoulin et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib40)\)\. In bioacoustics, commonly used base models include BirdNET\([Kahl et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib91)\), the Perch family of birdsong embeddings\([Ghani et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib50);[Burns et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib56)\), self\-supervised animal vocalisation encoders such as AVES and its avian variant BirdAVES\([Hagiwara, 2023](https://arxiv.org/html/2609.27409#bib.bib92)\), animal2vec\([Schäfer\-Zimmermann et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib93)\), masked\-autoencoder models adapted to audio\([Rauch et al\., 2025a](https://arxiv.org/html/2609.27409#bib.bib52)\), and general\-purpose audio encoders such as BEATs\([Chen et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib94)\); a recent comparative review contrasts these models\([Schwinger et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib45)\)\. For image\-based camera\-trap monitoring, MegaDetector provides a widely used animal, person, and vehicle detector\([Beery et al\., 2019](https://arxiv.org/html/2609.27409#bib.bib95);[Hernandez et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib96)\), complemented by visual foundation models such as BioCLIP\([Stevens et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib89)\)and multimodal models for zero\-shot species recognition\([Fabian et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib74)\)\. These systems differ in role, as classifiers, embedding models, or detectors, but all provide reusable predictions or representations\. Pretrained embeddings can transfer across locations, species, and in some cases taxa, making them an effective basis for downstream tasks under limited supervision\([Ghani et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib50);[Ghani et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib55);[Williams et al\., 2025b](https://arxiv.org/html/2609.27409#bib.bib48);[Burns et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib56)\)\.
Two recurring technical routes are visible in the literature\. The first is end\-to\-end active learning, in which the full network is trained or fine\-tuned within the loop\([Rauch et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib49)\)\. The second uses a frozen pretrained encoder to embed the candidate pool and applies AL to a lightweight classification head fitted on the resulting embeddings\([Rauch et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib30);[Kath et al\., 2024d](https://arxiv.org/html/2609.27409#bib.bib31);[Kath et al\., 2024b](https://arxiv.org/html/2609.27409#bib.bib41);[Kath et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib38);[Zhang and Virtanen, 2025](https://arxiv.org/html/2609.27409#bib.bib42);[Lindholm et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib33);[Bernard et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib51)\)\. The frozen\-encoder route is increasingly common: for a fixed candidate pool, embeddings can be cached, reducing the cost of repeated encoder inference\. The acquisition step and classifier retraining can still be expensive for very large pools, but the approach is complementary with transfer learning because both reduce the quantity of labelled data required\. Embeddings and predictions generated by pretrained models have therefore become an important starting point for AL in this field\.
When AL is built on a pretrained model, the task becomes selecting, from a large unlabelled pool, the most informative subset to annotate for adapting that pretrained model under a fixed budget\([Xie et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib97)\)\. The encoder may be frozen at its initial values or updated together with the classification head, but when the labelled seed set is absent or very small, the uncertainty signals that drive pool\-based AL are weak\. Cold\-start selection therefore often relies on the geometry of the embedding space through coverage, diversity, and typicality criteria before uncertainty\-based or hybrid strategies become reliable\([Sener and Savarese, 2018](https://arxiv.org/html/2609.27409#bib.bib99);[Hacohen et al\., 2022](https://arxiv.org/html/2609.27409#bib.bib98)\)\. Related work reframes efficient AL through proxies and feature alignment in the pretrained\-model setting\([Wen et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib88)\)and examines how AL interacts with pretrained models\([Tamkin et al\., 2022](https://arxiv.org/html/2609.27409#bib.bib78);[Bodesheim et al\., 2022](https://arxiv.org/html/2609.27409#bib.bib69)\)\. As pretrained models become a common substrate for biodiversity monitoring, the open problem shifts in part from which classifier to train towards which data to fine\-tune on\. Downstream performance is sensitive to the choice of embedding\([Dumoulin et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib40)\), yet principled criteria for selecting an embedding to drive AL remain to be established; we identify this as an open question\.
### 4\.3Query Strategies in Use
In applied studies, uncertainty sampling is the default method for selecting samples\([Qian et al\., 2017](https://arxiv.org/html/2609.27409#bib.bib46);[Kholghi et al\., 2018](https://arxiv.org/html/2609.27409#bib.bib23);[Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66)\): it is simple to implement and requires only model confidences as input\. It carries a hidden cost, however: entropy\- and confidence\-based criteria tend to select acoustically or visually confusing samples, and precisely these samples are the hardest to annotate, sometimes having to be skipped or labelled as uncertain\([Mononen et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib70);[van Osta et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib18)\)\. Maximising the expected information per label can therefore simultaneously maximise the annotation cost per label\.
A clear pattern is that monitoring applications benefit from hybrid methods that combine high\-utility sampling with diversification\. There are two reasons\. First, batch sizes in monitoring are large, so the implicit diversification provided by sequential or small\-batch settings is unavailable\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22)\)\. Second, as noted above, pure uncertainty criteria accumulate near\-duplicate, ambiguous samples within a batch\. Strategy comparisons on biodiversity data, with benchmarks on BirdSet, BioSED, and AnuraSet, support this observation\([Rauch et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib30);[Zhang and Virtanen, 2025](https://arxiv.org/html/2609.27409#bib.bib42);[Kath et al\., 2024c](https://arxiv.org/html/2609.27409#bib.bib37)\)\.
A third usage pattern falls outside the textbook, model training usecase but recurs in practice: exploratory sampling, in which the model guides experts to the target events themselves within a large volume of data, such as calls of rare or invasive species, or hard negatives\([McEwen et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib32);[Kather et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib27);[van Osta et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib57);[Alksne et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib26)\)\. Related as well are vector\-based similarity searches which is one component within the "agile modelling" workflow\([Dumoulin et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib40)\)\. The goal is not to improve a classifier evenly but to confirm as many target detections as possible in limited time\. This pattern is closest to the real motivation of many monitoring programmes, and it again illustrates that improvement in model accuracy is not the only quantity practitioners care about\.
### 4\.4Evaluation Practice
The dominant evaluation protocol reports alearning curve: a performance metric measured as a function of the annotation budget \(the number or proportion of labelled samples\) \(Figure[7](https://arxiv.org/html/2609.27409#S4.F7)\)\. When a passive baseline is included, the AL strategy is compared against random sampling under the same model, budget, and train/test split\. Performance is summarised using standard classification or detection metrics, including accuracy, macro\- and micro\-averagedF1F\_\{1\}, mean average precision, and area under the ROC or precision\-recall curve; rare\-class recall is sometimes reported where class imbalance is severe\([Kath et al\., 2024c](https://arxiv.org/html/2609.27409#bib.bib37);[Zhang and Virtanen, 2025](https://arxiv.org/html/2609.27409#bib.bib42);[Rauch et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib30)\)\. Two common summary values are widely reported: the reduction in labelling effort required to reach a target performance, and the performance gain at a fixed budget\. In applied studies, the headline figure is often the percentage reduction in annotation effort\([Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66);[Miao et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib25)\)\. Among the studies in our corpus that report a random sampling baseline, reductions reach 89%\([Moller et al\., 2017](https://arxiv.org/html/2609.27409#bib.bib83)\), with a mean of 64% \( n=7\)\. Gains reported at a fixed budget are not directly comparable across studies, since they are expressed in whichever metric the study adopts\.
Figure 7:Active passive learning curves, showing performance ceiling and Area Under the Learning Curve \(AULC\) for a fixed budget\. The horizontal axis is the number of labelled samples\|ℒt\|\|\\mathcal\{L\}\_\{t\}\|, which reaches the pool size\|𝒰\|\|\\mathcal\{U\}\|once the whole pool is labelled\.This headline figure can mislead when interpreted in isolation\. Reduction percentages are inflated on large, redundant datasets: when a multi\-million\-image camera\-trap corpus contains many near\-duplicate frames, repeated backgrounds, or otherwise easy samples, a small labelled fraction may suffice even under random sampling\. A reduction exceeding 99% may then reflect the redundancy of the pool and the chosen performance threshold as much as the efficacy of the query strategy\([Norouzzadeh et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib66)\); the absolute reduction conflates the informativeness of the selected samples with the redundancy of the candidate pool\. A second weakness is the absence of standardisation: studies report different metrics on different datasets, base models, budgets, and random seeds, which prevents comparison across papers\. A third weakness is that a single global metric obscures the quantities most relevant to ecological inference: per\-class performance for underrepresented taxa, model calibration, and generalisation under spatiotemporal domain shift\([McEwen et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib34)\)\.
Evaluation is inherently multi\-dimensional, and absolute annotation reduction captures only one dimension\.Selection efficiency, the computational cost of the acquisition step itself, is rarely reported, yet it determines practical applicability: many strategies scale poorly since they require repeated end\-to\-end retraining, Monte Carlo sampling, or pairwise distance computations over the candidate pool, and per\-cycle latency directly affects the usability of human\-in\-the\-loop systems\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22)\)\. Thescale of dataa method can operate over is a second dimension: a strategy validated on a small benchmark may not transfer to the pool sizes of real deployments\([Citovsky et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib22);[Beck et al\., 2023](https://arxiv.org/html/2609.27409#bib.bib77)\)\.Annotation cost realismis a third: counting samples ignores the variable cost of annotation, which differs across modalities and between weak and strong labels; we return to this in Section[5\.3](https://arxiv.org/html/2609.27409#S5.SS3)\.
The random \(passive\) baseline is the single most important reference point\. Its absence from a benchmarking study is a critical flaw because AL does not always outperform random sampling\([Kholghi et al\., 2018](https://arxiv.org/html/2609.27409#bib.bib23);[Moller et al\., 2017](https://arxiv.org/html/2609.27409#bib.bib83);[Kath et al\., 2024c](https://arxiv.org/html/2609.27409#bib.bib37);[Lowell et al\., 2019](https://arxiv.org/html/2609.27409#bib.bib17)\): under cold\-start conditions, large\-batch selection, or distribution shift, query strategies can match or underperform random selection\([Mittal et al\., 2019](https://arxiv.org/html/2609.27409#bib.bib100)\)\.[Kholghi et al\. \(2018\)](https://arxiv.org/html/2609.27409#bib.bib23)compared seven query strategies for annotating long\-duration environmental recordings and found margin sampling to be the only strategy that never outperformed the baseline;[Moller et al\. \(2017\)](https://arxiv.org/html/2609.27409#bib.bib83)report uncertainty sampling underperforming random selection over much of the learning curve; and[Kath et al\. \(2024c\)](https://arxiv.org/html/2609.27409#bib.bib37)report a speedup factor of 1\.0 for a purely diversity\-based strategy, indicating no gain over random selection\. In both directions these failures are single\-criterion strategies, while the largest gains come from hybrid acquisition functions combining informativeness with diversity, a pattern also seen in the recent BioDCASE challenge\([McEwen et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib3)\)\. Reporting performance relative to repeated random baselines is the minimum standard for demonstrating that a gain stems from the selection mechanism itself; a single\-run baseline is insufficient, since the reported performance of random sampling varies by as much as 13% between studies using the same dataset and settings\([Ren et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib81)\)\. A threshold\-free summary makes the comparison more portable: either the area under the learning curve expressed as a ratio to the random baseline over a stated budget range, or the speedup factor of[Kath et al\. \(2024c\)](https://arxiv.org/html/2609.27409#bib.bib37), the fraction of samples an AL strategy requires relative to random sampling, which is invariant to training set size but assumes both strategies converge to the same ceiling performance\. Threshold\-based reductions should state and motivate the threshold\. Publishing the paired learning curves themselves allows any summary statistic to be recomputed\. A representative validation or test set must likewise be kept separate from actively selected training samples, because active selection changes the sampled distribution and can bias validation, calibration, and threshold selection\([Farquhar et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib21)\)\. In the literature we reviewed, validation data are usually inherited from benchmark splits, and their annotation cost is almost never counted in the reported budget; by the standard of Section[2\.2](https://arxiv.org/html/2609.27409#S2.SS2), few if any studies are budget\-complete\.
### 4\.5Tools and Frameworks
Existing software falls into two groups\. The first comprises research\-oriented AL libraries such as scikit\-activeML\([Herde et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib90)\)and comparable toolkits \(BaseAL[McEwen and Zhang \(2026\)](https://arxiv.org/html/2609.27409#bib.bib7), modAL[Danka and Horvath \(2018\)](https://arxiv.org/html/2609.27409#bib.bib6), the DeepAL family[Huang \(2021\)](https://arxiv.org/html/2609.27409#bib.bib5)\): they implement a wide range of query strategies but rely on simulated oracle labelling and lack integration with real annotation workflows, likely because of the complexity of aligning diverse task requirements with different data formats\. The second comprises annotation platforms that embed AL or HITL assistance, such as AIDE for camera traps\([Kellenberger et al\., 2020](https://arxiv.org/html/2609.27409#bib.bib79)\)and annotation tooling for passive acoustic monitoring\([Kath et al\., 2024a](https://arxiv.org/html/2609.27409#bib.bib35);[McEwen et al\., 2024](https://arxiv.org/html/2609.27409#bib.bib32)\)\. The space between the two groups is the missing loop\-closing infrastructure noted in Section[2\.3](https://arxiv.org/html/2609.27409#S2.SS3): no platform has been widely adopted\. We argue that what is needed is not more tools but the integration of these methods into the annotation platforms ecologists already use\.
The methods surveyed here are applied across different datasets, base models, annotation budgets and modalities, which makes results difficult to compare between studies\. Shared tasks address this by fixing the data, the budget and the evaluation protocol so that acquisition strategies are compared under identical conditions\. The firstActive Learning for Bioacousticstask was run in 2026\([McEwen et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib3)\)as part of BioDCASE\([Stowell et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib4)\), with participants implementing acquisition functions\([Magaldi and Dubus, 2026](https://arxiv.org/html/2609.27409#bib.bib2);[Dubus et al\., 2026](https://arxiv.org/html/2609.27409#bib.bib1)\)over terrestrial and marine mammal datasets within the BaseAL framework\([McEwen and Zhang, 2026](https://arxiv.org/html/2609.27409#bib.bib7)\)\.222The authors organised this task and developed BaseAL\.Version 1\.2 of that framework reports annotation and computational cost alongside accuracy for each acquisition method, and supports active testing, so that the labels used for evaluation are selected under a budget rather than assumed to be available\. Nothing in the format is specific to bioacoustics, and the same structure could be applied to camera\-trap data and other data modalities\.
### 4\.6Summary
Taken together, the preceding subsections describe a field configured to demonstrate a single result: that a model of a given quality can be reached with fewer labels than random sampling would require\. The datasets are pre\-labelled benchmarks, the annotator is a simulated oracle, the dominant strategies optimise classifier improvement, and the dominant metric is the fraction of labels saved\. The questions a monitoring programme must answer before acting on model outputs, namely how good the model actually is on its operating distribution and how reliable the ecological estimates built on it are, sit outside this experimental design\.
## 5From Label Efficiency to Reliable Inference
### 5\.1Query\-Induced Sampling Bias
Query strategies improve label efficiency by oversampling informative regions of the input space, and this is exactly what makes the resulting labelled set unrepresentative\([Farquhar et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib21)\)\. Uncertainty sampling concentrates near decision boundaries, so difficult cases are overrepresented: performance estimated on these labels understates the model, and thresholds or calibration maps fitted to them misrepresent the score distributions of the full data stream\. The bias is not a side effect of poor implementation; it is the intended behaviour of the selection mechanism\.
Part of this bias can be corrected\. When samples are selected with known probabilities, importance\-weighted estimators recover unbiased risk estimates, and corrected estimators exist for pool\-based AL\([Farquhar et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib21)\)\. In practice the correction is fragile: many strategies are deterministic top\-kkselections without well\-defined inclusion probabilities, batch selection couples samples, and weights become extreme when the proposal distribution is far from the population\. The dependable alternative is structural: reserve an independently drawn, representative validation set, and treat its size as part of the budget rather than as an afterthought\.
A further problem is amplified in monitoring settings: data outlive models\. Actively selected datasets are coupled to the model and strategy that produced them, and the gains do not reliably transfer when a different architecture is trained on the same labels\([Lowell et al\., 2019](https://arxiv.org/html/2609.27409#bib.bib17)\)\. Monitoring archives are reused across model generations, so datasets built through AL should, at a minimum, record the selection mechanism in their metadata so that later users can assess and mitigate bias\.
### 5\.2Validating Under a Shared Budget
What validation requires is a labelled sample set whose distribution matches the conditions under which the model will operate, sized to give acceptable precision for the quantities that matter, and stratified enough to expose per\-class and per\-site failure\. The stopping criterion \(Section[2\.3](https://arxiv.org/html/2609.27409#S2.SS3)\), threshold selection, and every downstream analysis all require validation data\. Yet in the reviewed literature, validation data are usually inherited from benchmark splits and their label cost is almost never counted against the reported budget; outside pre\-labelled benchmarks, that cost falls squarely on a monitoring programme’s own experts\.
The validation budget can itself be spent efficiently\. Active testing selects which model outputs an expert should verify so that performance estimates reach a target precision with fewer verified labels\([Kossen et al\., 2021](https://arxiv.org/html/2609.27409#bib.bib14)\), and sequential variants stop verification once confidence intervals are acceptably narrow\([Perez et al\., 2024a](https://arxiv.org/html/2609.27409#bib.bib84);[Perez et al\., 2024b](https://arxiv.org/html/2609.27409#bib.bib80)\)\. Stratified designs allocate verification effort across locations, seasons, or classes\([McEwen et al\., 2025](https://arxiv.org/html/2609.27409#bib.bib34)\)\. A practical recipe follows from this: a two\-part budget,BtrainB\_\{\\mathrm\{train\}\}spent on actively selected labels for training andBvalB\_\{\\mathrm\{val\}\}on representatively drawn \(or actively tested\) labels for evaluation, with the split between the two reported\.
### 5\.3Active Annotation
The preceding sections treat AL as a question of which samples to label\. The cost of each label, however, also depends on how the annotator interacts with the selected sample\. It is therefore useful to decompose AL into two components\.Active samplingis the selection of samples for annotation, the component the rest of this review refers to simply as AL\.Active annotationconcerns how the selected samples are labelled efficiently, for example by presenting the expert with tentative machine labels, rough bounding boxes, or onset and offset times to correct\. Both components aim to reduce labelling effort; the AL literature concentrates on the first and treats the cost of each label as fixed\.
Active annotation methods include change\-point detection and tentative bounding boxes[Martinsson et al\. \(2024\)](https://arxiv.org/html/2609.27409#bib.bib29), extraction and ordering of segments[McEwen et al\. \(2024\)](https://arxiv.org/html/2609.27409#bib.bib32);[McEwen et al\. \(2023\)](https://arxiv.org/html/2609.27409#bib.bib39), human\-in\-the\-loop tools for annotating passive acoustic monitoring datasets[Kath et al\. \(2024a\)](https://arxiv.org/html/2609.27409#bib.bib35), and a broad category of human\-interface considerations that improve annotation efficiency\. Because the expert budget is measured in time, savings from active annotation add to savings from active sampling, and a budget accounted in expert time captures both\. A key challenge for human\-in\-the\-loop methods is the data infrastructure and tooling for closed\-loop, iterative labelling\. While various tools exist, they have not been widely adopted\.
## 6Conclusions
Machine learning has enabled radical acceleration of various types of monitoring in ecology\. However, it relies on trained models and thus retains a bottleneck of expert labelling for training and validating systems\. This bottleneck is important as it is often a rate\-limiting factor in applied biodiversity monitoring\. Active learning, reviewed here, offers the best way to resolve this issue, selecting data in ways which theoretically and empirically make the best use of ecologists’ time budgets\.
Active learning is most often tested in a traditional way: across acoustic and image applications, fewer labels can suffice to reach a given level of predictive performance\. This reduction can be dramatic and represents a key benefit of AL over classical training of recognition algorithms\. Our map of 164 studies, however, shows that this evidence comes almost entirely from simulations on pre\-labelled benchmarks\. The annotator is an oracle, the validation set is free, and the fraction of labels saved is the headline figure\. Monitoring programmes face a broader problem that a limited expert budget must simultaneously support model training, model validation, and the ecological estimates built on model outputs, and AL gains its efficiency precisely by distorting the sampling distribution, which disqualifies the labels it buys from serving the latter two purposes\.
Some of this can be acted on now\. Practitioners can reserve an independently drawn, representative validation set and count it in the budget, report results against repeated random baselines, record the selection mechanism in dataset metadata, and measure costs in expert time rather than sample counts\. The research agenda is also clear: query objectives tailored to ecological quantities, formal budget allocation across training and validation, correction methods for query\-induced bias in ecological inference, and the nearly empty territories of bats, insects, fish, and multimodal monitoring\. The standard by which AL becomes genuinely useful for biodiversity monitoring is not how many labels it saves, but whether, under a fixed expert budget, it delivers a trustworthy model and trustworthy ecological inference\.
## Acknowledgement\(s\)
BM and DS are funded by Biodiversa\+, the European Biodiversity Partnership, in the context of the Towards a Transnational Acoustic Biodiversity MOnitoring Network \(TABMON\) project under the 2022\-2023 BiodivMon joint call\. It was co\-funded by the European Commission \(GA ref\. 101052342\) and the following funding organisations: Norwegian Research Council \(project number 350977\), l’Agence Nationale de la Recherche \(ANR\-23\-EBIP\-0010\), l’Office français de la biodiversité \(OFB\-23\-1865\), Dutch Research Council \(2023/NWA/01580460\), and la Agencia Estatal de Investigación \(PCI2024\-153427\)\. SZ is funded by the Bioacoustic AI Doctoral Network \(European Union Marie Skłodowska\-Curie Action\) under grant agreement No 101116715\.
## Disclosure statement
The author\(s\) declare no competing interests\. Generative AI was used to aid in the extraction of information from literature and was used overall to reformulate some sentences and paragraphs\. All AI generate content has been manually checked by the authors\. Usage of AI is detailed in the manuscript\.
## Notes on contributor\(s\)
BM conceptualised the review, curated the literature and produced the visualisations\. BM and SZ conducted the analysis and wrote the manuscript, contributing equally to the writing\. DS reviewed and edited the manuscript\.
## References
- Alawad \(2022\)F\. AlawadAdaptive sampling for efficient acoustic noise monitoring: an incremental learning approach\.In2022 IEEE International Conferences on Internet of Things \(iThings\) and IEEE Green Computing & Communications \(GreenCom\) and IEEE Cyber, Physical & Social Computing \(CPSCom\) and IEEE Smart Data \(SmartData\) and IEEE Congress on Cybermatics \(Cybermatics\),pp\. 176–184\.Cited by:[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p2.1)\.
- Alksneet al\.\(2026\)M\. N\. Alksne, M\. A\. Roch, K\. E\. Frasier, J\. A\. Hildebrand, S\. Andres, D\. Adesanya, L\. M\. Baggett, J\. M\. Jones, X\. Liu, A\. Širović,et al\.Application of deep learning to estimate blue and fin whale call density in the southern california current ecosystem\.preprint\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p3.1)\.
- Allenet al\.\(2021\)A\. N\. Allen, M\. Harvey, L\. Harrell, A\. Jansen, K\. P\. Merkens, C\. C\. Wall, J\. Cattiau, and E\. M\. OlesonA convolutional neural network for automated detection of humpback whale song in a diverse, long\-term passive acoustic dataset\.Frontiers in Marine Science8,pp\. 607321\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Allenet al\.\(2024\)A\. N\. Allen, M\. Harvey, L\. Harrell, M\. Wood, A\. R\. Szesciorka, J\. L\. McCullough, and E\. M\. OlesonBryde’s whales produce biotwang calls, which occur seasonally in long\-term acoustic recordings from the central and western north pacific\.Frontiers in Marine Science11,pp\. 1394695\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Arthur and Vassilvitskii \(2007\)D\. Arthur and S\. VassilvitskiiK\-means\+\+: the advantages of careful seeding\.InSoda,Vol\.7,pp\. 1027–1035\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p2.1)\.
- Aueret al\.\(2021\)D\. Auer, P\. Bodesheim, C\. Fiderer, M\. Heurich, and J\. DenzlerMinimizing the annotation effort for detecting wildlife in camera trap images with active learning\.InINFORMATIK 2021,pp\. 547–564\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1)\.
- Ayerset al\.\(2021\)J\. Ayers, S\. Perry, V\. Tiwari, M\. Blue, N\. Balaji, C\. Schurgers, R\. Kastner, M\. Tobler, and I\. IngramReducing the barriers of acquiring ground\-truth from biodiversity rich audio datasets using intelligent sampling techniques\.InTackling Climate Change with AI Workshop, Conference on Neural Information Processing Systems,Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p2.1)\.
- Balantic and Donovan \(2019\)C\. Balantic and T\. DonovanTemporally adaptive acoustic sampling to maximize detection across a suite of focal wildlife species\.Ecology and Evolution9\(18\),pp\. 10582–10600\.Cited by:[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p2.1)\.
- Becket al\.\(2023\)N\. Beck, S\. Kothawade, P\. Shenoy, and R\. IyerStreamline: streaming active learning for realistic multi\-distributional settings\.arXiv preprint arXiv:2305\.10643\.Cited by:[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p3.1)\.
- Beeryet al\.\(2019\)S\. Beery, D\. Morris, and S\. YangEfficient pipeline for camera trap image review\.arXiv preprint arXiv:1907\.06772\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Bellafkiret al\.\(2023\)H\. Bellafkir, M\. Vogelbacher, D\. Schneider, M\. Mühling, N\. Korfhage, and B\. FreislebenEdge\-based bird species recognition via active learning\.InInternational Conference on Networked Systems,pp\. 17–34\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p2.1)\.
- Bernardet al\.\(2025\)C\. Bernard, B\. McEwen, B\. Cretois, H\. Glotin, D\. Stowell, and R\. MarxerData\-driven sampling strategies for fine\-tuning bird detection models\.bioRxiv,pp\. 2025–10\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1)\.
- Bodesheimet al\.\(2022\)P\. Bodesheim, J\. Blunk, M\. Koerschens, C\. Brust, C\. Kaeding, and J\. DenzlerPre\-trained models are not enough: active and lifelong learning is important for long\-term visual monitoring of mammals in biodiversity research—individual identification and attribute prediction with image features from deep neural networks and decoupled decision models applied to elephants and great apes\.Mammalian Biology102\(3\),pp\. 875–897\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p3.1)\.
- Boiński and Szymański \(2020\)T\. Boiński and J\. SzymańskiCollaborative data acquisition and learning support\.InInternational Conference on Computer Information Systems and Industrial Management,pp\. 220–229\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Bordouxet al\.\(2026\)V\. Bordoux, C\. Parcerisas, J\. Martin, D\. Elisabeth, A\. J\. Murk, and R\. M\. van der VenRapid fish sound detection using human\-in\-the\-loop deep learning\.Ecological Informatics,pp\. 103874\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Bothmannet al\.\(2023\)L\. Bothmann, L\. Wimmer, O\. Charrakh, T\. Weber, H\. Edelhoff, W\. Peters, H\. Nguyen, C\. Benjamin, and A\. MenzelAutomated wildlife image classification: an active learning tool for ecological applications\.Ecological Informatics77,pp\. 102231\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1)\.
- Brustet al\.\(2020\)C\. Brust, C\. Käding, and J\. DenzlerActive and incremental learning with weak supervision\.KI\-Künstliche Intelligenz34\(2\),pp\. 165–180\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1)\.
- Burnset al\.\(2025\)A\. Burns, L\. Harrell, B\. van Merriënboer, V\. Dumoulin, J\. Hamer, and T\. DentonPerch 2\.0 transfers ’whale’ to underwater tasks\.InThe Thirty\-Ninth Annual Conference on Neural Information Processing Systems workshop: AI for non\-human animal communication,Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Cañaset al\.\(2023\)J\. S\. Cañas, M\. P\. Toro\-Gómez, L\. S\. M\. Sugai, H\. D\. Benítez Restrepo, J\. Rudas, B\. Posso Bautista, L\. F\. Toledo, S\. Dena, A\. H\. R\. Domingos, F\. L\. de Souza,et al\.A dataset for benchmarking neotropical anuran calls identification in passive acoustic monitoring\.Scientific Data10\(1\),pp\. 771\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.p1.1)\.
- Chenet al\.\(2023\)S\. Chen, Y\. Wu, C\. Wang, S\. Liu, D\. Tompkins, Z\. Chen, W\. Che, X\. Yu, and F\. WeiBEATs: audio pre\-training with acoustic tokenizers\.InProceedings of the 40th International Conference on Machine Learning \(ICML\),Vol\.202,pp\. 5178–5193\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Chowdhuryet al\.\(2024\)S\. Chowdhury, G\. Hamerly, and M\. McGarrityActive learning strategy using contrastive learning and k\-means for aquatic invasive species recognition\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp\. 848–858\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p2.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Citovskyet al\.\(2021\)G\. Citovsky, G\. DeSalvo, C\. Gentile, L\. Karydas, A\. Rajagopalan, A\. Rostamizadeh, and S\. KumarBatch active learning at scale\.Advances in Neural Information Processing Systems34,pp\. 11933–11944\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p4.1),[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p3.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p3.1)\.
- Cretoiset al\.\(2026\)B\. Cretois, C\. M\. Rosten, J\. Wiel, C\. Barile, B\. McEwen, C\. Bernard, M\. P\. Boom, G\. Bota, L\. Brotons, E\. Serrano\-Davies,et al\.TABMON: design and deployment of a transnational passive acoustic monitoring network for european birds\.Methods in Ecology and Evolution\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p2.1)\.
- Danka and Horvath \(2018\)T\. Danka and P\. HorvathModAL: a modular active learning framework for python\.arXiv preprint arXiv:1805\.00979\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p1.1)\.
- Dubuset al\.\(2026\)G\. Dubus, H\. Magaldi, and A\. Gros\-MartialAdaptive diversity\-uncertainty active learning with redundancy control for bioacoustic event classification\.arXiv preprint arXiv:2607\.04868\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p2.1)\.
- Dumoulinet al\.\(2025\)V\. Dumoulin, O\. Stretcu, J\. Hamer, L\. Harrell, R\. Laber, H\. Larochelle, B\. van Merriënboer, A\. Navine, P\. Hart, B\. Williams,et al\.The search for squawk: agile modeling in bioacoustics\.arXiv preprint arXiv:2505\.03071\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p3.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p3.1)\.
- Erbset al\.\(2023\)F\. Erbs, M\. Gaona, M\. van der Schaar, S\. Zaugg, E\. Ramalho, D\. Houser, and M\. AndréTowards automated long\-term acoustic monitoring of endangered river dolphins: a case study in the brazilian amazon floodplains\.Scientific reports13\(1\),pp\. 10801\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Fabianet al\.\(2023\)Z\. Fabian, Z\. Miao, C\. Li, Y\. Zhang, Z\. Liu, A\. Hernández, A\. Montes\-Rojas, R\. Escucha, L\. Siabatto, A\. Link,et al\.Multimodal foundation models for zero\-shot animal species recognition in camera trap images\.arXiv preprint arXiv:2311\.01064\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Farquharet al\.\(2021\)S\. Farquhar, Y\. Gal, and T\. RainforthOn statistical bias in active learning: how and when to fix it\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p3.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1),[§5\.1](https://arxiv.org/html/2609.27409#S5.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.27409#S5.SS1.p2.1)\.
- Gemmekeet al\.\(2017\)J\. F\. Gemmeke, D\. P\. Ellis, D\. Freedman, A\. Jansen, W\. Lawrence, R\. C\. Moore, M\. Plakal, and M\. RitterAudio set: an ontology and human\-labeled dataset for audio events\.In2017 IEEE international conference on acoustics, speech and signal processing \(ICASSP\),pp\. 776–780\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.p1.1)\.
- Ghaniet al\.\(2023\)B\. Ghani, T\. Denton, S\. Kahl, and H\. KlinckGlobal birdsong embeddings enable superior transfer learning for bioacoustic classification\.Scientific Reports13\(1\),pp\. 22876\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Ghaniet al\.\(2025\)B\. Ghani, V\. J\. Kalkman, B\. Planqué, W\. Vellinga, L\. Gill, and D\. StowellImpact of transfer learning methods and dataset characteristics on generalization in birdsong classification\.Scientific Reports15\(1\),pp\. 16273\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Hacohenet al\.\(2022\)G\. Hacohen, A\. Dekel, and D\. WeinshallActive learning on a budget: opposite strategies suit high and low budgets\.InProceedings of the 39th International Conference on Machine Learning \(ICML\),Vol\.162,pp\. 8175–8195\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p2.1),[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p2.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p3.1)\.
- Hagiwara \(2023\)M\. HagiwaraAVES: animal vocalization encoder based on self\-supervision\.InICASSP 2023 – 2023 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 1–5\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Herdeet al\.\(2025\)M\. Herde, M\. T\. Pham, D\. Kottke, A\. Benz, L\. Lührs, P\. Mergard, C\. Sandrock, J\. Cheng, A\. Roghman, M\. Müjde,et al\.Scikit\-activeml: a comprehensive and user\-friendly active learning library\.pre\-print\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p1.1)\.
- Hernandezet al\.\(2024\)A\. Hernandez, Z\. Miao, L\. Vargas, S\. Beery, R\. Dodhia, P\. Arbelaez, and J\. M\. Lavista FerresPytorch\-Wildlife: a collaborative deep learning framework for conservation\.arXiv preprint arXiv:2405\.12930\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Huang \(2021\)K\. HuangDeepal: deep active learning in python\.arXiv preprint arXiv:2111\.15258\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p1.1)\.
- Kahlet al\.\(2021\)S\. Kahl, C\. M\. Wood, M\. Eibl, and H\. KlinckBirdNET: a deep learning solution for avian diversity monitoring\.Ecological Informatics61,pp\. 101236\.Cited by:[§2\.1](https://arxiv.org/html/2609.27409#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Kanget al\.\(2021\)D\. Kang, A\. Derhacobian, K\. Tsuji, T\. Hebert, P\. Bailis, T\. Fukami, T\. Hashimoto, Y\. Sun, and M\. ZahariaExploiting proximity search and easy examples to select rare events\.InNeurIPS Data\-Centric AI Workshop 2021,Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px2.p1.1)\.
- Kathet al\.\(2024a\)H\. Kath, T\. S\. Gouvêa, and D\. SonntagA human\-in\-the\-loop tool for annotating passive acoustic monitoring datasets\.InGerman Conference on Artificial Intelligence \(Künstliche Intelligenz\),pp\. 341–345\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p1.1),[§5\.3](https://arxiv.org/html/2609.27409#S5.SS3.p2.1)\.
- Kathet al\.\(2024b\)H\. Kath, T\. S\. Gouvêa, and D\. SonntagActive and transfer learning for efficient identification of species in multi\-label bioacoustic datasets\.InProceedings of the 2024 International Conference on Information Technology for Social Good,pp\. 22–25\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1)\.
- Kathet al\.\(2024c\)H\. Kath, T\. S\. Gouvêa, and D\. SonntagActive learning in multi\-label classification of bioacoustic data\.InGerman Conference on Artificial Intelligence \(Künstliche Intelligenz\),pp\. 114–127\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1),[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p3.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p1.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1)\.
- Kathet al\.\(2025\)H\. Kath, T\. S\. Gouvêa, and D\. SonntagSpeeding up bioacoustic data analysis: fine\-tuning deep models with active learning for efficient wildlife detection\.InProceedings of the 2025 International Conference on Information Technology for Social Good \(GoodIT\),pp\. 200–203\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1)\.
- Kathet al\.\(2024d\)H\. Kath, P\. P\. Serafini, I\. B\. Campos, T\. S\. Gouvêa, and D\. SonntagLeveraging transfer learning and active learning for data annotation in passive acoustic monitoring of wildlife\.Ecological Informatics82,pp\. 102710\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1)\.
- Katheret al\.\(2024\)V\. Kather, F\. Seipel, B\. Berges, G\. Davis, C\. Gibson, M\. Harvey, L\. Henry, A\. Stevenson, and D\. RischDevelopment of a machine learning detector for north atlantic humpback whale song\.The Journal of the Acoustical Society of America155\(3\),pp\. 2050–2064\.Cited by:[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p3.1)\.
- Kellenbergeret al\.\(2020\)B\. Kellenberger, D\. Tuia, and D\. MorrisAIDE: accelerating image\-based ecological surveys with interactive machine learning\.Methods in Ecology and Evolution11\(12\),pp\. 1716–1727\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p1.1)\.
- Kholghiet al\.\(2018\)M\. Kholghi, Y\. Phillips, M\. Towsey, L\. Sitbon, and P\. RoeActive learning for classifying long\-duration audio recordings of the environment\.Methods in Ecology and Evolution9\(9\),pp\. 1948–1958\.Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p1.1),[§1](https://arxiv.org/html/2609.27409#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p1.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1)\.
- Kitzeset al\.\(2026\)J\. Kitzes, L\. Chronister, C\. Czarnecki, C\. Fiss, L\. Freeland\-Haynes, B\. D\. Goodman, S\. Lapp, R\. P\. Lyon, H\. Nossan, T\. A\. Rhinehart, S\. Ruiz Guzman, A\. Syunkova, and L\. ViottiIntegrating AI models into ecological research workflows: the case of terrestrial bioacoustics\.Methods in Ecology and Evolution17\(2\),pp\. 257–271\.External Links:[Document](https://dx.doi.org/10.1111/2041-210X.70133)Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p3.1),[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p4.1)\.
- Kossenet al\.\(2021\)J\. Kossen, S\. Farquhar, Y\. Gal, and T\. RainforthActive testing: sample\-efficient model evaluation\.InInternational Conference on Machine Learning,pp\. 5753–5763\.Cited by:[§2\.1](https://arxiv.org/html/2609.27409#S2.SS1.p8.1),[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p1.1),[§5\.2](https://arxiv.org/html/2609.27409#S5.SS2.p2.1)\.
- Kozlovaet al\.\(2025\)E\. Kozlova, A\. Bonnetto, and A\. MathisDLC2Action: a deep learning\-based toolbox for automated behavior segmentation\.bioRxiv,pp\. 2025–09\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px2.p1.1)\.
- Kulitset al\.\(2021\)P\. Kulits, J\. Wall, A\. Bedetti, M\. Henley, and S\. BeeryElephantBook: a semi\-automated human\-in\-the\-loop system for elephant re\-identification\.InProceedings of the 4th ACM SIGCAS Conference on Computing and Sustainable Societies,pp\. 88–98\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1)\.
- Kurinchi\-Vendhan and Beery \(2026\)R\. Kurinchi\-Vendhan and S\. BeeryFinding needles in the haystack: transductive active labeling in ecology\.arXiv preprint arXiv:2606\.03821\.Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p2.1)\.
- Langeet al\.\(2023\)C\. Lange, E\. Cole, G\. Van Horn, and O\. Mac AodhaActive learning\-based species range estimation\.Advances in neural information processing systems36,pp\. 41892–41913\.Cited by:[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p3.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px2.p1.1)\.
- Lindholmet al\.\(2025\)R\. Lindholm, O\. Marklund, O\. Mogren, and J\. MartinssonAggregation strategies for efficient annotation of bioacoustic sound events using active learning\.arXiv preprint arXiv:2503\.02422\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1)\.
- Lowellet al\.\(2019\)D\. Lowell, Z\. C\. Lipton, and B\. C\. WallacePractical obstacles to deploying active learning\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 21–30\.Cited by:[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p3.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1),[§5\.1](https://arxiv.org/html/2609.27409#S5.SS1.p3.1)\.
- Magaldi and Dubus \(2026\)H\. Magaldi and G\. DubusDeterminantal point process sampling for bioacoustic active learning\.arXiv preprint arXiv:2607\.06063\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p2.1)\.
- Martinssonet al\.\(2024\)J\. Martinsson, O\. Mogren, M\. Sandsten, and T\. VirtanenFrom weak to strong sound event labels using adaptive change\-point detection and active learning\.In2024 32nd European Signal Processing Conference \(EUSIPCO\),pp\. 902–906\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px4.p2.1),[§5\.3](https://arxiv.org/html/2609.27409#S5.SS3.p2.1)\.
- Martinssonet al\.\(2025\)J\. Martinsson, T\. Virtanen, M\. Sandsten, and O\. MogrenThe accuracy cost of weakness: a theoretical analysis of fixed\-segment weak labeling for events in time\.arXiv preprint arXiv:2502\.09363\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px4.p2.1)\.
- McEwenet al\.\(2025\)B\. McEwen, C\. Bernard, and D\. StowellStratified active learning for spatiotemporal generalisation in bioacoustic monitoring\.bioRxiv,pp\. 2025–09\.Cited by:[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p3.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p2.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p2.1),[§5\.2](https://arxiv.org/html/2609.27409#S5.SS2.p2.1)\.
- McEwenet al\.\(2026\)B\. McEwen, R\. Kurinchi\-Vendhan, S\. Zhang, L\. Rauch, M\. Herde, and S\. BeeryBioDCASE: active learning for bioacoustics\.arXiv preprint arXiv:2609\.15255\.Cited by:[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1),[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p2.1)\.
- McEwenet al\.\(2023\)B\. McEwen, K\. Soltero, S\. Gutschmidt, A\. Bainbridge\-Smith, J\. Atlas, and R\. GreenAn improved computational bioacoustic monitoring approach for sparse features detection\.InProceedings of Meetings on Acoustics,Vol\.52,pp\. 022002\.Cited by:[§5\.3](https://arxiv.org/html/2609.27409#S5.SS3.p2.1)\.
- McEwenet al\.\(2024\)B\. McEwen, K\. Soltero, S\. Gutschmidt, A\. Bainbridge\-Smith, J\. Atlas, and R\. GreenActive few\-shot learning for rare bioacoustic feature annotation\.Ecological Informatics82,pp\. 102734\.Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p2.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p3.1),[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p1.1),[§5\.3](https://arxiv.org/html/2609.27409#S5.SS3.p2.1)\.
- McEwen and Zhang \(2026\)BaseAL: release v1\.2\.0External Links:[Document](https://dx.doi.org/10.5281/zenodo.21806641),[Link](https://doi.org/10.5281/zenodo.21806641)Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p1.1),[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p2.1)\.
- Miaoet al\.\(2021\)Z\. Miao, Z\. Liu, K\. M\. Gaynor, M\. S\. Palmer, S\. X\. Yu, and W\. M\. GetzIterative human and automated identification of wildlife images\.Nature Machine Intelligence3\(10\),pp\. 885–895\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p1.1)\.
- Mittalet al\.\(2019\)S\. Mittal, M\. Tatarchenko, Ö\. Çiçek, and T\. BroxParting with illusions about deep active learning\.arXiv preprint arXiv:1912\.05361\.Cited by:[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1)\.
- Mohaimenuzzamanet al\.\(2023\)M\. Mohaimenuzzaman, C\. Bergmeir, and B\. MeyerDeep active audio feature learning in resource\-constrained environments\.arXiv preprint arXiv:2308\.13201\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px5.p1.1)\.
- Molleret al\.\(2017\)T\. Moller, I\. Nilssen, and T\. W\. NattkemperActive learning for the classification of species in underwater images from a fixed observatory\.InProceedings of the IEEE International Conference on Computer Vision Workshops,pp\. 2891–2897\.Cited by:[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p1.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1)\.
- Monarch \(2021\)R\. M\. MonarchHuman\-in\-the\-loop machine learning: active learning and annotation for human\-centered ai\.Simon and Schuster\.Cited by:[§2\.1](https://arxiv.org/html/2609.27409#S2.SS1.p7.1)\.
- Mononenet al\.\(2025\)T\. Mononen, B\. Hardwick, S\. Alcobia, A\. Barrett, G\. A\. Blagoev, S\. Boyer, P\. Gonçalves, B\. Gottsberger, E\. Groner, C\. C\. Ho,et al\.An active ensemble classifier for detecting animal sequences from global camera trap data\.Methods in Ecology and Evolution16\(10\),pp\. 2500–2516\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px4.p1.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p1.1)\.
- Nguyen and Nguyen \(2025\)T\. T\. T\. Nguyen and D\. T\. NguyenA model\-agnostic active learning approach for animal detection from camera traps\.arXiv preprint arXiv:2507\.06537\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1)\.
- Norouzzadehet al\.\(2021\)M\. S\. Norouzzadeh, D\. Morris, S\. Beery, N\. Joshi, N\. Jojic, and J\. CluneA deep active learning system for species identification and counting in camera trap images\.Methods in ecology and evolution12\(1\),pp\. 150–161\.Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p1.1),[§1](https://arxiv.org/html/2609.27409#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p1.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p1.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p2.1)\.
- Perezet al\.\(2024a\)G\. Perez, S\. Maji, and D\. SheldonDISCount: counting in large image collections with detector\-based importance sampling\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 22294–22302\.Cited by:[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p1.1),[§5\.2](https://arxiv.org/html/2609.27409#S5.SS2.p2.1)\.
- Perezet al\.\(2024b\)G\. Perez, D\. Sheldon, G\. Van Horn, and S\. MajiHuman\-in\-the\-loop visual re\-id for population size estimation\.InEuropean Conference on Computer Vision,pp\. 185–202\.Cited by:[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p1.1),[§5\.2](https://arxiv.org/html/2609.27409#S5.SS2.p2.1)\.
- Qianet al\.\(2017\)K\. Qian, Z\. Zhang, A\. Baird, and B\. SchullerActive learning for bird sound classification via a kernel\-based extreme learning machine\.The Journal of the Acoustical Society of America142\(4\),pp\. 1796–1804\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p2.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p1.1)\.
- Rauchet al\.\(2025a\)L\. Rauch, R\. Heinrich, I\. Moummad, A\. Joly, B\. Sick, and C\. ScholzCan masked autoencoders also listen to birds?\.arXiv preprint arXiv:2504\.12880\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Rauchet al\.\(2024\)L\. Rauch, D\. Huseljic, M\. Wirth, J\. Decke, B\. Sick, and C\. ScholzTowards deep active learning in avian bioacoustics\.arXiv preprint arXiv:2406\.18621\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p2.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p2.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p1.1)\.
- Rauchet al\.\(2025b\)L\. Rauch, R\. Schwinger, M\. Wirth, R\. Heinrich, D\. Huseljic, M\. Herde, J\. Lange, S\. Kahl, B\. Sick, S\. Tomforde, and C\. ScholzBirdSet: a large\-scale dataset for audio classification in avian bioacoustics\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=dRXxFEY8ZE)Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.p1.1)\.
- Rauchet al\.\(2023\)L\. Rauch, R\. Schwinger, M\. Wirth, B\. Sick, S\. Tomforde, and C\. ScholzActive bird2vec: towards end\-to\-end bird sound monitoring with transformers\.arXiv preprint arXiv:2308\.07121\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1)\.
- Renet al\.\(2021\)P\. Ren, Y\. Xiao, X\. Chang, P\. Huang, Z\. Li, B\. B\. Gupta, X\. Chen, and X\. WangA survey of deep active learning\.ACM computing surveys \(CSUR\)54\(9\),pp\. 1–40\.Cited by:[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p1.1),[§2](https://arxiv.org/html/2609.27409#S2.p1.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p4.1)\.
- Saniet al\.\(2025\)D\. Sani, M\. Khurana, and S\. AnandActive learning for animal re\-identification with ambiguity\-aware sampling\.arXiv preprint arXiv:2511\.06658\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p3.1)\.
- Schäfer\-Zimmermannet al\.\(2024\)J\. C\. Schäfer\-Zimmermann, V\. Demartsev, B\. Averly, K\. Dhanjal\-Adams, M\. Duteil, G\. Gall, M\. Faiß, L\. Johnson\-Ulrich, D\. Stowell, M\. B\. Manser,et al\.Animal2vec and MeerKAT: a self\-supervised transformer for rare\-event raw audio input and a large\-scale reference dataset for bioacoustics\.arXiv preprint arXiv:2406\.01253\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Schwingeret al\.\(2025\)R\. Schwinger, P\. V\. Zadeh, L\. Rauch, M\. Kurz, T\. Hauschild, S\. Lapp, and S\. TomfordeFoundation models for bioacoustics–a comparative review\.arXiv preprint arXiv:2508\.01277\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Sener and Savarese \(2018\)O\. Sener and S\. SavareseActive learning for convolutional neural networks: a core\-set approach\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p2.1),[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p2.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p3.1)\.
- Settles \(2009\)B\. SettlesActive learning literature survey\.Technical reportTechnical Report1648,University of Wisconsin\-Madison Department of Computer Sciences,Madison, WI\.Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.27409#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2609.27409#S2.SS2.p1.1),[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p1.1),[§2](https://arxiv.org/html/2609.27409#S2.p1.1)\.
- Sheardet al\.\(2024\)J\. K\. Sheard, T\. Adriaens, D\. E\. Bowler, A\. Büermann, C\. T\. Callaghan, E\. C\. Camprasse, S\. Chowdhury, T\. Engel, E\. A\. Finch, J\. Von Gönner,et al\.Emerging technologies in citizen science and potential for insect monitoring\.Philosophical Transactions of the Royal Society B379\(1904\),pp\. 20230106\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Stevenset al\.\(2024\)S\. Stevens, J\. Wu, M\. J\. Thompson, E\. G\. Campolongo, C\. H\. Song, D\. E\. Carlyn, L\. Dong, W\. M\. Dahdul, C\. Stewart, T\. Berger\-Wolf,et al\.Bioclip: a vision foundation model for the tree of life\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 19412–19424\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Stowellet al\.\(2026\)D\. Stowell, E\. Vidaña\-Vila, I\. Nolasco, B\. McEwen, L\. Jean\-Labadye, Y\. Benhamadi, G\. Dubus, B\. Hoffman, P\. Linhart, I\. Morandi,et al\.BioDCASE: using data challenges to make community advances in computational bioacoustics\.bioRxiv,pp\. 2026–04\.Cited by:[§4\.5](https://arxiv.org/html/2609.27409#S4.SS5.p2.1)\.
- Stowell \(2022\)D\. StowellComputational bioacoustics with deep learning: a review and roadmap\.PeerJ10,pp\. e13152\.Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p1.1)\.
- Swansonet al\.\(2015\)A\. Swanson, M\. Kosmala, C\. Lintott, R\. Simpson, A\. Smith, and C\. PackerSnapshot serengeti, high\-frequency annotated camera trap images of 40 mammalian species in an african savanna\.Scientific data2\(1\),pp\. 150026\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.p1.1)\.
- Tamkinet al\.\(2022\)A\. Tamkin, D\. Nguyen, S\. Deshpande, J\. Mu, and N\. GoodmanActive learning helps pretrained models learn the intended task\.Advances in Neural Information Processing Systems35,pp\. 28140–28153\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px1.p2.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px5.p1.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p3.1)\.
- Van Klinket al\.\(2022\)R\. Van Klink, T\. August, Y\. Bas, P\. Bodesheim, A\. Bonn, F\. Fossøy, T\. T\. Høye, E\. Jongejans, M\. H\. Menz, A\. Miraldo,et al\.Emerging technologies revolutionise insect ecology and monitoring\.Trends in ecology & evolution37\(10\),pp\. 872–885\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- van Ommen Kloekeet al\.\(2025\)E\. van Ommen Kloeke, W\. D\. Kissling, J\. Evans, C\. Huijbers, J\. Kamminga, and G\. SchoutenARISE: a dutch dataspace connecting nature and people\.InMoral design and green technology,pp\. 233–251\.Cited by:[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px2.p1.1)\.
- van Ostaet al\.\(2024\)J\. van Osta, B\. Dreis, L\. Grogan, J\. Castley, and J\. van OstaSeasonal habitat use of the southern black\-throated finch\.Statement of originality,pp\. 148\.Cited by:[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p3.1)\.
- van Ostaet al\.\(2023\)J\. M\. van Osta, B\. Dreis, E\. Meyer, L\. F\. Grogan, and J\. G\. CastleyAn active learning framework and assessment of inter\-annotator agreement facilitate automated recogniser development for vocalisations of a rare species, the southern black\-throated finch \(poephila cincta cincta\)\.Ecological Informatics77,pp\. 102233\.Cited by:[§1](https://arxiv.org/html/2609.27409#S1.p1.1),[§1](https://arxiv.org/html/2609.27409#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px4.p1.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p2.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p1.1)\.
- Wenet al\.\(2024\)Z\. Wen, O\. Pizarro, and S\. WilliamsFeature alignment: rethinking efficient active learning via proxy in the context of pre\-trained models\.arXiv preprint arXiv:2403\.01101\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p3.1)\.
- Williamset al\.\(2025a\)B\. Williams, M\. C\. R\. P\. M\. Team, A\. Naseem, G\. Nava, A\. Roberts, F\. Nicholson, A\. du Luart, D\. Erasmus, O\. Stoole, A\. Whittick,et al\.Evidence of ecosystem process recovery across a large\-scale coral reef restoration programme using ai accelerated soundscape analysis\.bioRxiv,pp\. 2025–09\.Cited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1)\.
- Williamset al\.\(2025b\)B\. Williams, B\. van Merriënboer, V\. Dumoulin, J\. Hamer, A\. B\. Fleishman, M\. McKown, J\. Munger, A\. N\. Rice, A\. Lillis, C\. White,et al\.Using tropical reef, bird and unrelated sounds for superior transfer learning in marine bioacoustics\.Philosophical Transactions B380\(1928\),pp\. 20240280\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p1.1)\.
- Xieet al\.\(2023\)Y\. Xie, H\. Lu, J\. Yan, X\. Yang, M\. Tomizuka, and W\. ZhanActive finetuning: exploiting annotation budget in the pretraining\-finetuning paradigm\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 23715–23724\.Cited by:[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p3.1)\.
- Zhanget al\.\(2025\)G\. Zhang, K\. Fujinami, and T\. ShimmuraEvaluating active learning and classifiers on laying hens’ motion data of 27\-behavioral classes\.Note:Preprints\.orgCited by:[§2\.3](https://arxiv.org/html/2609.27409#S2.SS3.SSS0.Px3.p3.1)\.
- Zhang and Virtanen \(2025\)S\. Zhang and T\. VirtanenHybrid disagreement\-diversity active learning for bioacoustic sound event detection\.arXiv preprint arXiv:2505\.20956\.Cited by:[§2\.4](https://arxiv.org/html/2609.27409#S2.SS4.p3.1),[§4\.1](https://arxiv.org/html/2609.27409#S4.SS1.SSS0.Px1.p4.1),[§4\.2](https://arxiv.org/html/2609.27409#S4.SS2.p2.1),[§4\.3](https://arxiv.org/html/2609.27409#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2609.27409#S4.SS4.p1.1)\.
## Appendix ASchema
Each paper was processed and extracted against the 24\-field schema below\. Fields use fixed response options where possible to allow aggregation across the corpus\. Papers judged not relevant \(Field 2\) were not processed further; a short note recording the reason for exclusion was kept instead\. Any field that could not be determined confidently from the text was flagged for manual review by the authors rather than inferred\.相似文章
预算协调:少标签联邦主动学习
本文研究低预算环境下的联邦主动学习,揭示了由于异质性反转,同质数据需要更强的协调。它提出了一种利用联邦表示学习的框架,以实现全局协调的主动选择,性能优于现有方法。
针对尺度感知关键材料回收的决策导向主动学习
本技术说明介绍了一种用于优化关键材料回收的决策导向主动学习方法,表明在自动化实验室工作流中,自适应策略能够以比非自适应方法更少的实验实现最大富集。
基于基础模型先验的主动学习:类别不平衡下的高效学习
本文提出了一种新颖的主动学习框架,利用基础模型先验来同时解决类别不平衡和标签噪声问题,在图像和文本领域相比基线方法节省了超过50%的标注成本。
困难案例与不良标签:在有限标签噪声下测试不确定性采样中的误差暴露与误差定位
本研究在有限标签噪声条件下测试主动学习中的不确定性采样,通过比较不同数据集上的误差暴露和误差定位效应,以评估其鲁棒性和性能。
在标签空间演化下的开放世界意图发现中基于不确定性感知的持续学习
本文提出一个统一的不确定性感知概率框架,用于在演化标签空间中进行持续新意图发现,采用自适应β-VAE和多信号决策机制,以实现可控的标签扩展,同时减轻灾难性遗忘。实验展示了高新颖性精度和稳定的适应性,遗忘程度有限。