Processing and classifying bird songs using wavelet techniques and supervised learning
Summary
The paper proposes a framework for processing and classifying invasive bird species vocalizations using Bayesian wavelet shrinkage and supervised learning models, with SVM achieving high accuracy in classification.
View Cached Full Text
Cached at: 09/11/26, 08:20 AM
# Processing and classifying bird songs using wavelet techniques and supervised learning
Source: [https://arxiv.org/html/2609.10826](https://arxiv.org/html/2609.10826)
Laura Lucia Dominguez Barrios[https://orcid.org/0000-0001-6235-4806](https://orcid.org/0000-0001-6235-4806)111✉:l175089@dac\.unicamp\.brFidel Aniano Causil Barrios[https://orcid.org/0009-0008-7241-4655](https://orcid.org/0009-0008-7241-4655)222✉:f244960@dac\.unicamp\.brAlex Rodrigo dos Santos Sousa[https://orcid.org/0000-0001-5887-3638](https://orcid.org/0000-0001-5887-3638)333✉:asousa@unicamp\.brMariana Rodrigues Motta[https://orcid.org/0000-0002-2657-3857](https://orcid.org/0000-0002-2657-3857)444✉:marirm@unicamp\.br
1,2,3,4Department of Statistics, State University of Campinas, Campinas, Brazil
## Abstract
This study proposes an integrated framework for the processing and classification of invasive bird species vocalizations within natural soundscapes, characterized by high levels of environmental noise\. We address the challenge of signal degradation by employing a Bayesian wavelet shrinkage methodology based on the Epanechnikov kernel prior, which offers a closed form decision rule and high computational efficiency for processing large bioacoustic datasets\. The methodology was applied to recordings of three species obtained from the iNaturalist platform:Euphonia violacea,Leiothrix lutea, andPasser domesticus\. After signal denoising, we extracted a comprehensive set of features, including Mel\-Frequency Cepstral Coefficients \(MFCCs\) and spectral indices such as entropy and zero\-crossing rate\. Several supervised learning models: Random Forest, Multinomial Logistic Regression and Support Vector Machine \(SVM\) were evaluated across different feature dimensionalities\. Our results demonstrate that the proposed wavelet based preprocessing significantly enhances classification performance, with the SVM model achieving the highest accuracy \(up to 0\.9398\) under a 10\-dimensional MFCC configuration\. This research provides a robust statistical tool for automated ecological monitoring and the management of biological invasions\.
##### Keywords
Bioacoustics; wavelet; bayesian; classification; shrinkage; invasive species; supervised learning\.
## 1Introduction
Animal acoustic signals are shaped by selection to convey information based on their tempo, intensity, and frequency\. However, sound signals degrade as they transmit over space and across physical obstacles \(e\.g\., vegetation or infrastructure\), which affects communication potential\. Therefore, propagation experiments are designed to quantify changes in signal structure in a given habitat by broadcasting and re\-recording animal sounds at increasing distances\([Araya\-Salas et al\.,, 2025](https://arxiv.org/html/2609.10826#bib.bib1)\)\.
The extinction of species presents an irreversible loss to humanity, and preventing biodiversity loss is one of the biggest challenges our society faces\. There are challenges to sound\-based recognition of individual species in any natural soundscape, such as due to a variable distance of the sound source \(the animal\) to the sensor \(microphone\); inter\-species and inter\-individual variation in vocalizations; multiple unique vocalizations \(sonotypes\) per species, including mimicry; or biases due to equipment\([Darras et al\.,, 2020](https://arxiv.org/html/2609.10826#bib.bib5)\)\.
Wavelet\-based methods are widely applied in a range of fields, such as mathematics,signal and image processing, geophysics, and many others\. In statistics, applications of wavelets arise mainly in the areas of non\-parametric regression, density estimation,functional data analysis and stochastic processes\. These methods basically utilize the possibility of representing functions that belong to certain functional spaces as expansions in wavelet basis, similar to other expansions such as polynomials, splines and Fourier\. In particular, wavelet basis expansions are attractive due to their sparsity and localization properties; that is, most wavelet coefficients are equal to zero, while the few significant coefficients are concentrated around important features of the function, such as peaks, discontinuities, and oscillations\. See[Daubechies, \(1992\)](https://arxiv.org/html/2609.10826#bib.bib6)and[Mallat, \(2009\)](https://arxiv.org/html/2609.10826#bib.bib16)for the theoretical background on wavelets, and[Vidakovic, \(1999\)](https://arxiv.org/html/2609.10826#bib.bib30)and[Nason, \(2008\)](https://arxiv.org/html/2609.10826#bib.bib20)for general statistical modeling using wavelets\. For further applications of wavelet\-based methods to birdsong analysis and classification, see[Selin et al\., \(2006\)](https://arxiv.org/html/2609.10826#bib.bib26),[Priyadarshani et al\., \(2016\)](https://arxiv.org/html/2609.10826#bib.bib23),[Hsu et al\., \(2018\)](https://arxiv.org/html/2609.10826#bib.bib12),[Priyadarshani et al\., \(2020\)](https://arxiv.org/html/2609.10826#bib.bib24), and, more recently,[Li et al\., \(2025\)](https://arxiv.org/html/2609.10826#bib.bib14)\.
Bayesian approaches have also been extensively explored because they allow the incorporation of prior information, such as the sparsity property and the support of the coefficients, which is helpful in improving the estimates\. The standard prior to the wavelet coefficients is of spike and slab type\. Several prior distributions for the wavelet coefficients have been proposed in the literature\. Most of them are composed of a mixture of a point mass at zero and a symmetric unimodal distribution, such as the normal\([Chipman et al\.,, 1997](https://arxiv.org/html/2609.10826#bib.bib4)\), double\-exponential\([Vidakovic and Ruggeri,, 2001](https://arxiv.org/html/2609.10826#bib.bib31)\), double\-Weibull\([Reményi and Vidakovic,, 2015](https://arxiv.org/html/2609.10826#bib.bib25)\), beta\([Sousa et al\.,, 2020](https://arxiv.org/html/2609.10826#bib.bib27)\)and logistic\([dos Santos Sousa,, 2022](https://arxiv.org/html/2609.10826#bib.bib10)\)distributions\. Although the Bayesian shrinkage rules available in the literature have been successfully applied to several real data problems, they often do not perform well in data with high noise levels\. Recently,\([Barrios and dos Santos Sousa,, 2025](https://arxiv.org/html/2609.10826#bib.bib2)\)proposed the use of a mixture of a point mass at zero and the Epanechnikov distribution as a prior for the wavelet coefficients, and the associated shrinkage rule showed good performance under high noise levels in simulation studies\.
Our main objective is to develop a model for the classification of invasive birds based on bioacoustic signals obtained from natural soundscapes\. These signals usually contain high levels of environmental noise from the recording sites, such as wind, insect sounds, and vocalizations from other species\. Initially, the signals undergo a preprocessing stage in which silent time intervals are removed, the sampling rate \(Hz\) is standardized, and a wavelet shrinkage rule is applied for noise reduction and recovery of the true bird song\.
After preprocessing, covariates strongly associated with the bioacoustic signals are extracted, such as Mel\-Frequency Cepstral Coefficients \(MFCC\)\. These features are then used to train different supervised classification models available in the literature, allowing the subsequent evaluation and comparison of their performance in the identification of invasive birds\.
[Sun et al\., \(2022\)](https://arxiv.org/html/2609.10826#bib.bib28)obtained results suggest that transfer learning and data augmentation could make the use of CNNs to classify species’ vocalizations feasible even for small soundscape\-based projects with many rare species\.
The remainder of this paper is organized as follows\. Section 2 presents a brief overview of wavelet shrinkage\. Section 3 describes the materials and methods, including the datasets and the supervised learning methods\. Section 4 presents and discusses the results of the statistical analysis\. Finally, Section 5 concludes the paper with final remarks\.
## 2Wavelet\-based estimation methodology
### 2\.1Statistical model and the discrete wavelet transform
Assume a sample ofn=2Jn=2^\{J\}observations\(x1,y1\),…,\(xn,yn\)\(x\_\{1\},y\_\{1\}\),\\ldots,\(x\_\{n\},y\_\{n\}\), withJ∈NJ\\in N, and consider the following nonparametric regression model
yi=f\(xi\)\+ϵi,i=1,…,n,\\displaystyle y\_\{i\}=f\(x\_\{i\}\)\+\\epsilon\_\{i\},\\hskip 20\.00003pti=1,\\ldots,n,\(1\)
wheref∈𝕃2\(ℝ\)=\{f:∫f2\(x\)𝑑x<∞\}f\\in\\mathbb\{L\}^\{2\}\(\\mathbb\{R\}\)=\\\{f:\\int f^\{2\}\(x\)dx<\\infty\\\}is an unknown function andϵi\\epsilon\_\{i\}are independent and identically distributed \(IID\) random variables with a normal distribution with zero mean and unknown common varianceσ2\\sigma^\{2\}, withσ\>0\\sigma\>0\. The goal is to estimate the functionffwithout making any assumptions about its functional structure\. The standard nonparametric procedure is to expandffin basis functions for𝕃2\(ℝ\)\\mathbb\{L\}^\{2\}\(\\mathbb\{R\}\), such as polynomials, splines, and their variants, the Fourier basis, wavelets, and others, and estimate the coefficients of the linear combination[Barrios and dos Santos Sousa, \(2025\)](https://arxiv.org/html/2609.10826#bib.bib2);[dos Santos Sousa, \(2024\)](https://arxiv.org/html/2609.10826#bib.bib11)\.
We consider in this work the expansion offfin a wavelet basis,
f\(x\)=∑j,k∈ℤθj,kψj,k\(x\),\\displaystyle f\(x\)=\\sum\_\{j,k\\in\\mathbb\{Z\}\}\\theta\_\{j,k\}\\psi\_\{j,k\}\(x\),\(2\)
where\{ψj,k\(x\)=2j2ψ\(2jx−k\),j,k∈ℤ\}\\\{\\psi\_\{j,k\}\(x\)=2^\{\\frac\{j\}\{2\}\}\\psi\(2^\{j\}x\-k\),j,k\\in\\mathbb\{Z\}\\\}is an orthonormal wavelet basis for𝕃2\(ℝ\)\\mathbb\{L\}^\{2\}\(\\mathbb\{R\}\)constructed by dilationsjjand translateskkof a functionψ\\psicalled a wavelet or mother wavelet andθj,k∈ℝ\\theta\_\{j,k\}\\in\\mathbb\{R\}are wavelet coefficients that describe features offfat spatial locations2−jk2^\{\-j\}kand scales2j2^\{j\}or resolution levelsjj\. Thus, according to the representation[2](https://arxiv.org/html/2609.10826#S2.E2), the problem of estimating the function f is reduced to the problem of estimating the wavelet coefficientsθj,k\\theta\_\{j,k\}\([dos Santos Sousa,, 2024](https://arxiv.org/html/2609.10826#bib.bib11)\)\.
In vector notation, we have
𝒚=𝒇\+𝒆\\boldsymbol\{y\}=\\boldsymbol\{f\}\+\\boldsymbol\{e\}\(3\)where𝒚=\[y1,…,yn\]⊤,𝒇=\[f\(x1\),…,f\(xn\)\]⊤\\boldsymbol\{y\}=\[y\_\{1\},\\ldots,y\_\{n\}\]^\{\\top\},\\hskip 10\.00002pt\\boldsymbol\{f\}=\[f\(x\_\{1\}\),\\ldots,f\(x\_\{n\}\)\]^\{\\top\}, and𝒆=\[e1,…,en\]⊤\\boldsymbol\{e\}=\[e\_\{1\},\\ldots,e\_\{n\}\]^\{\\top\}\. We apply a discrete wavelet transform \(DWT\) on the original data to take them to the wavelet domain\. Although DWT is usually performed by fast algorithms such as the pyramidal algorithm, it is possible to represent it by a transformation matrix𝑾\{\\boldsymbol\{W\}\}with dimensionn×nn\\times nwhich is applied on both sides of \(Equation[3](https://arxiv.org/html/2609.10826#S2.E3)\)\. Since DWT is linear, we obtain the following model in the wavelet domain
𝒅=𝜽\+ϵ,\\boldsymbol\{d\}=\\boldsymbol\{\\theta\}\+\\boldsymbol\{\\epsilon\},\(4\)where𝒅=𝑾𝒚=\[d1,…,dn\]⊤\\boldsymbol\{d\}=\\boldsymbol\{Wy\}=\[d\_\{1\},\\ldots,d\_\{n\}\]^\{\\top\}is the vector of empirical \(observed\) wavelet coefficients,𝜽=𝑾𝒇=\[θ1,…,θn\]⊤\\boldsymbol\{\\theta\}=\\boldsymbol\{Wf\}=\[\\theta\_\{1\},\\ldots,\\theta\_\{n\}\]^\{\\top\}is the sparse vector of unknown wavelet coefficients of f, andϵ=𝑾𝒆=\[ϵ1,…,ϵn\]⊤\\boldsymbol\{\\epsilon\}=\\boldsymbol\{We\}=\[\\epsilon\_\{1\},\\ldots,\\epsilon\_\{n\}\]^\{\\top\}is the vector of random errors\. Thus, we can see the empirical wavelet coefficients𝒅\\boldsymbol\{d\}as noised versions of the unknown wavelet coefficients𝜽\\boldsymbol\{\\theta\}\. Further, the random errors in the wavelet domainϵ1,…,ϵn\\epsilon\_\{1\},\\ldots,\\epsilon\_\{n\}remain independent zero mean normal with common varianceσ2\\sigma^\{2\}since the DWT is orthogonal\. We will drop the subindices and consider the modeld=θ\+ϵd=\\theta\+\\epsilonfor a single wavelet coefficient along the text for simplicity\.
The estimate𝜽^\\boldsymbol\{\\hat\{\\theta\}\}of the vector of wavelet coefficients𝜽\\boldsymbol\{\\theta\}will be done coefficient by coefficient under a Bayesian framework\. Subsequently, the functionffof \([1](https://arxiv.org/html/2609.10826#S2.E1)\) is estimated at the pointsx1,⋯,xnx\_\{1\},\\cdots,x\_\{n\}by the inverse discrete wavelet transform \(IDWT\),
𝒇^=𝑾′𝜽^\.\\boldsymbol\{\\hat\{f\}\}=\\boldsymbol\{W^\{\\prime\}\}\\boldsymbol\{\\hat\{\\theta\}\}\.
### 2\.2Bayesian wavelet shrinkage rule
Wavelet shrinkage estimation is typically performed by reducing the magnitude of the empirical wavelet coefficients to obtain estimates of the true coefficients\. A variety of shrinkage rules have been proposed in the literature, most of which are based on thresholding procedures\. These methods set empirical coefficients to zero whenever their absolute values fall below a prescribed threshold, while retaining or shrinking larger coefficients\. Seminal contributions to this area include the works of\([Donoho and Johnstone,, 1994](https://arxiv.org/html/2609.10826#bib.bib8)\)and\([Donoho and Johnstone,, 1995](https://arxiv.org/html/2609.10826#bib.bib9)\)\.
Under a Bayesian perspective, it is possible to incorporate prior information about the parameters by a prior distribution\. A common priorπ\(⋅\)\\pi\(\\cdot\)for a single wavelet coefficientθ\\thetais a mixture of a point mass function at zeroδ0\(⋅\)\\delta\_\{0\}\(\\cdot\)and a symmetric around zero and unimodal probability density functiong\(⋅,β\)g\(\\cdot;\\beta\), i\.e
π\(θ,α,β\)=αδ0\(θ\)\+\(1−α\)g\(θ,β\),\\pi\(\\theta;\\alpha,\\beta\)=\\alpha\\delta\_\{0\}\(\\theta\)\+\(1\-\\alpha\)g\(\\theta;\\beta\),\(5\)whereα∈\(0,1\)\\alpha\\in\(0,1\)andβ\\betaare hyperparameters\. In this work, we consider the prior distribution based on the Epanechnikov kernel function proposed by\([Barrios and dos Santos Sousa,, 2025](https://arxiv.org/html/2609.10826#bib.bib2)\)given by \([5](https://arxiv.org/html/2609.10826#S2.E5)\) withg\(⋅\)g\(\\cdot\)given by
g\(θ,β\)=34β3\(β2−θ2\)𝕀\(−β,β\)\(θ\),\\displaystyle g\(\\theta;\\beta\)=\\frac\{3\}\{4\\beta^\{3\}\}\(\\beta^\{2\}\-\\theta^\{2\}\)\\mathbb\{I\}\_\{\(\-\\beta,\\beta\)\}\(\\theta\),\(6\)
whereβ\>0\\beta\>0and𝕀\{A\}\(⋅\)\\mathbb\{I\}\_\{\\\{A\\\}\}\(\\cdot\)is the indicator function on the setAA\. Further, it is assumed an exponential distribution as prior distribution forσ2\\sigma^\{2\},
π\(σ2,λ\)=λe−λσ2𝕀\(0,∞\)\(σ2\),\\pi\(\\sigma^\{2\};\\lambda\)=\\lambda e^\{\-\\lambda\\sigma^\{2\}\}\\mathbb\{I\}\_\{\(0,\\infty\)\}\(\\sigma^\{2\}\),\(7\)λ\>0\\lambda\>0\. Thus, under the models \([4](https://arxiv.org/html/2609.10826#S2.E4)\), \([5](https://arxiv.org/html/2609.10826#S2.E5)\), \([6](https://arxiv.org/html/2609.10826#S2.E6)\) and \([7](https://arxiv.org/html/2609.10826#S2.E7)\) and the squared loss functionL\(θ,δ\)=\(δ−θ\)2L\(\\theta,\\delta\)=\(\\delta\-\\theta\)^\{2\}, the associated shrinkage rule is the posterior mean ofθ\\theta, i\.e,
δ\(d\)=\\displaystyle\\delta\(d\)=𝔼\(θ∣d\)\\displaystyle\\mathbb\{E\}\(\\theta\\mid d\)\(8\)=\\displaystyle=\(1−α\)32λ8β3\[2λβ2\+32λβ\+32λ2\(e−2λ\(β−d\)−e−2λ\(β\+d\)\)\+\(λβ2−3\)2λd−λ2λd3λ2\]αℰ𝒟\(0,12λ\)\+\(1−α\)32λ8β3\[βλ\(e−2λ\(β\+d\)\+e−2λ\(β−d\)\)\+22λ\(β2−d2−1λ\)\],\\displaystyle\\frac\{\(1\-\\alpha\)\\frac\{3\\sqrt\{2\\lambda\}\}\{8\\beta^\{3\}\}\\left\[\\frac\{2\\lambda\\beta^\{2\}\+3\\sqrt\{2\\lambda\}\\beta\+3\}\{2\\lambda^\{2\}\}\\left\(e^\{\-\\sqrt\{2\\lambda\}\(\\beta\-d\)\}\-e^\{\-\\sqrt\{2\\lambda\}\(\\beta\+d\)\}\\right\)\+\\frac\{\(\\lambda\\beta^\{2\}\-3\)\\sqrt\{2\\lambda\}d\-\\lambda\\sqrt\{2\\lambda\}d^\{3\}\}\{\\lambda^\{2\}\}\\right\]\}\{\\alpha\\mathcal\{ED\}\\left\(0,\\frac\{1\}\{\\sqrt\{2\\lambda\}\}\\right\)\+\(1\-\\alpha\)\\frac\{3\\sqrt\{2\\lambda\}\}\{8\\beta^\{3\}\}\\left\[\\frac\{\\beta\}\{\\lambda\}\\left\(e^\{\-\\sqrt\{2\\lambda\}\(\\beta\+d\)\}\+e^\{\-\\sqrt\{2\\lambda\}\(\\beta\-d\)\}\\right\)\+\\frac\{2\}\{\\sqrt\{2\\lambda\}\}\\left\(\\beta^\{2\}\-d^\{2\}\-\\frac\{1\}\{\\lambda\}\\right\)\\right\]\},
whereℰ𝒟\(0,12λ\)\\mathcal\{ED\}\\left\(0,\\frac\{1\}\{\\sqrt\{2\\lambda\}\}\\right\)is the probability density function of the double exponential distribution with mean equals to zero and scale parameter equals to12λ\\frac\{1\}\{\\sqrt\{2\\lambda\}\}\. Figure[1](https://arxiv.org/html/2609.10826#S2.F1)shows the shrinkage rule \([8](https://arxiv.org/html/2609.10826#S2.E8)\) forβ=12\\beta=12,λ=1\.3\\lambda=1\.3and several values ofα\\alpha\. In fact, the hyperparameterα\\alphahas an impact on the severity of the shrinkage rule, which reduces the magnitude of the empirical coefficient\. See\([Vidakovic,, 1999](https://arxiv.org/html/2609.10826#bib.bib30)\)and\([Nason,, 2008](https://arxiv.org/html/2609.10826#bib.bib20)\)for a general overview on Bayesian wavelet shrinkage\.
Figure 1:Wavelet shrinkage rule \([8](https://arxiv.org/html/2609.10826#S2.E8)\) forβ=12\\beta=12,λ=1\.3\\lambda=1\.3and several values ofα\\alpha\.The choice of the Epanechnikov rule in this study was motivated by both its theoretical properties and its empirical performance in function estimation problems under noisy conditions\. This rule is based on a simple prior distribution that has been widely used in the statistical literature, facilitating interpretation and application across different settings\. Furthermore, the adopted formulation leads to an explicit decision rule, avoiding complex iterative procedures and resulting in high computational efficiency a particularly important feature when processing large collections of bioacoustic recordings\. Another relevant aspect is that previous studies have shown that the Epanechnikov rule performs competitively and often outperforms alternative approaches in low signal\-to\-noise ratio scenarios, a condition frequently encountered in field recordings of bird songs\. Across several simulation studies and practical applications, this method has surpassed both classical estimation rules and alternative Bayesian procedures available in the literature, including methods specifically designed for low signal\-to\-noise ratio settings, such as the Gamma Minimax rule based on three\-point prior distributions\. Therefore, its use in the present study is well justified for recovering the underlying bioacoustic signal, combining estimation accuracy, robustness to noise, and computational efficiency\.
### 2\.3Hyperparameters elicitation
The specification of the hyperparameters associated with the Epanechnikov rule plays a crucial role in the estimation procedure, since these parameters determine the amount of shrinkage applied to the wavelet coefficients and directly affect the ability of the method to separate the underlying bioacoustic signal from background noise\. In this study, the adopted specifications were based on previous literature and on evidence of good performance in low signal\-to\-noise ratio settings, a common characteristic of bird\-song recordings\. The hyperparameterα\\alphawas defined as a resolution\-level dependent quantity according to
α=α\(j\)=1−1\(j−J0\+l\)γ,\\alpha=\\alpha\(j\)=1\-\\frac\{1\}\{\(j\-J\_\{0\}\+l\)^\{\\gamma\}\},\(9\)
whereJ0J\_\{0\}is the primary resolution level andJ0≤j≤J−1J\_\{0\}\\leq j\\leq J\-1\. Throughout the study, the shrinkage rule was applied to the entire vector of empirical coefficients, i\.e\.,J0=0J\_\{0\}=0\. Following the recommendation of[Vimalajeewa et al\., \(2023\)](https://arxiv.org/html/2609.10826#bib.bib32), the valuesl=1l=1andγ=2\\gamma=2were adopted\. This specification yields increasing values ofα\(j\)\\alpha\(j\)as the resolution level increases, resulting in stronger shrinkage of coefficients associated with finer scales, where a substantial portion of the noise is typically concentrated\. The hyperparameterβ\\betawas selected adaptively at each resolution level according to
β=β\(j\)=max𝑘\|dj,k\|,\\beta=\\beta\(j\)=\\underset\{k\}\{\\mathrm\{max\}\}\\left\\lvert d\_\{j,k\}\\right\\rvert,\(10\)
where the maximization is performed over all wavelet coefficients belonging to leveljj\. This choice allows the support of the prior distribution, defined on the interval\(−β,β\)\(\-\\beta,\\beta\), to adjust automatically to the observed magnitude of the coefficients at each scale, providing greater flexibility in the estimation process\. Finally, the hyperparameterλ\\lambdawas determined from the variability of the wavelet coefficients at the finest resolution level through
λ=λ\(s\)=1s2\+cτexp\(−1τs\),\\lambda=\\lambda\(s\)=\\frac\{1\}\{s^\{2\}\}\+\\frac\{c\}\{\\tau\}\\exp\\left\(\-\\frac\{1\}\{\\tau\}s\\right\),\(11\)
wheressdenotes the sample standard deviation of the wavelet coefficients at the highest resolution level\. In the absence of prior information, the valuesc=1c=1andτ=2\\tau=2were adopted\. This specification enablesλ\\lambdato adapt to the noise level present in the recording, providing stronger smoothing when the observed variability is high while preserving relevant acoustic structures when the signal\-to\-noise ratio is more favorable\. See[Barrios and dos Santos Sousa, \(2025\)](https://arxiv.org/html/2609.10826#bib.bib2)for more details\.
## 3Materials and methods
### 3\.1Euphonia violacea
The violaceous euphonia \(Euphonia violacea\) lives up to its name in terms of beauty and sound, possessing colorful plumage and a melodious song\. It is a resident bird widely distributed throughout northeastern and eastern South America, from Venezuela and Trinidad south to Paraguay, northeastern Argentina, and southeastern Brazil\. It inhabits humid forests and forest edges, as well as parks and gardens, cocoa plantations, and citrus orchards\. Males have bright violet to bluish\-black upperparts and deep golden\-yellow underpart and forehead, while females and juveniles are duller, mostly olive\-colored on the upperparts and olive\-yellow on the underparts\. Violet\-bellied Euphonia are primarily frugivorous, foraging alone or in small groups \(including mixed flocks\), but they also eat nectar and insects when seasonally available\([Martinez,, 2020](https://arxiv.org/html/2609.10826#bib.bib18)\)\.
### 3\.2Leiothrix lutea
During the last decades, the Asian\-native red\-billed leiothrix \(Leiothrix lutea\) has become established in Europe due to escapes or deliberate introductions\. Despite a potential negative effect on ecosystems identified in other invaded regions, its situation in Europe and their potential effect on native birds are poorly known\.leiothrixmay displace some native species as result of its superior dominance\.[Pereira et al\., \(2022\)](https://arxiv.org/html/2609.10826#bib.bib22)obtained records for 37 regions in 10 countries, and identified established populations in France, Italy, Spain, and Portugal\. Its distribution range in Europe almost doubled in less than 20 years\. A species distribution model showed that species presence probability increased with increasing combined values of human population density, spatial trend of occurrences, precipitation seasonality, precipitation of the driest quarter, and minimum temperature of the coldest month\.
#### 3\.2\.1Risk analysis
Public and academic recognition of the problems associated with biological invasions has grown exponentially over the past decade\. The reasons for this growth are three\-fold\. First, the negative effects of some non\-native species have grown too large to ignore\. Thus, increasing numbers of scientists are studying and managing non\-native species in an effort to minimize the effects of biological invaders on native species and human economies\. Second, the number of species being moved out of their native ranges and into novel locations is itself growing\. Therefore, not only are the problems caused by non\-native species becoming blatantly obvious, but also the overall number of problems appears to be growing\. Third, with so many invasive species, it is very hard to do ecological field research without encountering invaders and eventually including them in investigations even if those investigations are for basic research\. Invaders offer some new insights, and it is very difficult for curious scientists to pass up the opportunity to explore these new avenues\([Lockwood et al\.,, 2006](https://arxiv.org/html/2609.10826#bib.bib15)\)\. With[form](https://drive.google.com/file/d/1wemsLrvr7uYfKNZjJa5QwjL2qEesVjYC/view?usp=sharing)we made the Risk analysis ofLeiothrix lutea\.
Table 1:Risk assessment summary for the invasive bird speciesLeiothrix lutea\. The table presents the final risk score obtained from the assessment protocol, including the number of responses in each evaluation section: \(A\) Biological and ecological characteristics, \(B\) Biogeographical aspects, \(C\) Social and economic aspects, and \(D\) Risk\-enhancing characteristics\. The species was classified as presenting a very high invasion risk, resulting in a recommendation for rejection\.Figure 2:Geographic distribution ofLeiothrix luteaobservations in the world, based on[iNaturalist](https://www.inaturalist.org/photos/62222232)data \(n = 5000 records\)\.The risk assessment conducted forLeiothrix lutearesulted in a final score of 81\.5 \(see Table[1](https://arxiv.org/html/2609.10826#S3.T1)\), classifying the species as presenting a very high invasion risk\. The species met the minimum criteria required for a valid assessment, with responses distributed across all evaluation sections\. High scores in the biological and ecological criteria indicate a strong capacity for establishment, survival, and spread in non\-native environments\. In addition, the species showed relevant biogeographical and socio\-economic risk factors, as well as several characteristics associated with invasion success \(Figure[2](https://arxiv.org/html/2609.10826#S3.F2)\)\. Based on the assessment protocol, the species received a recommendation for rejection, indicating that its introduction or management should be restricted due to its high invasive potential and possible negative impacts on native biodiversity and ecosystems\.
### 3\.3Passer domesticus
The house sparrow \(Passer domesticus\) is a bird of the sparrow family Passeridae, found in most parts of the world\. One of about 25 species in the genus Passer, the house sparrow is native to most of Europe, the Mediterranean Basin, and a large part of Asia\. Its intentional or accidental introductions to many regions, including parts of Australasia, Africa, and the Americas, make it the most widely distributed wild bird\. The house sparrow is host to a huge number of parasites and diseases, and the effect of most is unknown\. The commonly recorded bacterial pathogens of the house sparrow are often those common in humans, and include Salmonella and Escherichia coli\. Salmonella is common in the house sparrow, and a comprehensive study of house sparrow disease found it in 13% of sparrows tested\([Todd,, 2013](https://arxiv.org/html/2609.10826#bib.bib29)\)\.
#### 3\.3\.1Risk analysis
Table[2](https://arxiv.org/html/2609.10826#S3.T2)summarizes the risk assessment results for the invasive bird speciesPasser domesticus\. The species obtained a final score of 91\.5, classifying it as a species withvery highinvasion risk according to the adopted assessment protocol\. This elevated score indicates a strong combination of biological, ecological, biogeographical, and socio\-economic characteristics associated with invasion success and potential environmental impact\. Importantly, the assessment satisfied the minimum criteria required for a valid risk analysis \(“Valid RA”\), demonstrating that the available evidence was sufficient to support a reliable classification\. Based on the overall score and the associated risk category, the final recommendation was “Reject”, indicating that the introduction, maintenance, or further spread of the species should not be encouraged due to its considerable invasive potential and associated ecological risks\.
Table 2:Risk assessment summary for the invasive bird speciesPasser domesticus\. The table presents the final risk score obtained from the assessment protocol, including the number of responses in each evaluation section: \(A\) Biological and ecological characteristics, \(B\) Biogeographical aspects, \(C\) Social and economic aspects, and \(D\) Risk\-enhancing characteristics\.Figure 3:Geographic distribution ofPasser domesticusobservations in the Americas, based on[iNaturalist](https://www.inaturalist.org/photos/62222232)data \(n = 5000 records, years 2000 until 2025\)\.[Chasmai et al\., \(2024\)](https://arxiv.org/html/2609.10826#bib.bib3)the iNaturalist Sounds Dataset \(iNatSounds\) is a collection of230,000230,000audio files capturing sounds from over5,5005,500species, contributed by more than 27,000 recordists worldwide\. The dataset encompasses sounds from birds, mammals, insects, reptiles, and amphibians, with audio and species labels derived from observations submitted to iNaturalist, a global citizen science platform\.
Acoustic data were collected from the iNaturalist platform using the R packagerinat\. Observations were retrieved for three major species groups: birdsEuphonia violacea, Leiothrix lutea, andPasser domesticus, based on the query term “song” \(See Figure[4](https://arxiv.org/html/2609.10826#S3.F4)\)\. Bird recordings were restricted to the year 2025\. Records without valid audio files were excluded\. All datasets were combined into a single database, and the variable iconic\_\\\_taxon\_\\\_name was used to define the classification groups\.
Audio files were downloaded directly from their URLs using thehttrpackage\. Given the heterogeneity of file formats, a two\-stage procedure was adopted to ensure successful decoding\. First, audio files were converted to WAV format usingav; when this step failed, a fallback approach based on MP3 decoding viatuneRwas employed\. All signals were resampled to 16 kHz to ensure consistency across recordings\. Silence segments were removed using functions from theseewavepackage, thereby reducing non\-informative portions of the signals\. Each processed signal was then truncated to a length equal to the largest power of two not exceeding its original size, ensuring compatibility with wavelet\-based transformations\.
Figure 4:Animals with frequency in R by[iNaturalist\-Euphonia violacea](https://www.inaturalist.org/photos/28179600),[iNaturalist\-Leiothrix lutea](https://www.inaturalist.org/photos/62222232)and[iNaturalist\-Passer domesticus](https://www.inaturalist.org/photos/452209454)Interquartile Range \(IQR\)Is a measure of statistical dispersion that describes the spread of the middle 50 % of a dataset, calculated by subtracting the first quartile \(Q1\) from the third quartile \(Q3\)\. It represents the range between the 25th and 75th percentiles and is a robust measure of variability often used to identify outliers\.
Mel\-frequency cepstral coefficients \(MFCC\)Are widely used in bioacoustics to represent bird vocalizations by extracting key spectral features that mimic how sound is perceived\. They are crucial for automated bird species recognition, species classification and monitoring\([Davis and Mermelstein,, 1980](https://arxiv.org/html/2609.10826#bib.bib7)\)\.
Root Mean Square \(RMS\)Is a statistical measure of the magnitude of a varying quantity, calculated as the square root of the mean of the squares of a set of values\.
Spectral centroid \(SC\)In birds, it is an acoustic measure that indicates the ”center of gravity” of the power spectrum of a vocalization, essentially determining where most of the sound energy is concentrated \(whether it is higher or lower pitch\)\.
Spectral Entropy \(Entropy\-SP\)In birdsong analysis is a measure of the complexity, disorder, or ”noisiness” of a sound based on its power spectrum, often calculated using the Shannon entropy formula\. It quantifies how energy is distributed across frequencies in a signal\([Nunes et al\.,, 2004](https://arxiv.org/html/2609.10826#bib.bib21)\)\.
Spectral centroid \(Centroid\)In birdsong analysis represents thecenter of massor weighted average frequency of a vocalization, acting as a proxy for perceived soundbrightnessorwarmth\. It indicates where most of the sound energy is concentrated in the frequency spectrum, with higher centroid values corresponding to brighter, higher\-pitched sounds\([Mehdi et al\.,, 2026](https://arxiv.org/html/2609.10826#bib.bib19)\)\.
Zero Crossing Rate \(ZCR\)In bird sounds measures the frequency at which a bird’s audio signal passes through the zero\-amplitude axis, providing a computational metric to analyze the spectral content of a call\. It is widely used in ornithology and bioacoustics as a time\-domain feature for analyzing the rapid, high\-frequency nature of bird songs\([Marck et al\.,, 2022](https://arxiv.org/html/2609.10826#bib.bib17)\)\.
### 3\.4Statistical analysis
To reduce noise while preserving relevant acoustic structures, a Bayesian Wavelet denoising approach was applied through a custom implementation\. This method operates by Shrinking Wavelet coefficients under a prior structure controlled by hyperparameters, allowing adaptive smoothing of the signal while maintaining important temporal features\. The resulting denoised signals formed the basis for subsequent feature extraction\.
Feature extraction was performed using discrete Wavelet decomposition implemented in the Wavelets\. A Daubechies wavelet filter \(DaubExPhase, filter number 10\) was applied to each signal\. For each decomposition level, energy\-based descriptors were computed, including both the total energy and the mean energy of the detail coefficients\. Additionally, the standard deviation of each denoised signal was included as a global measure of variability\. These features were combined to form a structured dataset representing the multiscale characteristics of each acoustic signal\.
The resulting dataset was partitioned into training and testing subsets using stratified sampling to preserve class proportions, as implemented in the caret package\. Specifically, 80% of the data were allocated to model training and 20% to testing\.
Several classification models were then trained and compared\. A Random Forest classifier was fitted using therandomForestpackage with hyperparameter tuning performed via 10\-fold cross\-validation\. A Multinomial Logistic regression model was also fitted, with predictors standardized prior to estimation and model selection conducted using 5\-fold cross\-validation\. In addition, a Relevance Vector Machine was implemented using thekernlab, providing a sparse Bayesian classification framework with a radial basis kernel\. Finally, a Support Vector Machine with radial kernel was trained using thee1071, with hyperparameters selected via cross\-validation and predictors standardized\.
Model performance was evaluated on the test dataset using confusion matrices, allowing the assessment of overall classification accuracy and class\-specific predictive performance\. All analyses were conducted within theRstatistical computing environment\. All codes are in the repository of[GitHub](https://github.com/lldb14/Bird-Wavelet)\.
\(a\)Data 1 \-Passer domesticusandLeiothrix lutea\(b\)Data 2 \-Euphonia violaceaandLeiothrix lutea
Figure 5:Typical SNR Distribution in Birdsong Recordings\. The signal\-to\-noise ratio \(SNR\) in bioacoustic signalsFigure[5](https://arxiv.org/html/2609.10826#S3.F5)illustrates the temporal distribution of the signal\-to\-noise ratio \(SNR\) across birdsong recordings for the two datasets\. In both datasets, the SNR values exhibit substantial variability, characterized by intermittent peaks interspersed with periods of relatively low signal intensity\. This pattern reflects the heterogeneous acoustic conditions commonly observed in bioacoustic recordings, where vocalizations are influenced by environmental noise, recording distance, and background interference\.
For Data 1 \(Passer domesticusandLeiothrix lutea\), the SNR distribution shows a broader range of fluctuations, including several pronounced peaks exceeding 10 dB\. These results suggest the presence of recordings with highly distinguishable vocal signals, although many segments still exhibit low SNR values, indicating challenging acoustic environments\. In contrast, Data 2 \(Euphonia violaceaandLeiothrix lutea\) presents a comparatively narrower SNR range, with most values concentrated at lower levels and fewer extreme peaks\. Nevertheless, sporadic increases in SNR are also evident, demonstrating occasional segments with clearer vocal activity\.
#### 3\.4\.1Logistic Regression \(LR\)
Logistic Regression \(LR\) was used as a baseline statistical model to provide a standard reference for performance evaluation\. Operating within a linear framework, LR incorporated all selected candidate features\. After hyperparameter optimization by grid search, its performance was compared with the proposed machine learning models to assess the relative improvement in predictive ability\.
#### 3\.4\.2Random Forest \(RF\)
Random Forest \(RF\) is an ensemble model that aggregates predictions from multiple decision trees through majority voting\. This approach introduces robustness via feature and data randomness during the training process\. Using random search, the model was optimized with an entropy criterion for fine\-grained splitting\. The search ranges for max\_\\\_depth and n\_\\\_estimators \(number of trees\) were set to 10–35 and 60–110, respectively, to balance predictive accuracy against computational overhead\.
#### 3\.4\.3Support Vector Machine \(SVM\)
Support Vector Machine \(SVM\) aims to identify a hyperplane that maximizes the margin between classes\. For nonlinearly separable data, kernel functions map samples into a higher\-dimensional space\. In this study, the radial basis function \(RBF\) kernel was adopted\. The regularization parameter C controls the balance between margin maximization and classification error, while gamma determines the influence range of individual samples in the RBF kernel\. To achieve optimal performance, C \(0\.1, 1, 10, 100\) and gamma \(0\.01, 0\.1, 1\) were jointly tuned\. The formula of the RBF kernel function is asK\(x,y\)=exp\{−γ‖‖x−y‖‖2\}K\(x,y\)=exp\\\{\-\\gamma\\left\\\|\\left\\\|x\-y\\right\\\|\\right\\\|^\{2\}\\\}, wherexxandyyare the feature vectors of any two arbitrary input samples\.
#### 3\.4\.4Evaluation metrics
Accuracy, Sensitivity, Specificity, Balanced Accuracy, Kappa, Mcnemar´s Test are used to evaluate the prediction model\. Accuracy is the proportion of the sample that is correctly classified to the total number of samples\. The confusion matrix is a two\-dimensional table that is often used to evaluate the performance of a classification model\. Each element in the confusion matrix represents the number of times a sample is predicted to be a category\. Wherein TP denotes True Positive, FN denotes False Negative, FP denotes False Positive and TN denotes True Negative\.
## 4Results and discussion
The Figure[6](https://arxiv.org/html/2609.10826#S4.F6)shows that the species exhibit distinct acoustic patterns that can be effectively captured by the first principal components\. The separation is particularly pronounced when comparing Euphonia violacea and Leiothrix lutea, while the comparison between Passer domesticus and Leiothrix lutea shows a slight overlap, indicating a greater relative similarity between their acoustic characteristics\. The stability of the observed patterns between m=4 and m=10 supports the robustness of the variable selection procedure and its ability to preserve biologically relevant information for species classification[J\. Sueur et al\., \(2008\)](https://arxiv.org/html/2609.10826#bib.bib13)\.
\(a\)m=4 Data 1 \-Passer domesticusandLeiothrix lutea\(b\)m=4 Data 2 \-Euphonia violaceaandLeiothrix lutea\(c\)m=10 Data 1 \-Passer domesticusandLeiothrix lutea\(d\)m=10 Data 2 \-Euphonia violaceaandLeiothrix lutea
Figure 6:Principal Component Analysis \(PCA\) projection of the acoustic feature space for species discrimination\. Plots \([6\(a\)](https://arxiv.org/html/2609.10826#S4.F6.sf1)\) and \([6\(c\)](https://arxiv.org/html/2609.10826#S4.F6.sf3)\) correspond to Data 1 \(Passer domesticus and Leiothrix lutea\) for feature dimensionalities m=4 and m=10, respectively\. Plots \([6\(b\)](https://arxiv.org/html/2609.10826#S4.F6.sf2)\) and \([6\(d\)](https://arxiv.org/html/2609.10826#S4.F6.sf4)\) correspond to Data 2 \(Euphonia violacea and Leiothrix lutea\) for m=4 and m=10\.In Figures \([6\(a\)](https://arxiv.org/html/2609.10826#S4.F6.sf1)\) and \([6\(c\)](https://arxiv.org/html/2609.10826#S4.F6.sf3)\), corresponding to m=4 and m=10, respectively, a clear separation between the two species is observed, primarily along the first principal component \(PC1\)\.Passer domesticusindividuals are concentrated at negative PC1 values, whileLeiothrix luteaindividuals are clustered at positive values\. Although there is a slight overlap in the central region of the graph, the overall structure of the groups remains well\-defined for both m values\. Increasing m from 4 to 10 does not substantially alter the separation between species, suggesting that the main discriminating information is already captured with a reduced number of features\. However, at m=10, greater dispersion is observed withinLeiothrix lutea, especially in the PC2 direction, which could indicate greater intraspecific variability or the inclusion of additional, less discriminating variables\.
Figures \([6\(b\)](https://arxiv.org/html/2609.10826#S4.F6.sf2)\) and \([6\(d\)](https://arxiv.org/html/2609.10826#S4.F6.sf4)\) show a more pronounced separation between groups than that observed in Data 1\. At both m values, the two species form clearly differentiated clusters with minimal overlap\. Discrimination occurs mainly along PC1, while PC2 captures internal variability within each species\. For m=4 \(Figure[6\(b\)](https://arxiv.org/html/2609.10826#S4.F6.sf2)\), the groups appear compact and relatively homogeneous\. As m increases to 10 \(Figure[6\(d\)](https://arxiv.org/html/2609.10826#S4.F6.sf4)\), the separation between centroids is maintained, although an increase in internal dispersion and the appearance of some extreme values are observed, particularly in the group corresponding toEuphonia violacea\. Despite this, the structure of the groups remains clearly distinguishable\.
The variable importance analysis for the m=4 \(Figure[7\(a\)](https://arxiv.org/html/2609.10826#S4.F7.sf1)\) configuration, the ranking structure changes moderately, although Entropy and ZCR remain among the most influential variables\. The reduced dimensional configuration appears to concentrate predictive information into a smaller subset of cepstral descriptors, particularly variables associated with medians, means, and maxima of lower\-orderMFCCcoefficients, such asMFCC\_med1,MFCC\_max1, andMFCC\_mean4\. Interestingly, quantile\-based acoustic measures such as Q75 gain relative importance in this scenario, suggesting that upper\-tail energy behavior becomes more informative under lower\-dimensional representations\. This may indicate that coarse spectral energy patterns are more robust than fine cepstral variations when the representation complexity is reduced\.
The variable importance analysis for the m=10 \(Figure[7\(b\)](https://arxiv.org/html/2609.10826#S4.F7.sf2)\) configuration reveals that Entropy and Zero Crossing Rate \(ZCR\) are the most influential predictors according to both Mean Decrease Accuracy and Mean Decrease Gini criteria\. Their dominant contribution suggests that spectral complexity and temporal signal irregularity are fundamental for discriminating between the vocalizations ofEuphonia violaceaandLeiothrix lutea\. In particular, Entropy captures the distributional complexity of the acoustic spectrum, whereas ZCR reflects rapid temporal oscillations associated with fine\-grained vocal structures\.
Several MFCC\-derived descriptors also exhibited substantial predictive relevance, especially lower\-order coefficients and higher\-order distributional moments such as skewness and kurtosis\. Variables includingMFCC\_min6,MFCC\_Skewness7,MFCC\_med2, andMFCC\_Kurtosis8indicate that both central tendency and asymmetry of cepstral representations contribute meaningfully to class separation\. The simultaneous relevance of skewness and kurtosis metrics suggests that non\-Gaussian properties of the cepstral distributions play an important role in distinguishing species\-specific acoustic signatures\.
\(a\)m=4 Data 1 \-Passer domesticusandLeiothrix lutea\(b\)m=10 Data 1 \-Passer domesticusandLeiothrix lutea
Figure 7:Importance variables by m for Data 1 \-Passer domesticusandLeiothrix lutea\.In the m=4 configuration \(Figure[8\(a\)](https://arxiv.org/html/2609.10826#S4.F8.sf1)\), the importance ranking becomes even more concentrated around a small subset ofMFCCdescriptors, withMFCC\_min2,MFCC\_mean2,MFCC\_med1, andMFCC\_max2emerging as the most influential predictors\. This pattern suggests that lower\-order cepstral components retain the majority of discriminative information under reduced dimensionality\. Additionally, spectral\-energy measures such as Centroid and RMS become comparatively more relevant in this configuration, indicating that global frequency distribution and signal energy partially compensate for the reduction in cepstral detail\. The inclusion of quantile measures \(Q25 and Q75\) further suggests that dispersion characteristics of the spectral distribution contribute to species differentiation\.
The m=10 model for Data 2 \(Figure[8\(b\)](https://arxiv.org/html/2609.10826#S4.F8.sf2)\) demonstrates a distinct importance structure dominated primarily byMFCC\-based descriptors rather than global acoustic statistics\. Variables such asMFCC\_min2,MFCC\_sd2,MFCC\_med2, andMFCC\_Skewness10exhibit the highest predictive contributions, indicating that fine\-scale cepstral variability is central to distinguishing the vocal emissions ofPasser domesticusandLeiothrix lutea\. The prominence of variance and skewness relatedMFCCmeasures suggests that temporal instability and asymmetry in spectral envelopes are particularly informative for this species pair\. Additionally, the importance of lower\-tail statistics \(MFCC\_min2,MFCC\_min5,MFCC\_min9\) indicates that low\-energy spectral components contribute significantly to classification, potentially reflecting subtle harmonic or timbral differences between species\.
Unlike Data 1 \(Figure[7](https://arxiv.org/html/2609.10826#S4.F7)\), traditional global descriptors such as Entropy and ZCR do not dominate the ranking, implying that species discrimination in this dataset relies more heavily on detailed cepstral structure than on broad spectral complexity measures\. This difference may reflect intrinsic acoustic similarities between the species, requiring higher\-resolution cepstral information to achieve adequate discrimination\.
\(a\)m=4 Data 2 \-Euphonia violaceaandLeiothrix lutea\(b\)m=10 Data 2 \-Euphonia violaceaandLeiothrix lutea
Figure 8:Importance variables by m for Data 2 \-Euphonia violaceaandLeiothrix lutea\.Table[3](https://arxiv.org/html/2609.10826#S4.T3)reveals a clear dominance of MFCC\-derived descriptors in discriminating between Passer domesticus and Leiothrix lutea across both model configurations \(m=4m=4andm=10m=10\)\. In particular, the variables associated with the lower\-order MFCC coefficients, especially the minimum, mean, and median summaries of the first and second coefficients \(e\.g\.,MFCC\_min2,MFCC\_mean1,MFCC\_mean2, andMFCC\_med2\), consistently exhibited the highest values for both Mean Decrease Accuracy \(MDA\) and Mean Decrease Gini \(MDG\)\. This pattern indicates that spectral envelope information contained in the low\-frequency cepstral structure is the primary source of discrimination between the two species\.
For them=4m=4model,MFCC\_min2emerged as the most influential predictor, presenting the largest MDA \(14\.56\) and MDG \(9\.14\) values in the entire analysis\. Similarly,MFCC\_mean1,MFCC\_mean2,MFCC\_med1, andMFCC\_med2also displayed remarkably high importance values, suggesting that central tendency measures of the first cepstral coefficients capture highly species\-specific acoustic signatures\. The prominence of these variables indicates that the distinction between the vocalizations of the two species is strongly associated with stable spectral patterns rather than with higher\-order spectral fluctuations\.
When increasing the dimensionality tom=10m=10, the overall structure of variable importance remained stable, although the importance magnitudes generally decreased\. This reduction is expected because the inclusion of additional MFCC dimensions distributes explanatory power across a larger number of correlated predictors\. Nevertheless, the same subset of variables remained dominant, particularlyMFCC\_min2\(MDA = 9\.54; MDG = 5\.00\),MFCC\_mean1,MFCC\_mean2,MFCC\_med5, andMFCC\_med1\. The persistence of these variables across model specifications demonstrates the robustness of the acoustic information extracted from the lower cepstral components\.
An additional important finding is that several higher\-order MFCC descriptors introduced in them=10m=10configuration exhibited low or even negative MDA values \(e\.g\.,MFCC\_Kurtosis7,MFCC\_Skewness8, andMFCC\_max9\)\. Negative MDA values suggest that these variables do not contribute meaningfully to classification performance and may introduce noise or redundancy into the model\. This result reinforces the notion that increasing the number of cepstral dimensions does not necessarily improve discrimination accuracy and may instead reduce interpretability by incorporating acoustically irrelevant information\.
Beyond MFCC derived descriptors, traditional temporal and spectral measures such as Centroid, RMS, ZCR, Q25, and Q75 showed moderate but consistent importance across both models\. Although these variables contributed to classification, their importance values were substantially lower than those observed for the leading MFCC features, indicating that broad spectral energy distribution and signal dispersion characteristics play a secondary role in species differentiation relative to cepstral structure\.
Table 3:Comparison of variable importance measures by Data 1 \-Passer domesticusandLeiothrix luteaform=4m=4andm=10m=10\. Blank cells indicate that the variable was not selected in the corresponding model\.Table[4](https://arxiv.org/html/2609.10826#S4.T4)demonstrates a markedly different variable importance structure for the classification between Euphonia violacea and Leiothrix lutea when compared with the previous dataset\. In this case, the discriminatory information was less concentrated in a small subset of MFCC descriptors and instead distributed across both classical acoustic indices and cepstral\-based variables\. This pattern suggests a more acoustically complex separation between the two species, where spectral energy distribution and temporal variability contribute jointly to classification performance\.
For them=4m=4configuration, Entropy emerged as the most influential non\-MFCC variable, presenting a notably high Mean Decrease Accuracy \(MDA = 9\.31\), followed by ZCR \(6\.57\),MFCC\_min1\(6\.71\), and Centroid \(5\.81\)\. The strong importance of Entropy indicates that differences in spectral disorder or acoustic complexity play a central role in distinguishing the vocalizations of the two species\. Likewise, the relevance of ZCR and Centroid suggests that temporal signal transitions and spectral balance are highly informative descriptors for this classification problem\. These findings indicate that, unlike the previous dataset, the separation between species is not dominated exclusively by low\-order cepstral structure\.
Among the MFCC derived predictors, lower\-order coefficients again showed greater importance than higher\-order coefficients, although their magnitudes were considerably smaller than those observed in thePasser domesticusversusLeiothrix luteacomparison\. Variables such asMFCC\_mean1,MFCC\_med1,MFCC\_min1, andMFCC\_min4consistently displayed moderate relevance, indicating that cepstral summaries still capture meaningful species\-specific information\. However, the relatively lower MDA and MDG values suggest weaker spectral separability between these species\.
When the number of cepstral coefficients increased tom=10m=10, the general importance pattern remained stable, but the explanatory contribution became even more diffuse across predictors\. Entropy \(MDA = 7\.14\), Centroid \(5\.71\), ZCR \(5\.17\), andMFCC\_min1\(5\.70\) continued to rank among the most influential variables, demonstrating the robustness of these descriptors across model configurations\. Additionally, some higher\-order MFCC variables, such asMFCC\_max6,MFCC\_min6, andMFCC\_Kurtosis10, gained moderate relevance in them=10m=10model, suggesting that finer spectral details may provide complementary discriminatory information when a larger cepstral representation is considered\.
A particularly important result is the high frequency of negative or near\-zero MDA values among several MFCC skewness, kurtosis, and standard deviation descriptors\. Variables such asMFCC\_Skewness2,MFCC\_max3,MFCC\_min5, andMFCC\_sd8negatively affected predictive performance, indicating the presence of redundant or noisy information\. This behavior reinforces the importance of variable selection procedures in high\-dimensional acoustic models, especially when additional cepstral coefficients are incorporated\.
Furthermore, the Mean Decrease Gini values were generally low across both configurations, rarely exceeding 2\.5, which contrasts strongly with the previous dataset where several variables showed MDG values above 5\. This finding suggests that the decision boundaries separatingEuphonia violaceaandLeiothrix luteaare less sharply defined, potentially reflecting greater overlap in their acoustic characteristics\. Consequently, classification in this dataset appears to rely on the combined contribution of multiple moderately informative predictors rather than on a small set of highly dominant acoustic features\.
Table 4:Comparison of variable importance measures by Data 2 \-Euphonia violaceaandLeiothrix luteaform=4m=4andm=10m=10\. Blank cells indicate that the variable was not selected in the corresponding model\.### 4\.1Selection of models
The comparative evaluation of the classification models for Data 1 \(Passer domesticusandLeiothrix lutea\- Table[5](https://arxiv.org/html/2609.10826#S4.T5)\) demonstrates substantial differences in predictive performance across algorithms and dimensional configurations \(m=4 and m=10\)\. The results indicate that increasing the dimensionality from m=4 to m=10 generally improved classification performance for the majority of models, particularly for Multinomial Logistic Regression and Support Vector Machines \(SVM\)\.
Among all evaluated methods, the SVM classifier achieved the best overall performance, yielding the highest Accuracy \(0\.85\), Balanced Accuracy \(0\.84\), and Cohen’s Kappa \(0\.69\) under the m=10 configuration\. These results suggest that the SVM was particularly effective in capturing the nonlinear decision boundaries associated with the acoustic structure of the species vocalizations\. Moreover, the simultaneous attainment of high Sensitivity \(0\.91\) and Specificity \(0\.77\) indicates a well\-balanced classification behavior, avoiding excessive bias toward either class\. The relatively high Kappa coefficient further confirms substantial agreement beyond chance, reinforcing the robustness of the SVM model for birdsong discrimination\. Taken together, these findings in Table[5](https://arxiv.org/html/2609.10826#S4.T5), indicate that the m=10 representation generally provided superior discriminatory capacity, particularly for classifiers capable of exploiting complex multivariate relationships\. The superior performance of SVM and Multinomial Logistic Regression highlights the importance of preserving richer acoustic information in the feature extraction stage, especially when dealing with highly variable bioacoustic signals\.
The comparative performance of the supervised models evaluated on Data 2 \- Table[6](https://arxiv.org/html/2609.10826#S4.T6), demonstrates that the Support Vector Machine \(SVM\) and Multinomial Logistic Regression models are the most robust frameworks for discriminating betweenEuphonia violaceaandLeiothrix lutea\. Under them=10m=10configuration, the SVM achieved the highest overall accuracy of 0\.94 \(Balanced Accuracy = 0\.89\), while the Multinomial Logistic Regression model yielded a highly competitive accuracy of 0\.94 and the top Cohen’s Kappa coefficient of 0\.81\. Moving fromm=4m=4tom=10m=10systematically improved sensitivity across both classifiers, increasing from 0\.72 to 0\.86 for Logistic Regression and from 0\.70 to 0\.80 for SVM while preserving exceptional specificity levels above 0\.95\. This indicates that the addition of higher\-order Mel\-Frequency Cepstral Coefficients \(MFCCs\) captures crucial species\-specific spectral nuances that are otherwise lost in lower\-dimensional representations\.
Conversely, other classifiers struggled with inherent structural limitations in this bioacoustic task\. The Relevance Vector Machine \(RVM\) collapsed due to a severe class\-prediction bias, resulting in a low accuracy of 0\.29 with high sensitivity \(0\.97\) but nearly non\-existent specificity \(0\.10\) atm=10m=10\. This systemic imbalance was statistically validated by McNemar’s test, which rejected marginal homogeneity with absolute significance \(p<0\.001p<0\.001\)\. Additionally, the K\-Nearest Neighbors \(KNN\) model proved highly conservative; while maintaining high specificity \(¿0\.96\), its sensitivity remained low \(0\.457 atm=10m=10\), indicating high susceptibility to overlapping acoustic signatures in distance\-based clustering\. Crucially, the high p\-values of McNemar’s test for both SVM \(p=0\.665p=0\.665\) and Logistic Regression \(p=0\.794p=0\.794\) atm=10m=10confirm the statistical symmetry of their errors, demonstrating that these top\-performing architectures provide unbiased and highly reliable predictions for ecological monitoring\.
Table 5:Comparison of classification results for m=4 and m=10 by Data 1 \-Passer domesticusandLeiothrix luteafor 200 iterations\.Table 6:Comparison of classification results for m=4 and m=10 by Data 2 \-Euphonia violaceaandLeiothrix luteafor 200 iterations
## 5Final considerations
The findings of this study underscore the critical importance of specialized preprocessing techniques in the analysis of bioacoustic signals from natural environments\. The implementation of the bayesian wavelet shrinkage rule with an Epanechnikov prior proved to be a decisive factor in recovering the underlying bird songs from recordings with low signal to noise ratios \(SNR\)\. Unlike traditional thresholding methods, this approach provides an explicit, computationally efficient decision rule that is particularly robust under high noise levels, a common obstacle in field recordings\.
Our comparative analysis o classification models reveals that the choice of both the algorithm and the feature space dimensionality is paramount\. The Support Vector Machine \(SVM\) consistently outperformed other models, suggesting its superior ability to handle the complex, nonlinear decision boundaries inherent in species\-specific acoustic signatures\. Furthermore, the transition from a 4 to 10 dimensional feature space generally improved predictive metrics, highlighting that higher order cepstral information contains vital discriminatory details that simpler models might overlook\.
The variable importance analysis provided significant ecological insights, showing that while global descriptors like spectral entropy and zero crossing rate are influential, the spectral envelope information captured by lower order MFCCs remains the primary source of discrimination between invasive species\. This suggests that species identification relies heavily on stable, low\-frequency cepstral structures\.
The integrated pipeline presented here combining advanced Bayesian denoising with robust supervised learning offers a high performance solution for the automated monitoring of invasive birds \. Future research could explore the scalability of this framework to hyper\-diverse soundscapes and the integration of deep learning architectures to further enhance the detection of rare or mimetic vocalizations in real\-time ecological surveillance\.
### Acknowledgments
Laura Lucia Dominguez Barrios was funded by the São Paulo Research Foundation \(FAPESP\), through grant No\. 23/18444\-4, linked to grant No\. 22/04006\-2 – Center for Molecular Plant Breeding, AP\.PCPE\. Fidel Aniano Causil Barrios received support from the Coordination for the Improvement of Higher Education Personnel – Brazil \(CAPES\) – Funding Code 001\.
## References
- Araya\-Salas et al\., \(2025\)Araya\-Salas, M\., Grabarczyk, E\. E\., Quiroz\-Oliva, M\., García\-Rodríguez, A\., and Rico\-Guevara, A\. \(2025\)\.Quantifying degradation in animal acoustic signals with the r package barulho\.Methods in Ecology and Evolution, 16\(3\):456–467\.
- Barrios and dos Santos Sousa, \(2025\)Barrios, F\. A\. C\. and dos Santos Sousa, A\. R\. \(2025\)\.Bayesian wavelet shrinkage for low snr data based on the epanechnikov kernel\.
- Chasmai et al\., \(2024\)Chasmai, M\., Shepard, A\., Maji, S\., and Van Horn, G\. \(2024\)\.The inaturalist sounds dataset\.Advances in Neural Information Processing Systems\.
- Chipman et al\., \(1997\)Chipman, H\., Kolaczyc, E\., and McCulloch, R\. \(1997\)\.Adaptive bayesian wavelet shrinkage\.J\. Am\. Statist\. Ass\., 92:1413–1421\.
- Darras et al\., \(2020\)Darras, K\. F\., Deppe, F\., Fabian, Y\., Kartono, A\. P\., Angulo, A\., Kolbrek, B\., Mulyani, Y\. A\., and Prawiradilaga, D\. M\. \(2020\)\.High microphone signal\-to\-noise ratio enhances acoustic sampling of wildlife\.PeerJ, 8:e9955\.
- Daubechies, \(1992\)Daubechies, I\. \(1992\)\.Ten Lectures on Wavelets\.Society for Industrial and Applied Mathematics \(SIAM\), Philadelphia, PA\.
- Davis and Mermelstein, \(1980\)Davis, S\. and Mermelstein, P\. \(1980\)\.Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences\.IEEE transactions on acoustics, speech, and signal processing, 28\(4\):357–366\.
- Donoho and Johnstone, \(1994\)Donoho, D\. and Johnstone, I\. \(1994\)\.Ideal spatial adaptation by wavelet shrinkage\.Biometrika, 81\(1\):425–455\.
- Donoho and Johnstone, \(1995\)Donoho, D\. L\. and Johnstone, I\. M\. \(1995\)\.Adapting to unknown smoothness via wavelet shrinkage\.Journal of the American Statistical Association, 90\(432\):1200–1224\.
- dos Santos Sousa, \(2022\)dos Santos Sousa, A\. R\. \(2022\)\.Bayesian wavelet shrinkage with logistic prior\.Communications in Statistics\-Simulation and Computation, 51\(8\):4700–4714\.
- dos Santos Sousa, \(2024\)dos Santos Sousa, A\. R\. \(2024\)\.A bayesian wavelet shrinkage rule under linex loss function\.Research in Statistics, 2\(1\):2362926\.
- Hsu et al\., \(2018\)Hsu, S\.\-B\., Lee, C\.\-H\., Chang, P\.\-C\., Han, C\.\-C\., and Fan, K\.\-C\. \(2018\)\.Local wavelet acoustic pattern: a novel time–frequency descriptor for birdsong recognition\.IEEE Transactions on Multimedia, 20\(12\):3187–3199\.
- J\. Sueur et al\., \(2008\)J\. Sueur, T\. Aubin, and C\. Simonis \(2008\)\.Seewave: a free modular tool for sound analysis and synthesis\.Bioacoustics, 18:213–226\.
- Li et al\., \(2025\)Li, W\., Lv, D\., Yu, Y\., Zhang, Y\., Gu, L\., Wang, Z\., and Zhu, Z\. \(2025\)\.Multi\-scale deep feature fusion with machine learning classifier for birdsong classification\.Applied Sciences, 15\(4\):1885\.
- Lockwood et al\., \(2006\)Lockwood, J\., Hoopes, M\., and Marchetti, M\. \(2006\)\.Invasion ecology\.
- Mallat, \(2009\)Mallat, S\. \(2009\)\.A Wavelet Tour of Signal Processing, Third Edition: The Sparse Way\.Academic Press, Inc\., USA, 3rd edition\.
- Marck et al\., \(2022\)Marck, A\., Vortman, Y\., Kolodny, O\., and Lavner, Y\. \(2022\)\.Identification, analysis and characterization of base units of bird vocal communication: the white spectacled bulbul \(pycnonotus xanthopygos\) as a case study\.Frontiers in Behavioral Neuroscience, 15:812939\.
- Martinez, \(2020\)Martinez, D\. \(2020\)\.Violaceous euphonia \(Euphonia violacea\), version 1\.0\.In Schulenberg, T\. S\., editor,Birds of the World\. Cornell Lab of Ornithology, Ithaca, NY, USA\.
- Mehdi et al\., \(2026\)Mehdi, N\. A\., Adeel, M\., and Larik, A\. A\. \(2026\)\.Soundplot: An open\-source framework for birdsong acoustic analysis and neural synthesis with interactive 3d visualization\.arXiv preprint arXiv:2601\.12752\.
- Nason, \(2008\)Nason, G\. P\. \(2008\)\.Wavelet methods in statistics with R\.Springer, New York\.
- Nunes et al\., \(2004\)Nunes, R\. R\., Almeida, M\. P\. d\., and Sleigh, J\. W\. \(2004\)\.Spectral entropy: a new method for anesthetic adequacy\.Revista brasileira de anestesiologia, 54:404–422\.
- Pereira et al\., \(2022\)Pereira, P\., Godinho, C\., Roque, I\., Rabaça, J\., Barbosa, M\., Salgueiro, P\., Silva, R\., and Lourenço, R\. \(2022\)\.Situation of red\-billed leiothrix \(leiothrix lutea\) in europe and interactions with native species\.
- Priyadarshani et al\., \(2016\)Priyadarshani, N\., Marsland, S\., Castro, I\., and Punchihewa, A\. \(2016\)\.Birdsong denoising using wavelets\.PloS one, 11\(1\):e0146790\.
- Priyadarshani et al\., \(2020\)Priyadarshani, N\., Marsland, S\., Juodakis, J\., Castro, I\., and Listanti, V\. \(2020\)\.Wavelet filters for automated recognition of birdsong in long\-time field recordings\.Methods in Ecology and Evolution, 11\(3\):403–417\.
- Reményi and Vidakovic, \(2015\)Reményi, N\. and Vidakovic, B\. \(2015\)\.Wavelet shrinkage with double weibull prior\.Communications in Statistics \- Simulation and Computation, 44:88–104\.
- Selin et al\., \(2006\)Selin, A\., Turunen, J\., and Tanttu, J\. T\. \(2006\)\.Wavelets in recognition of bird sounds\.EURASIP Journal on Advances in Signal Processing, 2007\(1\):051806\.
- Sousa et al\., \(2020\)Sousa, A\., Garcia, N\., and Vidakovic, B\. \(2020\)\.Bayesian wavelet shrinkage with beta prior\.Computational Statistics, 36\(2\):1341–1363\.
- Sun et al\., \(2022\)Sun, Y\., Maeda, T\. M\., Solís\-Lemus, C\., Pimentel\-Alarcón, D\., and Buřivalová, Z\. \(2022\)\.Classification of animal sounds in a hyperdiverse rainforest using convolutional neural networks with data augmentation\.Ecological Indicators, 145:109621\.
- Todd, \(2013\)Todd, K\. \(2013\)\.Sparrow\.Reaktion Books\.
- Vidakovic, \(1999\)Vidakovic, B\. \(1999\)\.Statistical modeling by wavelets\.Wiley, New York\.
- Vidakovic and Ruggeri, \(2001\)Vidakovic, B\. and Ruggeri, F\. \(2001\)\.Bams method: Theory and simulations\.Sankhyā: The Indian Journal of Statistics, Series B, 63:234–249\.
- Vimalajeewa et al\., \(2023\)Vimalajeewa, D\., DasGupta, A\., Ruggeri, F\., and Vidakovic, B\. \(2023\)\.Gamma\-minimax wavelet shrinkage for signals with low snr\.The New England Journal of Statistics in Data Science, 1:159–171\.Similar Articles
Seeing Birdsong
Seeing Birdsong is a framework that transforms avian vocalizations into geometric forms and 3D visualizations, bridging art and acoustic science. It uses spectral descriptors to create data-rich structures for research, education, and artistic performance.
Classification and detection of multiple UAVs using rational Gaussian wavelet neural networks
This paper proposes a cost-effective UAV detection and classification system using sound signals processed by rational Gaussian wavelet neural networks, achieving interpretable and robust performance for single and multiple UAVs including swarms, outperforming traditional methods.
Detecting Alarming Student Verbal Responses using Text and Audio Classifier
This paper presents a hybrid framework for detecting alarming or distressed student verbal responses by combining a text classifier (content-based) and an audio classifier (prosodic features), aimed at expediting human review in Automated Verbal Response Scoring systems. The approach addresses a safety gap in automated scoring pipelines where at-risk student responses may otherwise go unnoticed.
Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications
This paper proposes an alternative signal-centric method for processing sonar data in remote sensing applications, using CSV format and acoustic processing to reduce processing time by 91.18% and improve machine learning-driven object detection.
I turned my security cameras into an automatic bird identification system
This article details how to integrate security cameras with the BirdNet-Go tool to create a real-time bird identification system using local AI models for audio detection and community data sharing.