BuddyBird Parrot Sound Classifier v2.0.0
English
Model description
This binary audio classifier labels an audio clip as parrot or non-parrot
based on whether it contains a parrot sound. It combines a frozen Perch 2.0 ONNX
encoder with a standardized L2 logistic-regression probe (C=0.0003) trained with
balanced class weights and inverse source-group weights.
Intended use
- Classify WAV audio as
parrotornon-parrot. - Unlike v1, v2 targets small and medium parrot species only. Large parrots are out of scope (see Limitations).
- It assumes an upstream step-1 loudness gate has already dropped very quiet clips (max 100 ms-window RMS below โ39 dBFS). It always returns one clip-level decision, even for a silent clip, but its decisions are unreliable on very quiet clips.
Installation
This model uses Python 3.12 and uv. It was tested on macOS 26.5.1 running on
Apple Silicon. Exact dependency versions are recorded in uv.lock.
uv sync --frozen
The first prediction downloads the pinned Perch 2.0 ONNX file (about 390 MB)
from justinchuby/Perch-onnx. Both the probe and Perch artifacts are checked
against the SHA-256 values in model_manifest.json before they are loaded.
Quick start
from inference import ParrotSoundGate
gate = ParrotSoundGate.from_pretrained(
"buddybird-ai/parrot-sound-gate",
revision="v2.0.0",
)
result = gate.predict_file("clip.wav")
print(result)
# {"probability": 0.9312..., "is_parrot": True}
Command-line use:
uv run --frozen python inference.py clip.wav --revision v2.0.0
Input and output contract
The model accepts either a file path or a NumPy waveform with its sample rate. Multi-channel input is averaged to mono and resampled to 32 kHz. Unlike v1, v2 does not rescale the audio's volume to a fixed peak level; the original loudness is kept as is. The waveform is split into 5.0-second windows with a 2.5-second hop, encoded to 1,536 dimensions, and mean-pooled across windows. Short clips are zero-padded.
probability: the logistic-regression positive-class score. It is not a calibrated confidence score.is_parrot:TrueforparrotandFalsefornon-parrot, determined byprobability >= 0.7122538089752197.
Training data summary
The model was trained on 18,175 clips, 6,580 of which were positive (parrot)
examples and 11,595 negative (non-parrot) examples.
Evaluation
The reported numbers come from development evaluation with repeated nested cross-validation (outer 5-fold, inner 4-fold, averaged over seeds 0, 1, 2).
The operating threshold was selected from grouped out-of-fold scores using bucket false-accept-rate (FAR) caps at 50%, targeting a service-guardian FAR of 1%, an ambient human-voice FAR of 1%, and a noise FAR of 1.5%.
| Metric | Development mean |
|---|---|
| Parrot recall | 91.3% |
| Human-voice FAR | 1.0% |
| Noise FAR | 1.4% |
Limitations
- The model returns one clip-level decision. It does not localize sound events, identify species, identify individuals, or judge successful word mimicry.
- v2 targets small and medium parrots; large parrots are out of scope and their recall is low (4/85 per seed).
Credit
This work was supported by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Ministry of Science and ICT (MSIT) (IITP-2026-AIยทSW Maestro)
ํ๊ตญ์ด
๋ชจ๋ธ ์ค๋ช
์ด ๋ชจ๋ธ์ audio clip์ ์ต๋ฌด์ ์๋ฆฌ๊ฐ ํฌํจ๋๋์ง ํ์ ํด parrot ๋๋ non-parrot์ผ๋ก
๋ถ๋ฅํ๋ binary audio classifier๋ค. ๊ณ ์ ๋ Perch 2.0 ONNX encoder์ ํ์คํ๋ L2
logistic-regression probe(C=0.0003)๋ฅผ ๊ฒฐํฉํ๊ณ , balanced ํด๋์ค ๊ฐ์ค์น์ ์ถ์ฒ ๊ทธ๋ฃน
์ญ์ ๊ฐ์ค์น๋ก ํ์ตํ๋ค.
์ฌ์ฉ ๋ชฉ์
- WAV audio๋ฅผ
parrot๋๋non-parrot์ผ๋ก ๋ถ๋ฅ - v1๊ณผ ๋ฌ๋ฆฌ v2๋ ์ํยท์คํ ์ต๋ฌด๋ง ๋์์ผ๋ก ํ๋ค. ๋ํ ์ต๋ฌด๋ ๋์ ๋ฒ์ ๋ฐ์ด๋ค(ํ๊ณ ์ฐธ๊ณ ).
- ๋๋ฌด ์กฐ์ฉํ ํด๋ฆฝ(์ต๋ 100ms ์ฐฝ RMS๊ฐ โ39 dBFS ๋ฏธ๋ง)์ด ์ด๋ฏธ ๊ฑธ๋ฌ์ง ๊ฒ์ ๊ฐ์ ํ๋ค. ๋ฌด์ clip์ด ๋ค์ด์๋ ํญ์ clip ๋จ์ ํ์ ํ๋๋ฅผ ๋ฐํํ์ง๋ง, ๋งค์ฐ ์กฐ์ฉํ ํด๋ฆฝ์์๋ ํ์ ์ด ๋ถ์์ ํ๋ค.
์ค์น
Python 3.12์ uv๋ฅผ ์ฌ์ฉํ๋ค. Apple Silicon ๊ธฐ๋ฐ macOS 26.5.1์์ ๊ฒ์ฆ๋์๋ค. ์ ํํ
์์กด์ฑ ๋ฒ์ ์ uv.lock์ ๊ธฐ๋ก๋์ด ์๋ค.
uv sync --frozen
์ต์ด ์ถ๋ก ๋ justinchuby/Perch-onnx์์ ๊ณ ์ ๋ Perch 2.0 ONNX file ์ฝ 390 MB๋ฅผ
๋ด๋ ค๋ฐ๋๋ค. loader๋ probe์ Perch artifact๋ฅผ loadํ๊ธฐ ์ ์
model_manifest.json์ ๊ณ ์ ๋ SHA-256๊ณผ ๋น๊ตํ๋ค.
๋น ๋ฅธ ์์
from inference import ParrotSoundGate
gate = ParrotSoundGate.from_pretrained(
"buddybird-ai/parrot-sound-gate",
revision="v2.0.0",
)
result = gate.predict_file("clip.wav")
print(result)
# {"probability": 0.9312..., "is_parrot": True}
command line์์๋ ๋ค์์ฒ๋ผ ์คํํ๋ค.
uv run --frozen python inference.py clip.wav --revision v2.0.0
์ ๋ ฅ๊ณผ ์ถ๋ ฅ ๊ณ์ฝ
file path ๋๋ sample rate์ NumPy waveform์ ๋ฐ๋๋ค. multi-channel ์ ๋ ฅ์ mono ํ๊ท , 32 kHz resample ์์ผ๋ก ์ฒ๋ฆฌํ๋ค.
v1๊ณผ ๋ฌ๋ฆฌ v2๋ ์ ๋ ฅ ์๋์ ์ผ์ ํ ์ต๋๊ฐ์ผ๋ก ๋ง์ถ๋ ์ ๊ทํ๋ฅผ ํ์ง ์๊ณ ์๋ณธ ์๋์ ๊ทธ๋๋ก ์ฌ์ฉํ๋ค. ์ดํ 5.0์ด window์ 2.5์ด hop ๋ถํ , 1,536์ฐจ์ embedding, window ํ๊ท pooling ์์ผ๋ก ์ฒ๋ฆฌํ๋ค. ์งง์ clip์ 0์ผ๋ก paddingํ๋ค.
probability: logistic-regression์ positive-class score์ด๋ฉฐ ๋ณด์ ๋ confidence๊ฐ ์๋๋ค.is_parrot:probability >= 0.7122538089752197์ด๋ฉดparrot์ ๋ปํ๋True, ์๋๋ฉดnon-parrot์ ๋ปํ๋False๋ค.
ํ์ต ๋ฐ์ดํฐ ์์ฝ
๋ชจ๋ธ์ 18,175๊ฐ clip์ผ๋ก ํ์ตํ์ผ๋ฉฐ, ์ด ์ค 6,580๊ฐ๊ฐ ์ต๋ฌด์ positive clip, 11,595๊ฐ๊ฐ ๋น์ต๋ฌด์ negative clip์ด๋ค.
ํ๊ฐ
์๋ ์์น๋ repeated nested ๊ต์ฐจ๊ฒ์ฆ(outer 5-fold, inner 4-fold, seed 0ยท1ยท2 ํ๊ท )์ ๊ฐ๋ฐ ํ๊ฐ ๊ฒฐ๊ณผ๋ค.
์ด์ threshold๋ ๊ทธ๋ฃน out-of-fold ์ ์์์ bucket FAR ์ํ 50%๋ก ์ ํํ์ผ๋ฉฐ, ์๋น์ค ๋ณดํธ์ FAR 1%, ambient ์ฌ๋ ์์ฑ FAR 1%, ์ก์ FAR 1.5%๋ฅผ ๋ชฉํ๋ก ํ๋ค.
| ์งํ | ๊ฐ๋ฐ ํ๊ท |
|---|---|
| ์ต๋ฌด์ recall | 91.3% |
| ์ฌ๋ ์์ฑ FAR | 1.0% |
| ์ก์ FAR | 1.4% |
ํ๊ณ
- clip ๋จ์ ํ์ ๋ง ์ ๊ณตํ๋ค. sound event ์์น, ์ข , ๊ฐ์ฒด, ๋จ์ด ๋ชจ์ฌ ์ฑ๊ณต ์ฌ๋ถ๋ ํ์ ํ์ง ์๋๋ค.
- v2๋ ์ํยท์คํ ์ต๋ฌด ๋์์ด๋ฉฐ ๋ํ ์ต๋ฌด๋ ๋์ ๋ฒ์ ๋ฐ์ด๊ณ recall์ด ๋ฎ๋ค(๊ฐ seed 4/85).
Credit
์ด ์ฑ๊ณผ๋ 2026๋ ๋ ๊ณผํ๊ธฐ์ ์ ๋ณดํต์ ๋ถ์ ์ฌ์์ผ๋ก ์ ๋ณดํต์ ๊ธฐํํ๊ฐ์์ ์ง์์ ๋ฐ์ ์ํ๋ ๊ฒฐ๊ณผ๋ฌผ์ (IITP-2026-AIยทSW๋ง์์คํธ๋ก)