BuddyBird Parrot Sound Classifier v2.0.0

English

Model description

This binary audio classifier labels an audio clip as parrot or non-parrot based on whether it contains a parrot sound. It combines a frozen Perch 2.0 ONNX encoder with a standardized L2 logistic-regression probe (C=0.0003) trained with balanced class weights and inverse source-group weights.

Intended use

  • Classify WAV audio as parrot or non-parrot.
  • Unlike v1, v2 targets small and medium parrot species only. Large parrots are out of scope (see Limitations).
  • It assumes an upstream step-1 loudness gate has already dropped very quiet clips (max 100 ms-window RMS below โˆ’39 dBFS). It always returns one clip-level decision, even for a silent clip, but its decisions are unreliable on very quiet clips.

Installation

This model uses Python 3.12 and uv. It was tested on macOS 26.5.1 running on Apple Silicon. Exact dependency versions are recorded in uv.lock.

uv sync --frozen

The first prediction downloads the pinned Perch 2.0 ONNX file (about 390 MB) from justinchuby/Perch-onnx. Both the probe and Perch artifacts are checked against the SHA-256 values in model_manifest.json before they are loaded.

Quick start

from inference import ParrotSoundGate

gate = ParrotSoundGate.from_pretrained(
    "buddybird-ai/parrot-sound-gate",
    revision="v2.0.0",
)
result = gate.predict_file("clip.wav")
print(result)
# {"probability": 0.9312..., "is_parrot": True}

Command-line use:

uv run --frozen python inference.py clip.wav --revision v2.0.0

Input and output contract

The model accepts either a file path or a NumPy waveform with its sample rate. Multi-channel input is averaged to mono and resampled to 32 kHz. Unlike v1, v2 does not rescale the audio's volume to a fixed peak level; the original loudness is kept as is. The waveform is split into 5.0-second windows with a 2.5-second hop, encoded to 1,536 dimensions, and mean-pooled across windows. Short clips are zero-padded.

  • probability: the logistic-regression positive-class score. It is not a calibrated confidence score.
  • is_parrot: True for parrot and False for non-parrot, determined by probability >= 0.7122538089752197.

Training data summary

The model was trained on 18,175 clips, 6,580 of which were positive (parrot) examples and 11,595 negative (non-parrot) examples.

Evaluation

The reported numbers come from development evaluation with repeated nested cross-validation (outer 5-fold, inner 4-fold, averaged over seeds 0, 1, 2).

The operating threshold was selected from grouped out-of-fold scores using bucket false-accept-rate (FAR) caps at 50%, targeting a service-guardian FAR of 1%, an ambient human-voice FAR of 1%, and a noise FAR of 1.5%.

Metric Development mean
Parrot recall 91.3%
Human-voice FAR 1.0%
Noise FAR 1.4%

Limitations

  • The model returns one clip-level decision. It does not localize sound events, identify species, identify individuals, or judge successful word mimicry.
  • v2 targets small and medium parrots; large parrots are out of scope and their recall is low (4/85 per seed).

Credit

This work was supported by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Ministry of Science and ICT (MSIT) (IITP-2026-AIยทSW Maestro)

ํ•œ๊ตญ์–ด

๋ชจ๋ธ ์„ค๋ช…

์ด ๋ชจ๋ธ์€ audio clip์— ์•ต๋ฌด์ƒˆ ์†Œ๋ฆฌ๊ฐ€ ํฌํ•จ๋๋Š”์ง€ ํŒ์ •ํ•ด parrot ๋˜๋Š” non-parrot์œผ๋กœ ๋ถ„๋ฅ˜ํ•˜๋Š” binary audio classifier๋‹ค. ๊ณ ์ •๋œ Perch 2.0 ONNX encoder์™€ ํ‘œ์ค€ํ™”๋œ L2 logistic-regression probe(C=0.0003)๋ฅผ ๊ฒฐํ•ฉํ–ˆ๊ณ , balanced ํด๋ž˜์Šค ๊ฐ€์ค‘์น˜์™€ ์ถœ์ฒ˜ ๊ทธ๋ฃน ์—ญ์ˆ˜ ๊ฐ€์ค‘์น˜๋กœ ํ•™์Šตํ–ˆ๋‹ค.

์‚ฌ์šฉ ๋ชฉ์ 

  • WAV audio๋ฅผ parrot ๋˜๋Š” non-parrot์œผ๋กœ ๋ถ„๋ฅ˜
  • v1๊ณผ ๋‹ฌ๋ฆฌ v2๋Š” ์†Œํ˜•ยท์ค‘ํ˜• ์•ต๋ฌด๋งŒ ๋Œ€์ƒ์œผ๋กœ ํ•œ๋‹ค. ๋Œ€ํ˜• ์•ต๋ฌด๋Š” ๋Œ€์ƒ ๋ฒ”์œ„ ๋ฐ–์ด๋‹ค(ํ•œ๊ณ„ ์ฐธ๊ณ ).
  • ๋„ˆ๋ฌด ์กฐ์šฉํ•œ ํด๋ฆฝ(์ตœ๋Œ€ 100ms ์ฐฝ RMS๊ฐ€ โˆ’39 dBFS ๋ฏธ๋งŒ)์ด ์ด๋ฏธ ๊ฑธ๋Ÿฌ์ง„ ๊ฒƒ์„ ๊ฐ€์ •ํ•œ๋‹ค. ๋ฌด์Œ clip์ด ๋“ค์–ด์™€๋„ ํ•ญ์ƒ clip ๋‹จ์œ„ ํŒ์ • ํ•˜๋‚˜๋ฅผ ๋ฐ˜ํ™˜ํ•˜์ง€๋งŒ, ๋งค์šฐ ์กฐ์šฉํ•œ ํด๋ฆฝ์—์„œ๋Š” ํŒ์ •์ด ๋ถˆ์•ˆ์ •ํ•˜๋‹ค.

์„ค์น˜

Python 3.12์™€ uv๋ฅผ ์‚ฌ์šฉํ•œ๋‹ค. Apple Silicon ๊ธฐ๋ฐ˜ macOS 26.5.1์—์„œ ๊ฒ€์ฆ๋˜์—ˆ๋‹ค. ์ •ํ™•ํ•œ ์˜์กด์„ฑ ๋ฒ„์ „์€ uv.lock์— ๊ธฐ๋ก๋˜์–ด ์žˆ๋‹ค.

uv sync --frozen

์ตœ์ดˆ ์ถ”๋ก  ๋•Œ justinchuby/Perch-onnx์—์„œ ๊ณ ์ •๋œ Perch 2.0 ONNX file ์•ฝ 390 MB๋ฅผ ๋‚ด๋ ค๋ฐ›๋Š”๋‹ค. loader๋Š” probe์™€ Perch artifact๋ฅผ loadํ•˜๊ธฐ ์ „์— model_manifest.json์— ๊ณ ์ •๋œ SHA-256๊ณผ ๋น„๊ตํ•œ๋‹ค.

๋น ๋ฅธ ์‹œ์ž‘

from inference import ParrotSoundGate

gate = ParrotSoundGate.from_pretrained(
    "buddybird-ai/parrot-sound-gate",
    revision="v2.0.0",
)
result = gate.predict_file("clip.wav")
print(result)
# {"probability": 0.9312..., "is_parrot": True}

command line์—์„œ๋Š” ๋‹ค์Œ์ฒ˜๋Ÿผ ์‹คํ–‰ํ•œ๋‹ค.

uv run --frozen python inference.py clip.wav --revision v2.0.0

์ž…๋ ฅ๊ณผ ์ถœ๋ ฅ ๊ณ„์•ฝ

file path ๋˜๋Š” sample rate์™€ NumPy waveform์„ ๋ฐ›๋Š”๋‹ค. multi-channel ์ž…๋ ฅ์€ mono ํ‰๊ท , 32 kHz resample ์ˆœ์œผ๋กœ ์ฒ˜๋ฆฌํ•œ๋‹ค.

v1๊ณผ ๋‹ฌ๋ฆฌ v2๋Š” ์ž…๋ ฅ ์Œ๋Ÿ‰์„ ์ผ์ •ํ•œ ์ตœ๋Œ“๊ฐ’์œผ๋กœ ๋งž์ถ”๋Š” ์ •๊ทœํ™”๋ฅผ ํ•˜์ง€ ์•Š๊ณ  ์›๋ณธ ์Œ๋Ÿ‰์„ ๊ทธ๋Œ€๋กœ ์‚ฌ์šฉํ•œ๋‹ค. ์ดํ›„ 5.0์ดˆ window์™€ 2.5์ดˆ hop ๋ถ„ํ• , 1,536์ฐจ์› embedding, window ํ‰๊ท  pooling ์ˆœ์œผ๋กœ ์ฒ˜๋ฆฌํ•œ๋‹ค. ์งง์€ clip์€ 0์œผ๋กœ paddingํ•œ๋‹ค.

  • probability: logistic-regression์˜ positive-class score์ด๋ฉฐ ๋ณด์ •๋œ confidence๊ฐ€ ์•„๋‹ˆ๋‹ค.
  • is_parrot: probability >= 0.7122538089752197์ด๋ฉด parrot์„ ๋œปํ•˜๋Š” True, ์•„๋‹ˆ๋ฉด non-parrot์„ ๋œปํ•˜๋Š” False๋‹ค.

ํ•™์Šต ๋ฐ์ดํ„ฐ ์š”์•ฝ

๋ชจ๋ธ์€ 18,175๊ฐœ clip์œผ๋กœ ํ•™์Šตํ–ˆ์œผ๋ฉฐ, ์ด ์ค‘ 6,580๊ฐœ๊ฐ€ ์•ต๋ฌด์ƒˆ positive clip, 11,595๊ฐœ๊ฐ€ ๋น„์•ต๋ฌด์ƒˆ negative clip์ด๋‹ค.

ํ‰๊ฐ€

์•„๋ž˜ ์ˆ˜์น˜๋Š” repeated nested ๊ต์ฐจ๊ฒ€์ฆ(outer 5-fold, inner 4-fold, seed 0ยท1ยท2 ํ‰๊ท )์˜ ๊ฐœ๋ฐœ ํ‰๊ฐ€ ๊ฒฐ๊ณผ๋‹ค.

์šด์˜ threshold๋Š” ๊ทธ๋ฃน out-of-fold ์ ์ˆ˜์—์„œ bucket FAR ์ƒํ•œ 50%๋กœ ์„ ํƒํ–ˆ์œผ๋ฉฐ, ์„œ๋น„์Šค ๋ณดํ˜ธ์ž FAR 1%, ambient ์‚ฌ๋žŒ ์Œ์„ฑ FAR 1%, ์žก์Œ FAR 1.5%๋ฅผ ๋ชฉํ‘œ๋กœ ํ–ˆ๋‹ค.

์ง€ํ‘œ ๊ฐœ๋ฐœ ํ‰๊ท 
์•ต๋ฌด์ƒˆ recall 91.3%
์‚ฌ๋žŒ ์Œ์„ฑ FAR 1.0%
์žก์Œ FAR 1.4%

ํ•œ๊ณ„

  • clip ๋‹จ์œ„ ํŒ์ •๋งŒ ์ œ๊ณตํ•œ๋‹ค. sound event ์œ„์น˜, ์ข…, ๊ฐœ์ฒด, ๋‹จ์–ด ๋ชจ์‚ฌ ์„ฑ๊ณต ์—ฌ๋ถ€๋Š” ํŒ์ •ํ•˜์ง€ ์•Š๋Š”๋‹ค.
  • v2๋Š” ์†Œํ˜•ยท์ค‘ํ˜• ์•ต๋ฌด ๋Œ€์ƒ์ด๋ฉฐ ๋Œ€ํ˜• ์•ต๋ฌด๋Š” ๋Œ€์ƒ ๋ฒ”์œ„ ๋ฐ–์ด๊ณ  recall์ด ๋‚ฎ๋‹ค(๊ฐ seed 4/85).

Credit

์ด ์„ฑ๊ณผ๋Š” 2026๋…„๋„ ๊ณผํ•™๊ธฐ์ˆ ์ •๋ณดํ†ต์‹ ๋ถ€์˜ ์žฌ์›์œผ๋กœ ์ •๋ณดํ†ต์‹ ๊ธฐํšํ‰๊ฐ€์›์˜ ์ง€์›์„ ๋ฐ›์•„ ์ˆ˜ํ–‰๋œ ๊ฒฐ๊ณผ๋ฌผ์ž„ (IITP-2026-AIยทSW๋งˆ์—์ŠคํŠธ๋กœ)

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for buddybird-ai/parrot-sound-gate

Base model

cgeorgiaw/Perch
Finetuned
(2)
this model