Instructions to use TLSJ/whisper-wesol-ko with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TLSJ/whisper-wesol-ko with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="TLSJ/whisper-wesol-ko")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("TLSJ/whisper-wesol-ko") model = AutoModelForSpeechSeq2Seq.from_pretrained("TLSJ/whisper-wesol-ko", device_map="auto") - Notebooks
- Google Colab
- Kaggle
whisper-wesol-ko
whisper-wesol-ko๋ ํ๊ตญ์ด ํ์ยท์๋ดยท์ ํยท๊ฐ์ ์ ์ฌ๋ฅผ ๋ชฉํ๋ก
seastar105/whisper-medium-komixv2๋ฅผ ์ถ๊ฐ ํ์ตํ Whisper Medium ํฌ๊ธฐ์
ํ๊ตญ์ด ์์ฑ์ธ์ ๋ชจ๋ธ์
๋๋ค.
๋์ผํ ๋ก์ปฌ ํ๊ฐ ํ์ดํ๋ผ์ธ์์ ์ง์ ๊ธฐ๋ฐ ๋ชจ๋ธ์ 8๊ฐ ํ๊ฐ์ ํ๊ท CER์ 7.03%์์ 6.45%๋ก ๋ฎ์ท์ต๋๋ค. ๋ค๋ง FLEURS Korean์์๋ ๊ธฐ๋ฐ ๋ชจ๋ธ์ด ๋ ์ข์๊ณ , ์ธ๋ถ ๋ชจ๋ธ ๋ฐ ์์ฉ API์ ๊ณต๊ฐ ์ ์์๋ ์คํ ํ๊ฒฝ์ด ๋ค๋ฅด๋ฏ๋ก ์ด ๋ชจ๋ธ์ด ๊ทธ ๋ชจ๋ธ๋ค๋ณด๋ค ์ฐ์ํ๋ค๊ณ ๋จ์ ํ์ง ์์ต๋๋ค.
CER์ ๋ฎ์์๋ก ์ข์ต๋๋ค.
Model details
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| Task | Korean automatic speech recognition |
| Architecture | Whisper encoder-decoder Transformer |
| Parameters | ์ฝ 769M |
| Direct base model | seastar105/whisper-medium-komixv2 |
| Original ancestor | openai/whisper-medium |
| Fine-tuning | ์ ์ฒด ํ๋ผ๋ฏธํฐ ํ์ธํ๋(full fine-tuning), LoRA ์๋ |
| Training compute dtype | BF16 |
| Distributed weights dtype | FP16 |
| Input | 16 kHz mono audio, Whisper ๊ธฐ๋ณธ ์ ๋ ฅ ๊ธธ์ด ์ต๋ 30์ด |
| Output | ํ๊ตญ์ด ์ ์ฌ๋ฌธ |
๋ชจ๋ธ ๊ณ๋ณด๋ ๋ค์๊ณผ ๊ฐ์ต๋๋ค.
openai/whisper-medium โ seastar105/whisper-medium-komixv2 โ
whisper-wesol-ko
์ง์ ๊ธฐ๋ฐ ๋ชจ๋ธ์ OpenAI Whisper Medium์ ์ฌ๋ฌ ํ๊ตญ์ด ๋ฐ์ดํฐ์ ์ผ๋ก ๋จผ์ ํ์ธํ๋ํ ๋ชจ๋ธ์ ๋๋ค. ๋ฐ๋ผ์ ๋ณธ ๋ชจ๋ธ์ ์๋์ ์ถ๊ฐ ํ์ต ๋ฐ์ดํฐ๋ฟ ์๋๋ผ ๊ธฐ๋ฐ ๋ชจ๋ธ๊ณผ ์๋ณธ Whisper์ ํ์ต ๋ฐ์ดํฐ ๋ฐ ํ๊ณ๋ ํจ๊ป ์์ํฉ๋๋ค.
Intended use
๊ถ์ฅ ์ฉ๋:
- ํ๊ตญ์ด ํ์, ์๋ด, ์ฝ์ผํฐ, ์ ํ๋ง, ๊ฐ์ ์์ฑ ์ ์ฌ
- ํ๊ตญ์ด ๋ํ์ฒด ๋ฐ ์ผ๋ฐ ์์ฑ์ ๋ฐฐ์น ์ ์ฌ
- ํ๊ตญ์ด ASR ์ฐ๊ตฌ ๋ฐ ์ถ๊ฐ ํ์ธํ๋์ ์์์
๊ฒ์ฆ๋์ง ์์๊ฑฐ๋ ๊ถ์ฅํ์ง ์๋ ์ฉ๋:
- ์๋ฃยท๋ฒ๋ฅ ยท์ฑ์ฉ ๋ฑ ์ค์ธ์์ด ์ฌ๋์๊ฒ ์ค๋ํ ์ํฅ์ ์ฃผ๋ ์์ฌ๊ฒฐ์
- ํ์ ์๋ณ, ๊ฐ์ ยท์ฑ๊ฒฉยท๋ฏผ๊ฐ ์์ฑ ์ถ๋ก
- ๋์ ์์ด ์์งํ ์์ฑ์ ๊ฐ์ ๋๋ ๋๊ท๋ชจ ์ ์ฌ
- ๋ณ๋ VAD/ํ์ ๋ถ๋ฆฌ ์์ด ๊ธด ๋คํ์ ๋ น์์ ๊ทธ๋๋ก ์ฒ๋ฆฌํ๋ ์ฉ๋
- ํ๊ตญ์ด ์ด์ธ ์ธ์ด์ ํ์ง์ด ์ค์ํ ์๋น์ค
Additional training data
๋ณธ ์ถ๊ฐ ํ์ต ๋จ๊ณ์์๋ ๋ค์ ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ์ต๋๋ค.
| ๋ฐ์ดํฐ | ์ฌ์ฉ ๋ชฉ์ | ์ฌ์ฉ๋ |
|---|---|---|
| Zeroth Korean train | ํ๊ตญ์ด ๋ญ๋ ์ฒด | 51.7์๊ฐ / 22,263๋ฐํ |
| AIHub ์ ์์ง ์ ํ๋ง ์์ฑ์ธ์ ๋ฐ์ดํฐ | ์ ํ ์์ฑ | 99.4์๊ฐ / 46,944๋ฐํ |
| AIHub ์๋ด ์์ฑ ๋ฐ์ดํฐ | ์๋ด ๋ํ | 99.4์๊ฐ / 47,425๋ฐํ |
| AIHub ํ์ ์์ฑ ๋ฐ์ดํฐ | ํ์ ์์ฑ | 99.6์๊ฐ / 62,476๋ฐํ |
| AIHub ํ๊ตญ์ด ๊ฐ์ ์์ฑ ๋ฐ์ดํฐ | ๊ฐ์ ์์ฑ | 99.5์๊ฐ / 50,766๋ฐํ |
| ํฉ๊ณ | ์ถ๊ฐ ํ์ต ์์ฑ | ์ฝ 449.6์๊ฐ / 229,874๋ฐํ |
| AIHub ์ํํ๊ฒฝ์์ ๋ฐ์ดํฐ | ์์ ์ฆ๊ฐ | ๊ณต์ฌ์ฅ ์์ 4,699๊ฐ ํด๋ฆฝ |
์๊ฐ์ ์์์ ์ฒซ์งธ ์๋ฆฌ๋ก ๋ฐ์ฌ๋ฆผํ์ต๋๋ค. ํ์ต ๋ฐ์ดํฐ๋ ํ๊ฐ ๋ฐ ๊ฐ๋ฐ ๋ฐ์ดํฐ์ IDยท์ธ์ ๋จ์๋ก ๋ถ๋ฆฌํ์ต๋๋ค. Zeroth test๋ Zeroth train์ ๋ณ๋ split์ ๋๋ค.
Training procedure
| ์ค์ | ๊ฐ |
|---|---|
| Objective | ์ผ๋ฐ ๋ฐฐ์น: teacher-forced cross entropy |
| Noise-consistency step | ์ค์ ์ ์ ์ฒด iteration์ 15% |
| Noise SNR | 0โ20 dB์์ ๊ฒฐ์ ์ ์ผ๋ก ์ ํ |
| Noise loss | noisy CE + 0.5 ร KL(detached clean distribution โ noisy distribution) |
| Optimizer | AdamW, weight decay 0.01 |
| Maximum learning rate | 2e-5 |
| LR schedule | 300 iteration linear warm-up ํ ReduceLROnPlateau(factor 0.5, patience 2, min LR 1e-6) |
| Micro batch / accumulation | 6 / 3 |
| Effective batch | 18๊ฐ ๋ฐํ/optimizer update |
| Gradient clipping | global norm 1.0 |
| Seed | 42 |
| Precision | BF16, SDPA |
| SpecAugment | ์ฌ์ฉํ์ง ์์ |
| Checkpoint evaluation | 1,000 iteration๋ง๋ค |
| Selected checkpoint | 6,000 iteration(์ฝ 2,000 optimizer updates) |
| Early stop | 11,000 iteration, patience 5 |
์ ์ฒด ๋๋ฉ์ธ์ ๋ฐํ๋ฅผ ํ๋์ pool๋ก ํฉ์ณ ๋น๋ณต์ ์ ํํ์ต๋๋ค. ์ผ๋ฐ iteration์ ์ ๋ต ์ ์ฌ์ ๋ํ CE๋ฅผ ์ฌ์ฉํ์ต๋๋ค. ์์-consistency iteration์์๋ ๊ฐ์ ์์ฑ์ clean/noisy ๋ view๋ฅผ ํจ๊ป ๊ณ์ฐํ๋, noisy view์ ์ ๋ต CE์ clean view๋ฅผ ๊ธฐ์ค์ผ๋ก ํ KL loss๋ฅผ ์ฌ์ฉํ์ต๋๋ค. ์ง์์ฆ๋ฅ ๊ต์ฌ ๋ชจ๋ธ์ ์ฌ์ฉํ์ง ์์์ต๋๋ค.
Evaluation
ํ๊ฐ๋ 2026-07-28์ ์ํํ์ต๋๋ค.
Protocol
- decoding: greedy(
do_sample=False),language=ko,task=transcribe - Transformers ์ถ๋ก dtype: BF16
- ์ ๋ ฅ: 16 kHz, ์ต๋ 30์ด
- ์ ๊ทํ: Whisper
BasicTextNormalizer์ ์ฉ ํ ๋ชจ๋ ๊ณต๋ฐฑ ์ ๊ฑฐ - ์ธํธ ์ ์: ๋ฐํ๋ณ CER์ ์ฐ์ ํ๊ท (macro CER)
- ์ ์ฒด ํ๊ท : ๊ฐ ํ๊ฐ์ CER์ ๋์ผ ๊ฐ์ค ์ฐ์ ํ๊ท
- ์ด ๋ฐํ ์: 18,610๊ฐ
์ ์ฒด ํ๊ท ์ 18,610๊ฐ ๋ฌธ์๋ฅผ ํ๊บผ๋ฒ์ ํฉ์ฐํ corpus CER๊ฐ ์๋๋๋ค. 228๊ฐ์ธ CV15๋ 3,000๊ฐ์ธ AIHub ์ธํธ์ ํ๊ท ์์ ๊ฐ์ ๊ฐ์ค์น๋ฅผ ๋ฐ์ต๋๋ค.
| ํ๊ฐ์ | ๋ฐํ ์ |
|---|---|
| Common Voice 15 Korean test | 228 |
| FLEURS Korean test | 382 |
| AIHub ์ ์์ง ์ ํ๋ง test | 3,000 |
| AIHub ํ์ test | 3,000 |
| AIHub ์๋ด test | 3,000 |
| AIHub ํ๊ตญ์ด ๊ฐ์ test | 3,000 |
| KsponSpeech eval clean | 3,000 |
| KsponSpeech eval other | 3,000 |
Controlled local comparison
์๋ ์ธ ๋ชจ๋ธ์ ๊ฐ์ ์ค๋์ค, ์ ์ฌ, ์ ๊ทํ, decoding ์กฐ๊ฑด์์ ์ง์ ์ธก์ ํ์ต๋๋ค.
| Model | 8-set Avg. | CV15 | FLEURS | ์ ํ | ํ์ | ์๋ด | ๊ฐ์ | Kspon clean | Kspon other |
|---|---|---|---|---|---|---|---|---|---|
| seastar105/whisper-medium-komixv2 (direct base) | 7.03 | 6.75 | 4.49 | 5.82 | 9.45 | 5.54 | 8.42 | 8.08 | 7.66 |
| whisper-wesol-ko | 6.45 | 5.86 | 4.73 | 5.20 | 8.90 | 4.21 | 7.88 | 7.74 | 7.07 |
| whisper-wesol-ko-int8 | 6.46 | 5.87 | 4.71 | 5.25 | 8.83 | 4.26 | 7.91 | 7.63 | 7.24 |
๋์ผ ์คํ์์ ๋ณธ ๋ชจ๋ธ์ ๊ธฐ๋ฐ ๋ชจ๋ธ๋ณด๋ค 8-set ํ๊ท CER์ด 0.58%p ๋ฎ์์ต๋๋ค.
8๊ฐ ์ค 7๊ฐ ์ธํธ์์ ๊ฐ์ ๋์ง๋ง FLEURS๋ 4.49 โ 4.73์ผ๋ก ์
ํ๋์ต๋๋ค.
CT2 int8 ๋ฐฐํฌ๋ณธ์ 8-set ํ๊ท ์ด 6.46%๋ก FP16 ์๋ณธ๊ณผ +0.02%p ์ฐจ์ด์์ต๋๋ค.
๋จ, ์ธํธ๋ณ๋ก๋ ์ต๋ 0.17%p ์ฐจ์ด๊ฐ ์์ผ๋ฏ๋ก ๋ชจ๋ ๋ฐ์ดํฐ์์ ๋ฌด์์ค์ด๋ผ๊ณ
๋ณด์ฅํ ์๋ ์์ต๋๋ค.
Additional evaluations
| Model | KOpenAudioBench (2,835) | Zeroth Korean test (457) |
|---|---|---|
| seastar105/whisper-medium-komixv2 | 3.94 | 6.24 |
| whisper-wesol-ko | 3.53 | 2.78 |
KOpenAudioBench๋ ๋ณธ ์ถ๊ฐ ํ์ต ๋จ๊ณ์์ ์ฌ์ฉํ์ง ์์์ต๋๋ค. Zeroth test๋ ํ์ต์ ์ฌ์ฉํ Zeroth train์ held-out split์ ๋๋ค.
Published reference values
์๋ ๊ฐ์ ๋น๊ต๋ฅผ ์ํด
seastar105/whisper-medium-komixv2 ๋ชจ๋ธ ์นด๋์์
๊ฐ์ ธ์จ ๊ณต๊ฐ ์์น์
๋๋ค. ์ธ๋ถ ๋ชจ๋ธ์ ๊ฐ์ ํ๊ฒฝ์์ ์ฌ์ธก์ ํ์ง ์์์ผ๋ฉฐ, ์ฐ๋ฆฌ
๋ชจ๋ธ ํ๋ง ์ ๋ก์ปฌ ์คํ ๊ฒฐ๊ณผ์
๋๋ค.
| Model | Average | CV15 | FLEURS | ์ ํ | ํ์ | ์๋ด | ๊ฐ์ | Kspon clean | Kspon other |
|---|---|---|---|---|---|---|---|---|---|
| whisper-small-komixv2 (published) | 7.36 | 7.07 | 4.19 | 5.60 | 9.67 | 5.50 | 8.55 | 9.26 | 9.07 |
| whisper-medium-komixv2 (published) | 7.30 | 6.62 | 4.52 | 5.85 | 9.42 | 5.47 | 8.38 | 9.19 | 8.97 |
| whisper-large-v3 (published) | 7.99 | 5.11 | 3.72 | 5.45 | 9.35 | 3.83 | 8.46 | 15.08 | 12.89 |
| whisper-large-v3-turbo (published) | 10.75 | 5.38 | 3.99 | 10.93 | 10.27 | 4.21 | 9.42 | 26.66 | 15.16 |
| whisper-wesol-ko (local) | 6.45 | 5.86 | 4.73 | 5.20 | 8.90 | 4.21 | 7.88 | 7.74 | 7.07 |
๊ณต๊ฐ ์์น์ ๋ก์ปฌ ์์น๋ ๋์ผ ์คํ ๊ฒฐ๊ณผ๊ฐ ์๋๋ฏ๋ก ์ง์ ์ ์ธ ์์๋ก ํด์ํ๋ฉด ์ ๋ฉ๋๋ค.
Commercial/API reference
rtzr/Awesome-Korean-Speech-Recognition์
๊ณต๊ฐํ์ ๊ฒน์น๋ 6๊ฐ ์ด(ํ์ยท์๋ดยท์ ํยท๊ฐ์ยทKspon clean/other)๋ง ๋ค์ ํ๊ท ํ๋ฉด
๋ค์๊ณผ ๊ฐ์ต๋๋ค.
| System | Common 6-set Avg. CER |
|---|---|
| ๋ฆฌํด์ ๋ก | 5.90 |
| ๋ฆฌํด์ ๋ก Whisper | 6.55 |
| whisper-wesol-ko | 6.83 |
| Naver ClovaSpeech | 7.02 |
์์ฉ/API ๊ฐ์ ๋ชจ๋ธ ๋ฒ์ , ํ์ฒ๋ฆฌ, ์คํ ์์ ์ด ๋ฌ๋ผ ํต์ ๋ ์์๋ก ๋ณผ ์ ์์ต๋๋ค.
๋ณธ ๋ชจ๋ธ์ ํด๋น ๋ฒค์น๋งํฌ์ ์ฃผ์ ์์ญ๋ณ ํ์ ์ธํธ๋ฅผ ํ๊ฐํ์ง ์์์ต๋๋ค.
Usage
์๋ MODEL_ID๋ฅผ ์ค์ Hugging Face ์ ์ฅ์ ์ด๋ฆ์ผ๋ก ๋ฐ๊พธ์ญ์์ค.
import torch
from transformers import AutoProcessor, WhisperForConditionalGeneration
MODEL_ID = "tlstjdwns/whisper-wesol-ko"
processor = AutoProcessor.from_pretrained(MODEL_ID)
processor.tokenizer.set_prefix_tokens(language="ko", task="transcribe")
model = WhisperForConditionalGeneration.from_pretrained(
MODEL_ID,
dtype=torch.float16,
low_cpu_mem_usage=True,
).to("cuda")
model.eval()
model.generation_config.language = "ko"
model.generation_config.task = "transcribe"
model.generation_config.forced_decoder_ids = None
# audio: 16 kHz mono float array
inputs = processor.feature_extractor(
audio,
sampling_rate=16_000,
return_tensors="pt",
)
input_features = inputs.input_features.to("cuda", dtype=torch.float16)
with torch.inference_mode():
token_ids = model.generate(
input_features,
max_new_tokens=256,
do_sample=False,
)
text = processor.batch_decode(token_ids, skip_special_tokens=True)[0]
print(text.strip())
30์ด๋ณด๋ค ๊ธด ํ์ผ์ chunking, VAD ๋๋ faster-whisper์ ๊ฐ์ ๋ณ๋ long-form
์ฒ๋ฆฌ ๋ฐฉ์์ ์ฌ์ฉํ์ญ์์ค. ํ ๋ฒ์ ์ฒ์ 30์ด๋ง ๋ฃ๋ ์ฝ๋๋ ์ ์ฒด ํ์ผ ์ ์ฌ๋ฅผ
๋ณด์ฅํ์ง ์์ต๋๋ค.
Limitations
- FLEURS Korean์์๋ ์ง์ ๊ธฐ๋ฐ ๋ชจ๋ธ๋ณด๋ค CER์ด
0.24%p๋์์ต๋๋ค. - ํ์ต ๋๋ฉ์ธ๊ณผ ๊ฐ๊น์ด ํ์ยท์๋ดยท์ ํยท๊ฐ์์ ์ ๋ฆฌํ in-domain ๋ชจ๋ธ์ ๋๋ค. ๋ฏ์ ๋ฐฉ์ก, ๋ฐฉ์ธ, ์๋, ๋ ธ์ธ, ์๋ฃยท๋ฒ๋ฅ ์ ๋ฌธ์ฉ์ด, ์ฝ๋ ์ค์์นญ์ ๊ฐ์ ์ฑ๋ฅ์ ๊ธฐ๋ํ๋ฉด ์ ๋ฉ๋๋ค.
- ๋ณ๋์ ํ์ ๋ถ๋ฆฌ, VAD, ์ค์๊ฐ streaming ๊ธฐ๋ฅ์ ํฌํจํ์ง ์์ต๋๋ค.
- Whisper ๊ณ์ด ํน์ฑ์ ๋ฌด์, ๋งค์ฐ ์งง์ ์ธ์นจ, ์ฌํ ์์ ๋๋ ๋๋ฉ์ธ ๋ฐ ์์ฑ์์ ์ค์ ๋ก ๋งํ์ง ์์ ๋ฌธ์ฅ์ ์์ฑํ๊ฑฐ๋ ๊ฐ์ ๋ฌธ์ฅ์ ๋ฐ๋ณตํ ์ ์์ต๋๋ค.
- ๊ณต์ฌ์ฅ ์์ ํฉ์ฑ์ ์ฌ์ฉํ์ง๋ง, ๋ชจ๋ ์ค์ ๊ฑด์ค ํ์ฅ ํ๊ฒฝ์ ๋ํ ๊ฐ๊ฑด์ฑ์ ๋ณด์ฅํ์ง ์์ต๋๋ค.
- ์ฑ๋ณ, ์ฐ๋ น, ์ง์ญ, ์ต์ ๋ฑ ์ธ๊ตฌํต๊ณ ํ์์ง๋จ๋ณ ํธํฅ ํ๊ฐ๋ ์ํํ์ง ์์์ต๋๋ค.
- ๋ ๋ฆฝ์ ์ธ ์ 3์ ํ๊ฐ๋ ์ํํ์ง ์์์ต๋๋ค.
License and data notice
์ด ์ ์ฅ์์ ๋ผ์ด์ ์ค ํ์๋ other์
๋๋ค.
- ์๋ณธ
openai/whisper-medium์ Apache-2.0์ผ๋ก ๊ณต๊ฐ๋์ต๋๋ค. - ์ง์ ๊ธฐ๋ฐ ๋ชจ๋ธ
seastar105/whisper-medium-komixv2์ ํ์ฌ ๋ชจ๋ธ ์นด๋์๋ ๋ณ๋์ ๋ผ์ด์ ์ค๊ฐ ๋ช ์๋ผ ์์ง ์์ต๋๋ค. - ์ถ๊ฐ ํ์ต์๋ AIHub ๋ฐ์ดํฐ๊ฐ ์ฌ์ฉ๋์ผ๋ฉฐ ๋ฐ์ดํฐ์ ๋ณ ์ด์ฉ์กฐ๊ฑด์ด ์ ์ฉ๋ ์ ์์ต๋๋ค.
๊ณต๊ฐ ๋๋ ์์ ์ด์ฉ ์ ์ ์ง์ ๊ธฐ๋ฐ ๋ชจ๋ธ๊ณผ ๊ฐ AIHub ๋ฐ์ดํฐ์ ์ ์ต์ ์ด์ฉ์กฐ๊ฑด์ ํ์ธํด์ผ ํฉ๋๋ค.
Acknowledgements
- OpenAI Whisper
- seastar105/whisper-medium-komixv2
- rtzr/Awesome-Korean-Speech-Recognition
- KRAFTON/KOpenAudioBench
Citation
์ด ๋ชจ๋ธ ์์ฒด์ ๋ํ ๋ ผ๋ฌธ์ ์์ต๋๋ค. Whisper๋ฅผ ์ธ์ฉํ ๋๋ ๋ค์ ๋ ผ๋ฌธ์ ์ฌ์ฉํ์ญ์์ค.
@article{radford2022whisper,
title={Robust Speech Recognition via Large-Scale Weak Supervision},
author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
journal={arXiv preprint arXiv:2212.04356},
year={2022}
}
- Downloads last month
- -