Instructions to use Vision21Tech/VELA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Vision21Tech/VELA with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Vision21Tech/VELA")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Vision21Tech/VELA") model = AutoModelForMultimodalLM.from_pretrained("Vision21Tech/VELA", device_map="auto") - Notebooks
- Google Colab
- Kaggle
VELA โ ํ๊ตญ์ด ASR (Qwen3-ASR-1.7B ํ์ธํ๋)
Qwen/Qwen3-ASR-1.7B ๋ฅผ AI Hub ๊ทนํ ์์ ์์ฑ ์ธ์ ๋ฐ์ดํฐ(136-1) ์๋ณธ ์ค๋์ค 100์๊ฐ์ผ๋ก
ํ์ธํ๋ํ ์ฒดํฌํฌ์ธํธ์
๋๋ค. ์ ์ยท์๋ ์์ฑ ์ธ์ ์ฑ๋ฅ์ ๋ชฉํ๋ก ํ ์คํ ๊ณ์ด ์ค ํ๋์
๋๋ค.
- ํ์ต ์ฒดํฌํฌ์ธํธ:
checkpoint-7257(3 epoch ์ข ๋ฃ ์์ ) - ๋ฐฐํฌ ํ์: bf16 ๋จ์ผ
model.safetensors(3.8 GB) โ ์๋ณธ fp32 ํ์ต ์ฒดํฌํฌ์ธํธ(์ตํฐ๋ง์ด์ ํฌํจ 21 GB)์์ ์ถ๋ก ์ ํ์ํ ๊ฐ์ค์น๋ง ์ถ์ถํด bf16 ์ผ๋ก ์บ์คํ ํ ๊ฒ์ ๋๋ค. ์ถ๋ก ์ฝ๋๊ฐ ์ด์ฐจํผ bf16 ์ผ๋ก ๋ก๋ํ๋ฏ๋ก ๋์ผ ์กฐ๊ฑด์์ ์๋ณธ๊ณผ ์์ธก์ด ์์ ํ ์ผ์นํจ์ 20๊ฐ ํด๋ฆฝ์ผ๋ก ๋์กฐ ํ์ธํ์ต๋๋ค.
์ฑ๋ฅ
ํ๊ฐ์ ์ 5์ข STT(gpt-4o ยท chirp-3 ยท voxtral ยท mai ยท soniox) ํฉ์๋ก ์ ์ ํ canonical ํ ์คํธ์ ๊ณผ ์ฌ๋์ด ๊ต์ ํ ์ ์ฌ์ ๋๋ค. ์์น๋ ๋ชจ๋ ๋ฎ์์๋ก ์ข์.
| ํ๊ฐ์ | ๊ตฌ์ฑ | WER (%) | CER (%) |
|---|---|---|---|
| ์์ด์ข์ | 576 ํด๋ฆฝ (ํฉ์ ๋ผ๋ฒจ) | 37.19 | 21.33 |
| ์ํ์ ์น์ | 429 ํด๋ฆฝ (ํฉ์ ๋ผ๋ฒจ) | 38.68 | 19.45 |
| kids ํตํฉ | 1,005 ํด๋ฆฝ (์ ๋ ํฉ์ฐ) | 37.61 | 20.77 |
| ์ํ์ ์ฒด | 6,867 ๋ฐํ (์ฌ๋ ๊ต์ ์ ๋ต) | 51.40 | 31.80 |
์ฌ๋ ์ ๋ต ๊ธฐ์ค(์ํ์ ์ฒด)์์๋ ๊ฐ์ ๊ณ์ด ์คํ ์ค ๊ฐ์ฅ ์ข์ ๊ฐ์ด๊ณ , ํฉ์ ๋ผ๋ฒจ kids ๊ธฐ์ค์ผ๋ก๋ 50์๊ฐ ํ์ต๋ณธ(37.61 ๋๋น 36.11)์ด ์ฝ๊ฐ ์์ญ๋๋ค.
์ฃผ์: ์ ์์น๋ ์ด ์ฒดํฌํฌ์ธํธ์ ์ ๋ ์ฑ๋ฅ์ด๋ฉฐ, ๋ฒ ์ด์ค ๋ชจ๋ธ(
Qwen/Qwen3-ASR-1.7B) ๋๋น ๊ฐ์ ์ ์ฃผ์ฅํ๋ ๊ฐ์ด ์๋๋๋ค. ๋์ผ ํ๊ฐ์ ์์ ๋ฒ ์ด์ค ๋๋น ๋น๊ต๋ ๋ณ๋๋ก ๊ฒ์ฆ์ด ํ์ํฉ๋๋ค.
ํ์ต ์ค์
| ํญ๋ชฉ | ๊ฐ |
|---|---|
| ๋ฒ ์ด์ค | Qwen/Qwen3-ASR-1.7B |
| ๋ฐ์ดํฐ | AI Hub 136-1 ๊ทนํ ์์, ์๋ณธ ์ค๋์ค 100.1h โ ์ธ๊ทธ๋จผํธ 81,621๊ฐ |
| ๋ผ๋ฒจ | AI Hub ๋๋ด JSON ์ ๋ต (whisper ์ ์ฌ์ ์ ์ฌ๋ ๋งค์นญ ๋ณด์ , ์๊ณ 0.60) |
| ์ ์ฒ๋ฆฌ | pyannote ํ์๋ถ๋ฆฌ โ ์ธ๊ทธ๋จผํธ ๋ถํ โ 16 kHz mono |
| epochs / lr | 3 / 2e-5 (linear, warmup 0.02) |
| ์ ๋ฐ๋ | fp32 ํ์ต + ์ค๋์ค ์ธ์ฝ๋ ๋๊ฒฐ (bf16 ํ์ต์ grad NaN ๋ฐ์) |
| ์ ํจ ๋ฐฐ์น | 32 (per-device 2 ร grad-acc 8 ร 2 GPU DDP) |
์ฌ์ฉ๋ฒ
pip install -r requirements.txt
# torch ๋ GPU ์ ๋ง๋ ๋น๋๋ก:
# pip install torch torchaudio --index-url https://download.pytorch.org/whl/cu128
๋๋ดํ transcribe.py ๋ก ๋ฐ๋ก ์ ์ฌํ ์ ์์ต๋๋ค.
python transcribe.py ์ค๋์ค.wav
python transcribe.py ์ค๋์คํด๋/ --out ๊ฒฐ๊ณผ.jsonl --device cuda:0
์ง์ ๋ก๋ํ๋ ค๋ฉด:
import torch, soundfile as sf
from transformers import AutoModel, AutoProcessor
from qwen_asr import Qwen3ASRModel
model = AutoModel.from_pretrained("Vision21Tech/VELA", dtype=torch.bfloat16, device_map="cuda:0")
processor = AutoProcessor.from_pretrained("Vision21Tech/VELA", fix_mistral_regex=True)
m = Qwen3ASRModel(backend="transformers", model=model, processor=processor,
max_inference_batch_size=8, max_new_tokens=256)
audio, sr = sf.read("์ค๋์ค.wav", dtype="float32") # 16 kHz mono ๊ถ์ฅ
print(m.transcribe([(audio, 16000)], language="Korean")[0].text)
์๊ตฌ์ฌํญ: Python 3.12, qwen-asr==0.0.6 (transformers 4.57.6 / accelerate 1.12.0 ํ).
VRAM ์ bf16 ๊ธฐ์ค ์ฝ 4~5 GB. Ampere ๋ฏธ๋ง GPU ๋ fp16, CPU ๋ fp32 ๋ก ์๋ ํด๋ฐฑํฉ๋๋ค.
๋ผ์ด์ ์ค ยท ์ด์ฉ ์กฐ๊ฑด
ํ์ต ๋ฐ์ดํฐ๋ AI Hub ์ ๊ณต ๋ฐ์ดํฐ๋ก ๋ณ๋ ์ด์ฉ ์ฝ๊ด์ด ์ ์ฉ๋๋ฉฐ, ๋ฒ ์ด์ค ๋ชจ๋ธ
Qwen/Qwen3-ASR-1.7B ์ ๋ผ์ด์ ์ค๋ ํจ๊ป ์ ์ฉ๋ฉ๋๋ค. ์ฌ๋ฐฐํฌยท์์
์ ์ด์ฉ ์ ์ ์์ธก ์กฐ๊ฑด์
๋ฐ๋์ ํ์ธํ์ญ์์ค. ์ด ์ ์ฅ์์ license: other ํ๊ธฐ๋ ๊ทธ ์ ์ฝ์ ๋ฐ์ํ ๋ณด์์ ํ๊ธฐ์
๋๋ค.
- Downloads last month
- 20
Model tree for Vision21Tech/VELA
Base model
Qwen/Qwen3-ASR-1.7B