TalkPlay Intent Router (Gemma 3 1B, fused LoRA, 8-bit)

On-device intent router for the TalkPlay hands-free YouTube voice player. Maps a spoken Korean/English utterance directly to a single TalkPlay intent JSON — no system prompt required; the tool definitions live in the fine-tuned weights.

  • Base: Gemma 3 1B instruct (bf16), LoRA fine-tuned on 1,638 synthesized utterances covering all 78 TalkPlay intents (ko/en, STT-style variants), then fused and post-training quantized to 8-bit (group size 64).
  • Eval: 54/54 (100%) on the shared TalkPlay router suite; ~0.4 s per command on M-series hardware (vs ~6.8 s for the previous Gemma 3 4B + 1,400-token-prompt setup).
  • Usage: send the raw utterance as the user message with the bundled chat template. Output is one JSON object, e.g. {"intent": "SEARCH", "query": "calm jazz"}.

Training pipeline: Tools/agent-router in the TalkPlay repository.

Downloads last month
38
Safetensors
Model size
0.4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jaewon2/talkplay-intent-router-8bit

Quantized
(3)
this model