Instructions to use yiyuliu/voice-command-intent-qwen0.5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use yiyuliu/voice-command-intent-qwen0.5b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="yiyuliu/voice-command-intent-qwen0.5b", filename="voice-command-intent-qwen0.5b-Q3_K_M.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - llama-cpp-python
How to use yiyuliu/voice-command-intent-qwen0.5b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="yiyuliu/voice-command-intent-qwen0.5b", filename="voice-command-intent-qwen0.5b-Q3_K_M.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use yiyuliu/voice-command-intent-qwen0.5b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M # Run inference directly in the terminal: llama cli -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M # Run inference directly in the terminal: llama cli -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M # Run inference directly in the terminal: ./llama-cli -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Use Docker
docker model run hf.co/yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
- LM Studio
- Jan
- vLLM
How to use yiyuliu/voice-command-intent-qwen0.5b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yiyuliu/voice-command-intent-qwen0.5b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yiyuliu/voice-command-intent-qwen0.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
- Ollama
How to use yiyuliu/voice-command-intent-qwen0.5b with Ollama:
ollama run hf.co/yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
- Unsloth Studio
How to use yiyuliu/voice-command-intent-qwen0.5b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for yiyuliu/voice-command-intent-qwen0.5b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for yiyuliu/voice-command-intent-qwen0.5b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for yiyuliu/voice-command-intent-qwen0.5b to start chatting
- Pi
How to use yiyuliu/voice-command-intent-qwen0.5b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use yiyuliu/voice-command-intent-qwen0.5b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use yiyuliu/voice-command-intent-qwen0.5b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use yiyuliu/voice-command-intent-qwen0.5b with Docker Model Runner:
docker model run hf.co/yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
- Lemonade
How to use yiyuliu/voice-command-intent-qwen0.5b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull yiyuliu/voice-command-intent-qwen0.5b:Q3_K_M
Run and chat with the model
lemonade run user.voice-command-intent-qwen0.5b-Q3_K_M
List all available models
lemonade list
TINYLM Voice Command Intent Classifier (Qwen 0.5B, Q3_K_M)
Fine-tuned Qwen2.5-0.5B-Instruct for routing student speech during live assignment defense oral assessments. Runs on CPU via llama-cpp-python โ no GPU required.
Intents
| Intent | Platform action |
|---|---|
answer |
Do nothing โ student is answering (default) |
navigate_back |
Jump to target_question_id (e.g. q2) |
continue |
Resume forward after failed recall |
Model details
| Base model | Qwen/Qwen2.5-0.5B-Instruct |
| Fine-tuning | LoRA SFT (rank 8, ~4.4M trainable params) |
| Quantization | Q3_K_M |
| File size | ~339 MB |
| RAM at inference | ~410โ470 MB (n_ctx=1024) |
| Eval token accuracy | 99.2% |
Files in this repository
| File | Description |
|---|---|
voice-command-intent-qwen0.5b-Q3_K_M.gguf |
Production weights (ship this) |
inference/ |
Python wrapper (VoiceCommandClassifier, prompts, schema, guard) |
Quick start
Download GGUF only
from huggingface_hub import hf_hub_download
gguf = hf_hub_download(
repo_id="yiyuliu/voice-command-intent-qwen0.5b",
filename="voice-command-intent-qwen0.5b-Q3_K_M.gguf",
)
Full integration (GGUF + inference code)
- Clone this repo or copy
voice-command-intent-qwen0.5b-Q3_K_M.gguf+inference/into your project. - Install dependencies:
cd inference
pip install -e .
- Classify an utterance:
from tinylm.inference import VoiceCommandClassifier
from tinylm.prompts import QuestionRecord
clf = VoiceCommandClassifier("voice-command-intent-qwen0.5b-Q3_K_M.gguf")
result = clf.classify(
questions=[
QuestionRecord("q1", "remember", "answered"),
QuestionRecord("q2", "understand", "answered", focus="technical debt"),
QuestionRecord("q3", "apply", "current"),
],
current_question_id="q3",
utterance="could I correct my previous question",
subject="Computer Science",
)
print(result.to_action())
# {"action": "navigate", "target_question_id": "q2"}
Verify install
cd inference
python smoke_test.py
Runtime inputs
Required: student utterance (STT text), Bloom session position (q1โq6, status, optional focus term).
Not required: full lecturer question text or assignment content.
Questions follow Bloom's taxonomy: remember โ understand โ apply โ analyze โ evaluate โ create.
Environment variables (optional)
| Variable | Default |
|---|---|
TINYLM_GGUF_PATH |
auto-resolve GGUF in repo root |
TINYLM_N_CTX |
1024 |
TINYLM_MAX_TOKENS |
64 |
License
- Fine-tuned weights: Apache 2.0 (same as base Qwen2.5)
- Inference code: MIT
Base model license: Qwen/Qwen2.5-0.5B-Instruct
- Downloads last month
- 43
3-bit