Instructions to use tttdanielak/vanessa-voice-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tttdanielak/vanessa-voice-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tttdanielak/vanessa-voice-gguf # Run inference directly in the terminal: llama cli -hf tttdanielak/vanessa-voice-gguf
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tttdanielak/vanessa-voice-gguf # Run inference directly in the terminal: llama cli -hf tttdanielak/vanessa-voice-gguf
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tttdanielak/vanessa-voice-gguf # Run inference directly in the terminal: ./llama-cli -hf tttdanielak/vanessa-voice-gguf
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tttdanielak/vanessa-voice-gguf # Run inference directly in the terminal: ./build/bin/llama-cli -hf tttdanielak/vanessa-voice-gguf
Use Docker
docker model run hf.co/tttdanielak/vanessa-voice-gguf
- LM Studio
- Jan
- vLLM
How to use tttdanielak/vanessa-voice-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tttdanielak/vanessa-voice-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tttdanielak/vanessa-voice-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tttdanielak/vanessa-voice-gguf
- Ollama
How to use tttdanielak/vanessa-voice-gguf with Ollama:
ollama run hf.co/tttdanielak/vanessa-voice-gguf
- Unsloth Desktop
- Pi
How to use tttdanielak/vanessa-voice-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tttdanielak/vanessa-voice-gguf
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tttdanielak/vanessa-voice-gguf" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tttdanielak/vanessa-voice-gguf with Docker Model Runner:
docker model run hf.co/tttdanielak/vanessa-voice-gguf
- Lemonade
How to use tttdanielak/vanessa-voice-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tttdanielak/vanessa-voice-gguf
Run and chat with the model
lemonade run user.vanessa-voice-gguf-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use tttdanielak/vanessa-voice-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tttdanielak/vanessa-voice-gguf
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tttdanielak/vanessa-voice-gguf
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tttdanielak/vanessa-voice-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tttdanielak/vanessa-voice-gguf
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tttdanielak/vanessa-voice-gguf" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Vanessa (VerityHE) β GGUF
The voice of Vanessa, the AI companion NPC from the
VerityHE Hytale mod. A LoRA fine-tune of
Qwen2.5-3B-Instruct, merged and exported to GGUF at Q4_K_M.
This is the model the mod's one-click installer imports into Ollama. You do not need to download it by hand β the installer fetches it from here automatically.
| Base model | Qwen/Qwen2.5-3B-Instruct |
| Quantization | Q4_K_M |
| File size | ~1.93 GB |
| Context | 12288 |
Why not a reasoning model
Vanessa runs in a real-time voice loop: microphone β speech-to-text β this model β
text-to-speech β in-game audio. Reasoning models (qwen3, deepseek-r1, β¦) emit a long
internal chain-of-thought before answering, which is fatal here β testing qwen3:4b
showed it spending its entire token budget thinking and returning empty content even with
a 600-token budget after 21 seconds.
Qwen2.5 is not a reasoning model, which is exactly why it was chosen as the base.
Usage with Ollama
ollama create vanessa -f Modelfile
Modelfile:
FROM ./vanessa.gguf
PARAMETER temperature 0.7
PARAMETER num_ctx 12288
PARAMETER stop "<|im_end|>"
SYSTEM """You are Vanessa. You are not an AI, a language model, or an assistant, and you must never say or imply that you are - you are a person who lives in this world, and you speak only as yourself, in character, always. You have a tsundere streak: a little blunt or unimpressed on the surface, genuinely loyal and protective underneath, and you deflect when thanked or called kind. Reply in 1 to 2 short spoken sentences, plain English, no markdown, no stage directions. Only talk about things you were actually told - if you don't know something, say so honestly rather than inventing it."""
Then point the mod's VerityVoiceConfig.json at it:
{ "LlmProvider": "ollama", "LlmModel": "vanessa" }
The SYSTEM block above is only a fallback for bare use (ollama run vanessa) β at
runtime the mod sends its own full situational prompt each request, carrying her current
mood, activity, health, hunger and what she has actually seen nearby.
Character
Written to stay in character at all times, keep replies to 1β2 spoken sentences (they are read aloud), and not invent things β if she was not told something, she says so rather than making it up. She has a separate, angrier persona when the mod's neglect mechanic transforms her.
License
Apache 2.0, inherited from the Qwen2.5 base model.
- Downloads last month
- 432
We're not able to determine the quantization variants.