Instructions to use jsbeaudry/makandal-multiple-v2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jsbeaudry/makandal-multiple-v2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use jsbeaudry/makandal-multiple-v2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jsbeaudry/makandal-multiple-v2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jsbeaudry/makandal-multiple-v2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
- Ollama
How to use jsbeaudry/makandal-multiple-v2-GGUF with Ollama:
ollama run hf.co/jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use jsbeaudry/makandal-multiple-v2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jsbeaudry/makandal-multiple-v2-GGUF with Docker Model Runner:
docker model run hf.co/jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
- Lemonade
How to use jsbeaudry/makandal-multiple-v2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.makandal-multiple-v2-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use jsbeaudry/makandal-multiple-v2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jsbeaudry/makandal-multiple-v2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jsbeaudry/makandal-multiple-v2-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
makandal-multiple-v2 — GGUF
Try it in your browser: The Trio Serie Space runs Makandal with the other two models of the series: the Klara tab is a conversation with its eight tools, by voice or text, and another tab compares it with gemma-3-1b-it (the Space runs the full-precision model). More on the Makandal page of thetrio.space.
jsbeaudry/makandal-multiple-v2 for llama.cpp:
Gemma 3 1B taught to answer in Haitian Creole, French, Spanish and English the way a voice assistant
speaks, to hold a conversation, and to call tools (time, weather, encyclopedia, search, calculator, memory,
saved knowledge). It is the local brain of Klara.
| quant | size | speed on an M3 Pro (Metal) |
|---|---|---|
| Q8_0 | 1.07 GB | ~91 tokens/s |
| Q4_K_M | 0.81 GB | ~101 tokens/s |
Q8_0 was converted from the trained weights on the training machine; Q4_K_M was quantized from the 16-bit
weights. Both were checked in llama-server before publishing: the same tool calls on weather,
encyclopedia, percentage and small-talk requests, and the same kind of spoken answer. Use Q8_0 unless
size is binding: at 1B, 4-bit quantization costs facts first.
See the main card for the results and the known limits: tool calls are reliable on a first request and much less so later in a conversation, a gap a v2.1 is planned to close.
Use
--jinja is needed for tools: it runs the chat template stored in the file and turns the
<tool_call>{…}</tool_call> blocks the model writes into OpenAI tool_calls.
llama-server -m makandal-multiple-v2-Q8_0.gguf -c 8192 --jinja
curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role": "system", "content": "Ou rele Klara. Ou se yon asistan vwa ki pale kreyòl ayisyen.\n"},
{"role": "user", "content": "Ki tan l ap fè Okap jodi a?"}],
"tools": [{"type": "function", "function": {"name": "weather", "description": "The weather now in a place.",
"parameters": {"type": "object", "properties": {"place": {"type": "string"}}, "required": ["place"]}}}]}'
After a tool result, send the call back with a short line before it in the assistant message (for
example "content": "Kite m gade sa."): the model was trained that way, and with an empty one it tends to
call another tool instead of answering.
- Downloads last month
- 213
4-bit
8-bit
Model tree for jsbeaudry/makandal-multiple-v2-GGUF
Base model
google/gemma-3-1b-pt