makandal-multiple-v2 — GGUF

Try it in your browser: The Trio Serie Space runs Makandal with the other two models of the series: the Klara tab is a conversation with its eight tools, by voice or text, and another tab compares it with gemma-3-1b-it (the Space runs the full-precision model). More on the Makandal page of thetrio.space.

jsbeaudry/makandal-multiple-v2 for llama.cpp: Gemma 3 1B taught to answer in Haitian Creole, French, Spanish and English the way a voice assistant speaks, to hold a conversation, and to call tools (time, weather, encyclopedia, search, calculator, memory, saved knowledge). It is the local brain of Klara.

quant size speed on an M3 Pro (Metal)
Q8_0 1.07 GB ~91 tokens/s
Q4_K_M 0.81 GB ~101 tokens/s

Q8_0 was converted from the trained weights on the training machine; Q4_K_M was quantized from the 16-bit weights. Both were checked in llama-server before publishing: the same tool calls on weather, encyclopedia, percentage and small-talk requests, and the same kind of spoken answer. Use Q8_0 unless size is binding: at 1B, 4-bit quantization costs facts first.

See the main card for the results and the known limits: tool calls are reliable on a first request and much less so later in a conversation, a gap a v2.1 is planned to close.

Use

--jinja is needed for tools: it runs the chat template stored in the file and turns the <tool_call>{…}</tool_call> blocks the model writes into OpenAI tool_calls.

llama-server -m makandal-multiple-v2-Q8_0.gguf -c 8192 --jinja
curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "messages": [{"role": "system", "content": "Ou rele Klara. Ou se yon asistan vwa ki pale kreyòl ayisyen.\n"},
               {"role": "user", "content": "Ki tan l ap fè Okap jodi a?"}],
  "tools": [{"type": "function", "function": {"name": "weather", "description": "The weather now in a place.",
             "parameters": {"type": "object", "properties": {"place": {"type": "string"}}, "required": ["place"]}}}]}'

After a tool result, send the call back with a short line before it in the assistant message (for example "content": "Kite m gade sa."): the model was trained that way, and with an empty one it tends to call another tool instead of answering.

Downloads last month
213
GGUF
Model size
1.0B params
Architecture
gemma3
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jsbeaudry/makandal-multiple-v2-GGUF

Quantized
(2)
this model

Space using jsbeaudry/makandal-multiple-v2-GGUF 1