Main repo:

GGUF / llama.cpp

The tokenizer has no BPE merges and a pre-tokenizer unknown to the stock converter, so convert_hf_to_gguf.py fails on it. The make_gguf.py wrapper from the code repository patches this at conversion time without modifying llama.cpp (details in its header):

git clone https://github.com/ggml-org/llama.cpp
python make_gguf.py llama.cpp CalmaCatCoder-Next-mini calmacatcoder-f16.gguf f16     # or q8_0

Because <|im_end|> is plain text, llama.cpp does not stop on it by itself. Save the ChatML prompt to a file and use a reverse prompt:

llama-completion -m calmacatcoder-f16.gguf -no-cnv --no-escape -f prompt.txt -r "<|im_end|>" -n 300 --temp 0.7 --top-k 40

A warning special_eos_id is not in special_eog_ids is expected: the model has no EOS token. Prefer f16/bf16 or q8_0; this model is so small that aggressive quantization will likely hurt it.

Downloads last month
106
GGUF
Model size
96.6M params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ViorikaAI-org/CalmaCatCoder-next-mini-gguf

Quantized
(1)
this model

Collection including ViorikaAI-org/CalmaCatCoder-next-mini-gguf