Instructions to use ZEROLABS1/zero-loop with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ZEROLABS1/zero-loop with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZEROLABS1/zero-loop:Q4_K_M # Run inference directly in the terminal: llama cli -hf ZEROLABS1/zero-loop:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZEROLABS1/zero-loop:Q4_K_M # Run inference directly in the terminal: llama cli -hf ZEROLABS1/zero-loop:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ZEROLABS1/zero-loop:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ZEROLABS1/zero-loop:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ZEROLABS1/zero-loop:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ZEROLABS1/zero-loop:Q4_K_M
Use Docker
docker model run hf.co/ZEROLABS1/zero-loop:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ZEROLABS1/zero-loop with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZEROLABS1/zero-loop" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZEROLABS1/zero-loop", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ZEROLABS1/zero-loop:Q4_K_M
- Ollama
How to use ZEROLABS1/zero-loop with Ollama:
ollama run hf.co/ZEROLABS1/zero-loop:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use ZEROLABS1/zero-loop with Docker Model Runner:
docker model run hf.co/ZEROLABS1/zero-loop:Q4_K_M
- Lemonade
How to use ZEROLABS1/zero-loop with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ZEROLABS1/zero-loop:Q4_K_M
Run and chat with the model
lemonade run user.zero-loop-Q4_K_M
List all available models
lemonade list
- Atomic Chat

ZERO LOOP
ZERO LOOP is a tiny foundation language model from Zero Labs, trained from scratch. At 19.9M parameters it is built to run fully on-device: it thinks before it answers, can call a calculator and Wikipedia search, and is small enough to fine-tune on a CPU
Highlights
- Trained from scratch. Random initialisation, its own 16,384-token vocabulary, no pretrained weights reused.
- Think, then answer. Replies open with a short
<think>block, then the answer. - Tool-aware. Trained to emit JSON tool calls for exact arithmetic and Wikipedia lookup, so a 20M model does not have to memorise facts or do maths in its head.
- Tiny. The Q4_K_M file is 14 MB; Q8_0 is 22 MB. It runs anywhere llama.cpp runs: phones, single-board computers, laptops, browsers.
- Made to be fine-tuned. Clean base weights, chat weights, tokenizer and GGUF files, all in one repo.
What is in this repo
| Path | Contents |
|---|---|
model.safetensors, config.json, tokenizer files |
ZERO LOOP chat model (post-trained), Hugging Face format, fp32 |
base/ |
ZERO LOOP base model (pretrained only, no chat post-training), for continued pretraining and your own fine-tuning |
zero-loop-*.gguf |
GGUF files for llama.cpp, LM Studio, PocketPal, Ollama-style runtimes (table below) |
tools.py, agent.py |
Reference tool runner: executes the calculator and Wikipedia search for the model |
| GGUF file | Size | Use |
|---|---|---|
zero-loop-F16.gguf |
40 MB | full precision, reference |
zero-loop-Q8_0.gguf |
22 MB | recommended: best quality |
zero-loop-Q6_K.gguf |
17 MB | near-lossless |
zero-loop-Q5_K_M.gguf |
15 MB | very high quality |
zero-loop-Q4_K_M.gguf |
14 MB | small and fast |
zero-loop-Q3_K_M.gguf |
12 MB | smallest, lower quality |
Model details
| Developer | Zero Labs |
| Parameters | 19.9M |
| Architecture | Decoder-only transformer,Architecture in Transformers (grouped-query attention, RoPE, RMSNorm, SwiGLU, QK-norm, tied embeddings) |
| Layers / hidden size / FFN size | 16 / 256 / 1024 |
| Attention heads (query / key-value) | 4 / 2, head dimension 64 |
| Vocabulary | 16,384 tokens, byte-level BPE, trimmed from the Qwen2/3 vocabulary to the tokens that matter for this model |
| Context | pretrained on 1,024-token sequences, post-trained on up to 1,536 tokens (config allows 4,096; quality beyond the trained length is not guaranteed) |
| Pretraining data | about 0.6B tokens of educational web text and synthetic textbooks |
| Language | English |
| License | Apache-2.0 |
Quickstart
llama.cpp (recent build, uses the chat template stored in the GGUF):
llama-cli -m zero-loop-Q8_0.gguf -cnv --jinja --temp 0.4 --top-p 0.9 --repeat-penalty 1.1
PocketPal / LM Studio / other apps: load a .gguf, turn Reasoning model on, keep Add Generation Prompt on, leave the system prompt empty (a default is built in), set the stop word to <|end|>, BOS/EOS off. Suggested sampling: temperature 0.3 to 0.6, top_p 0.9, repeat penalty 1.1.
Transformers:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
repo = "ZEROLABS1/zero-loop"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.float32).eval()
msgs = [{"role": "user", "content": "Explain what a prime number is."}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
ids = tok(text, return_tensors="pt", add_special_tokens=False)
out = model.generate(**ids, max_new_tokens=300, do_sample=True, temperature=0.4, top_p=0.9,
repetition_penalty=1.1, eos_token_id=tok.convert_tokens_to_ids("<|end|>"))
print(tok.decode(out[0, ids["input_ids"].shape[1]:], skip_special_tokens=True))
Prompt format
<|system|>You are Zero Loop ...<|end|><|user|>Hello<|end|><|assistant|><think>
short reasoning
</think>
Answer<|end|>
The chat template opens the <think> block for the model. When the model needs a tool it ends its turn with a call, and the application returns the result in a <|tool|> turn:
<tool_call>{"name": "calculator", "arguments": {"expression": "2*400"}}</tool_call><|end|>
<|tool|>800<|end|><|assistant|><think> ... </think> 2 x 400 = 800.<|end|>
Available tools: search (argument query, Wikipedia) and calculator (argument expression). Chat apps only display the call; they do not run it. To execute tools, use python agent.py <model.gguf> (needs tools.py and llama-cpp-python) or implement the same loop in your app: parse the call, run it, append a tool message, generate again.
Intended uses
- Offline, private assistants on phones, wearables, kiosks and embedded devices.
- Voice and text front-ends that route intents or call tools on constrained hardware.
- A cheap first-pass model in larger systems (triage, short replies, formatting).
- A base for domain assistants: FAQ and support bots, command and form understanding, game NPCs, education tools.
- Research and teaching on small language models: a complete, reproducible-scale reference you can fully fine-tune on one GPU.
Fine-tuning
ZERO LOOP is small enough for full fine-tuning (no LoRA needed) on a single consumer GPU Start from base/ for new domains or languages, or from the chat weights to adapt behaviour.
from datasets import Dataset
from transformers import (AutoTokenizer, AutoModelForCausalLM, TrainingArguments,
Trainer, DataCollatorForLanguageModeling)
repo = "ZEROLABS1/zero-loop"
tok = AutoTokenizer.from_pretrained(repo, subfolder="base")
model = AutoModelForCausalLM.from_pretrained(repo, subfolder="base")
# write your data in the prompt format shown above
texts = ["<|user|>Your question<|end|><|assistant|><think>\nshort reasoning\n</think>\nYour answer<|end|>"]
ds = Dataset.from_dict({"text": texts}).map(
lambda b: tok(b["text"], truncation=True, max_length=1024, add_special_tokens=False),
batched=True, remove_columns=["text"])
args = TrainingArguments(output_dir="out", per_device_train_batch_size=16, learning_rate=2e-4,
lr_scheduler_type="cosine", warmup_ratio=0.05, num_train_epochs=3,
weight_decay=0.1, fp16=True, logging_steps=20, save_strategy="no", report_to=[])
Trainer(model=model, args=args, train_dataset=ds,
data_collator=DataCollatorForLanguageModeling(tok, mlm=False)).train()
model.save_pretrained("my-zero-loop"); tok.save_pretrained("my-zero-loop")
Tips: for SFT use a peak learning rate around 1e-4 to 3e-4 and compute the loss on assistant tokens only; for continued pretraining on raw domain text around 3e-4 to 5e-4 with cosine decay worked well for this model. Mix in some general data to limit forgetting. Keep the prompt format and the <think> structure if you want the same behaviour.
To make your own GGUF, use llama.cpp convert_hf_to_gguf.py. The tokenizer uses the Qwen2 pre-tokenizer regex, so if your converter does not recognise the tokenizer hash, set the pre-tokenizer type to qwen2; the GGUF files here were produced that way and their tokenization was checked against the Transformers tokenizer.
Training
Tokenizer. A byte-level BPE vocabulary to the 16,384 tokens most used in the training text, plus chat and tool markers.
Pretraining (from random initialisation). Next-token prediction on educational web text and synthetic textbook-style text (FineWeb-Edu and Cosmopedia from the SmolLM corpus), in several phases with cosine learning-rate schedules. Part of pretraining also used logit distillation from Qwen/Qwen3-1.7B-Base (soft targets restricted to this model's vocabulary). A final continued-pretraining phase mixes in conversational, reasoning and tool-call text so the base already knows the format. The checkpoint was only kept if held-out validation loss did not get worse.
Post-training (SFT). Supervised fine-tuning on assistant tokens only, on conversations written by Qwen3-family teacher models (Qwen3-1.7B and, where available, Qwen3-4B) and filtered:
- prompts from SmolTalk and UltraChat for conversation, including multi-turn follow-ups;
- GSM8K word problems and TriviaQA / Natural Questions / HotpotQA questions, kept only when the final answer matches the dataset's gold answer;
- tool calls whose results come from the real Wikipedia API or from exact calculator code, plus rows where the model says it is not sure when the search result does not contain the answer;
- chat answers self-graded by the teacher, low scores dropped;
- hand-written identity data, and typo / sloppy-phrasing augmentation so short or misspelled messages still work.
The best checkpoint on a held-out validation set was kept, and it was published only if it beat the previous release on that set.
Evaluation
| ZERO LOOP | ZERO LOOP Q4_K_M | SmolLM2-135M-Instruct | SmolLM-135M-Instruct | |
|---|---|---|---|---|
| Size And Speed | ||||
| Parameters | 19.9M | 19.9M | 134.5M | 134.5M |
| Download size | 40 MB | 14 MB | 269 MB | 269 MB |
| CPU speed (tokens/s, 4 threads) | 40 | 464 | 18 | 18 |
| Tool Use | ||||
| ZL-Calc: calculator tool | 0 | 0 | 6 | 0 |
| ZL-Calc-Noisy: typos, lower-case, × | 0 | 0 | 0 | 0 |
| ZL-Word: GSM8K problems + calculator | 0 | 3 | 1 | 0 |
| ZL-Search: TriviaQA + Wikipedia tool | 3 | 5 | 16 | 20 |
| Valid tool call on the first try | 0 | 0 | 16 | 4 |
| Assistant Behaviour | ||||
| ZL-Identity: name and creator | 92 | 85 | 46 | 31 |
ZERO LOOP is ahead of every 134M-class baseline on 1 of 6 benchmarks and tied on 1.
Same machine, greedy decoding and the same tool protocol for every model. The baselines were not trained on the protocol, so they get it in the system prompt plus two worked examples; ZERO LOOP uses its built-in prompt. Identity: baselines are told their name and creator in the system prompt. ZL-Calc / ZL-Calc-Noisy: 100 arithmetic questions with unseen numbers (the noisy set is lower-cased, with a typo and a multiplication sign). ZL-Word: 100 GSM8K test word problems, the model may call the calculator. ZL-Search: 100 TriviaQA validation questions, live Wikipedia search, correct if the final answer contains a gold alias. Valid tool call: the first reply to an arithmetic question is a parseable call. These are our own small tests of the skills ZERO LOOP is built for (100 questions each: differences under ~10 points are noise). They say nothing about general ability: on MMLU, HellaSwag or ARC the larger models trained on far more data score higher. The Q4_K_M column runs the shipped GGUF in llama.cpp; its speed comes from llama.cpp, the other columns from Transformers fp32, so speed is shown for it but not compared.
Held-out web-text perplexity of the base model: 25.0 (cross-entropy 3.217 nats per token on this model's own vocabulary, so not comparable across tokenizers). We recommend evaluating on your own task before relying on it.
Limitations and responsible use
- It is a 20M-parameter model. Facts it states from memory are often wrong; use the search tool or your own retrieval for anything factual.
- Arithmetic without the calculator tool is unreliable.
- Tool calls can be malformed or unnecessary. Validate every call before running it and never give tools more access than needed.
- The
<think>text is a short scratchpad, not a faithful explanation of the model's reasoning. - English only, short context, limited long-form coherence, and weak on code.
- Web data and teacher outputs carry biases and errors. There is no dedicated safety tuning beyond data filtering, so add your own filters for anything user-facing.
- Not for medical, legal, financial or other safety-critical decisions.
License and attribution
Weights and code in this repository are released under Apache-2.0. Training used public datasets (FineWeb-Edu, Cosmopedia, SmolTalk, UltraChat, GSM8K, TriviaQA, Natural Questions, HotpotQA, Wikipedia) that carry their own licenses, teacher outputs from Qwen3 models (Apache-2.0), and a vocabulary derived from Qwen2/3. GGUF conversion and quantisation use llama.cpp.
Citation
@misc{zeroloop2026,
title = {ZERO LOOP: a tiny foundation language model trained from scratch},
author = {Zero Labs},
year = {2026},
howpublished = {Hugging Face},
url = {https://huggingface.co/ZEROLABS1/zero-loop}
}
- Downloads last month
- -
