ZERO LOOP

ZERO LOOP

ZERO LOOP is a tiny foundation language model from Zero Labs, trained from scratch. At 19.9M parameters it is built to run fully on-device: it thinks before it answers, can call a calculator and Wikipedia search, and is small enough to fine-tune on a CPU

Highlights

  • Trained from scratch. Random initialisation, its own 16,384-token vocabulary, no pretrained weights reused.
  • Think, then answer. Replies open with a short <think> block, then the answer.
  • Tool-aware. Trained to emit JSON tool calls for exact arithmetic and Wikipedia lookup, so a 20M model does not have to memorise facts or do maths in its head.
  • Tiny. The Q4_K_M file is 14 MB; Q8_0 is 22 MB. It runs anywhere llama.cpp runs: phones, single-board computers, laptops, browsers.
  • Made to be fine-tuned. Clean base weights, chat weights, tokenizer and GGUF files, all in one repo.

What is in this repo

Path Contents
model.safetensors, config.json, tokenizer files ZERO LOOP chat model (post-trained), Hugging Face format, fp32
base/ ZERO LOOP base model (pretrained only, no chat post-training), for continued pretraining and your own fine-tuning
zero-loop-*.gguf GGUF files for llama.cpp, LM Studio, PocketPal, Ollama-style runtimes (table below)
tools.py, agent.py Reference tool runner: executes the calculator and Wikipedia search for the model
GGUF file Size Use
zero-loop-F16.gguf 40 MB full precision, reference
zero-loop-Q8_0.gguf 22 MB recommended: best quality
zero-loop-Q6_K.gguf 17 MB near-lossless
zero-loop-Q5_K_M.gguf 15 MB very high quality
zero-loop-Q4_K_M.gguf 14 MB small and fast
zero-loop-Q3_K_M.gguf 12 MB smallest, lower quality

Model details

Developer Zero Labs
Parameters 19.9M
Architecture Decoder-only transformer,Architecture in Transformers (grouped-query attention, RoPE, RMSNorm, SwiGLU, QK-norm, tied embeddings)
Layers / hidden size / FFN size 16 / 256 / 1024
Attention heads (query / key-value) 4 / 2, head dimension 64
Vocabulary 16,384 tokens, byte-level BPE, trimmed from the Qwen2/3 vocabulary to the tokens that matter for this model
Context pretrained on 1,024-token sequences, post-trained on up to 1,536 tokens (config allows 4,096; quality beyond the trained length is not guaranteed)
Pretraining data about 0.6B tokens of educational web text and synthetic textbooks
Language English
License Apache-2.0

Quickstart

llama.cpp (recent build, uses the chat template stored in the GGUF):

llama-cli -m zero-loop-Q8_0.gguf -cnv --jinja --temp 0.4 --top-p 0.9 --repeat-penalty 1.1

PocketPal / LM Studio / other apps: load a .gguf, turn Reasoning model on, keep Add Generation Prompt on, leave the system prompt empty (a default is built in), set the stop word to <|end|>, BOS/EOS off. Suggested sampling: temperature 0.3 to 0.6, top_p 0.9, repeat penalty 1.1.

Transformers:

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

repo = "ZEROLABS1/zero-loop"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.float32).eval()

msgs = [{"role": "user", "content": "Explain what a prime number is."}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
ids = tok(text, return_tensors="pt", add_special_tokens=False)
out = model.generate(**ids, max_new_tokens=300, do_sample=True, temperature=0.4, top_p=0.9,
                     repetition_penalty=1.1, eos_token_id=tok.convert_tokens_to_ids("<|end|>"))
print(tok.decode(out[0, ids["input_ids"].shape[1]:], skip_special_tokens=True))

Prompt format

<|system|>You are Zero Loop ...<|end|><|user|>Hello<|end|><|assistant|><think>
short reasoning
</think>
Answer<|end|>

The chat template opens the <think> block for the model. When the model needs a tool it ends its turn with a call, and the application returns the result in a <|tool|> turn:

<tool_call>{"name": "calculator", "arguments": {"expression": "2*400"}}</tool_call><|end|>
<|tool|>800<|end|><|assistant|><think> ... </think> 2 x 400 = 800.<|end|>

Available tools: search (argument query, Wikipedia) and calculator (argument expression). Chat apps only display the call; they do not run it. To execute tools, use python agent.py <model.gguf> (needs tools.py and llama-cpp-python) or implement the same loop in your app: parse the call, run it, append a tool message, generate again.

Intended uses

  • Offline, private assistants on phones, wearables, kiosks and embedded devices.
  • Voice and text front-ends that route intents or call tools on constrained hardware.
  • A cheap first-pass model in larger systems (triage, short replies, formatting).
  • A base for domain assistants: FAQ and support bots, command and form understanding, game NPCs, education tools.
  • Research and teaching on small language models: a complete, reproducible-scale reference you can fully fine-tune on one GPU.

Fine-tuning

ZERO LOOP is small enough for full fine-tuning (no LoRA needed) on a single consumer GPU Start from base/ for new domains or languages, or from the chat weights to adapt behaviour.

from datasets import Dataset
from transformers import (AutoTokenizer, AutoModelForCausalLM, TrainingArguments,
                          Trainer, DataCollatorForLanguageModeling)

repo = "ZEROLABS1/zero-loop"
tok = AutoTokenizer.from_pretrained(repo, subfolder="base")
model = AutoModelForCausalLM.from_pretrained(repo, subfolder="base")

# write your data in the prompt format shown above
texts = ["<|user|>Your question<|end|><|assistant|><think>\nshort reasoning\n</think>\nYour answer<|end|>"]
ds = Dataset.from_dict({"text": texts}).map(
    lambda b: tok(b["text"], truncation=True, max_length=1024, add_special_tokens=False),
    batched=True, remove_columns=["text"])

args = TrainingArguments(output_dir="out", per_device_train_batch_size=16, learning_rate=2e-4,
                         lr_scheduler_type="cosine", warmup_ratio=0.05, num_train_epochs=3,
                         weight_decay=0.1, fp16=True, logging_steps=20, save_strategy="no", report_to=[])
Trainer(model=model, args=args, train_dataset=ds,
        data_collator=DataCollatorForLanguageModeling(tok, mlm=False)).train()
model.save_pretrained("my-zero-loop"); tok.save_pretrained("my-zero-loop")

Tips: for SFT use a peak learning rate around 1e-4 to 3e-4 and compute the loss on assistant tokens only; for continued pretraining on raw domain text around 3e-4 to 5e-4 with cosine decay worked well for this model. Mix in some general data to limit forgetting. Keep the prompt format and the <think> structure if you want the same behaviour.

To make your own GGUF, use llama.cpp convert_hf_to_gguf.py. The tokenizer uses the Qwen2 pre-tokenizer regex, so if your converter does not recognise the tokenizer hash, set the pre-tokenizer type to qwen2; the GGUF files here were produced that way and their tokenization was checked against the Transformers tokenizer.

Training

Tokenizer. A byte-level BPE vocabulary to the 16,384 tokens most used in the training text, plus chat and tool markers.

Pretraining (from random initialisation). Next-token prediction on educational web text and synthetic textbook-style text (FineWeb-Edu and Cosmopedia from the SmolLM corpus), in several phases with cosine learning-rate schedules. Part of pretraining also used logit distillation from Qwen/Qwen3-1.7B-Base (soft targets restricted to this model's vocabulary). A final continued-pretraining phase mixes in conversational, reasoning and tool-call text so the base already knows the format. The checkpoint was only kept if held-out validation loss did not get worse.

Post-training (SFT). Supervised fine-tuning on assistant tokens only, on conversations written by Qwen3-family teacher models (Qwen3-1.7B and, where available, Qwen3-4B) and filtered:

  • prompts from SmolTalk and UltraChat for conversation, including multi-turn follow-ups;
  • GSM8K word problems and TriviaQA / Natural Questions / HotpotQA questions, kept only when the final answer matches the dataset's gold answer;
  • tool calls whose results come from the real Wikipedia API or from exact calculator code, plus rows where the model says it is not sure when the search result does not contain the answer;
  • chat answers self-graded by the teacher, low scores dropped;
  • hand-written identity data, and typo / sloppy-phrasing augmentation so short or misspelled messages still work.

The best checkpoint on a held-out validation set was kept, and it was published only if it beat the previous release on that set.

Evaluation

ZERO LOOP ZERO LOOP Q4_K_M SmolLM2-135M-Instruct SmolLM-135M-Instruct
Size And Speed
Parameters 19.9M 19.9M 134.5M 134.5M
Download size 40 MB 14 MB 269 MB 269 MB
CPU speed (tokens/s, 4 threads) 40 464 18 18
Tool Use
ZL-Calc: calculator tool 0 0 6 0
ZL-Calc-Noisy: typos, lower-case, × 0 0 0 0
ZL-Word: GSM8K problems + calculator 0 3 1 0
ZL-Search: TriviaQA + Wikipedia tool 3 5 16 20
Valid tool call on the first try 0 0 16 4
Assistant Behaviour
ZL-Identity: name and creator 92 85 46 31

ZERO LOOP benchmarks

ZERO LOOP is ahead of every 134M-class baseline on 1 of 6 benchmarks and tied on 1.

Same machine, greedy decoding and the same tool protocol for every model. The baselines were not trained on the protocol, so they get it in the system prompt plus two worked examples; ZERO LOOP uses its built-in prompt. Identity: baselines are told their name and creator in the system prompt. ZL-Calc / ZL-Calc-Noisy: 100 arithmetic questions with unseen numbers (the noisy set is lower-cased, with a typo and a multiplication sign). ZL-Word: 100 GSM8K test word problems, the model may call the calculator. ZL-Search: 100 TriviaQA validation questions, live Wikipedia search, correct if the final answer contains a gold alias. Valid tool call: the first reply to an arithmetic question is a parseable call. These are our own small tests of the skills ZERO LOOP is built for (100 questions each: differences under ~10 points are noise). They say nothing about general ability: on MMLU, HellaSwag or ARC the larger models trained on far more data score higher. The Q4_K_M column runs the shipped GGUF in llama.cpp; its speed comes from llama.cpp, the other columns from Transformers fp32, so speed is shown for it but not compared.

Held-out web-text perplexity of the base model: 25.0 (cross-entropy 3.217 nats per token on this model's own vocabulary, so not comparable across tokenizers). We recommend evaluating on your own task before relying on it.

Limitations and responsible use

  • It is a 20M-parameter model. Facts it states from memory are often wrong; use the search tool or your own retrieval for anything factual.
  • Arithmetic without the calculator tool is unreliable.
  • Tool calls can be malformed or unnecessary. Validate every call before running it and never give tools more access than needed.
  • The <think> text is a short scratchpad, not a faithful explanation of the model's reasoning.
  • English only, short context, limited long-form coherence, and weak on code.
  • Web data and teacher outputs carry biases and errors. There is no dedicated safety tuning beyond data filtering, so add your own filters for anything user-facing.
  • Not for medical, legal, financial or other safety-critical decisions.

License and attribution

Weights and code in this repository are released under Apache-2.0. Training used public datasets (FineWeb-Edu, Cosmopedia, SmolTalk, UltraChat, GSM8K, TriviaQA, Natural Questions, HotpotQA, Wikipedia) that carry their own licenses, teacher outputs from Qwen3 models (Apache-2.0), and a vocabulary derived from Qwen2/3. GGUF conversion and quantisation use llama.cpp.

Citation

@misc{zeroloop2026,
  title        = {ZERO LOOP: a tiny foundation language model trained from scratch},
  author       = {Zero Labs},
  year         = {2026},
  howpublished = {Hugging Face},
  url          = {https://huggingface.co/ZEROLABS1/zero-loop}
}
Downloads last month
-
Safetensors
Model size
19.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support