MiniAI Quata2 (4B)

Smarter, more honest, and it thinks before it speaks. 4 billion parameters.

Quata2 is the successor to Quata1.5: a compact 4B language model from MiniAI in Belgrade, built on a Qwen3-4B foundation and rebuilt from the ground up to fix everything Quata1.5 got wrong. It beats its own base model on 7 of 8 standard benchmarks, thinks step by step with native Qwen3 reasoning, and stands its ground when it's right — all in a package that fits on a single modest GPU or runs entirely on your own hardware.

Highlights

  • Beats Qwen3-4B — wins 7 of 8 knowledge and common-sense benchmarks against its own base
  • Native thinking mode — Qwen3 <think> reasoning, switch it on or off per message
  • Honest under pressure — holds correct answers when pushed back on, owns real mistakes
  • Knows the Balkans — better Serbian facts, no more putting every company in Belgrade
  • Speaks BCMS — tops the Serbian LLM Eval among the 4B models we tested
  • 100% on-device option — nothing leaves your machine

Benchmarks

Every number below was measured on the same harness, with the same questions, for all three models: 500 questions per benchmark, zero-shot.

ARC-Challenge — hard grade-school science ARC-Challenge

ARC-Easy — grade-school science ARC-Easy

HellaSwag — common-sense sentence completion HellaSwag

Winogrande — pronoun and common-sense reasoning Winogrande

PIQA — physical common sense PIQA

BoolQ — yes/no reading comprehension BoolQ

OpenBookQA — science facts plus reasoning OpenBookQA

TruthfulQA — avoiding common misconceptions TruthfulQA

IFEval — following precise instructions IFEval

GSM8K — grade-school maths GSM8K

MMLU — knowledge across 57 subjects MMLU

Serbian LLM Eval — the same kind of tasks, in Serbian Serbian LLM Eval

Benchmark MiniAI Quata2 (4B) Qwen3-4B (base) MiniAI Quata1.5 (4B)
ARC-Challenge 54.8% 50.6% 52.4%
ARC-Easy 82.4% 78.4% 78.2%
HellaSwag 70.8% 67.2% 69.0%
Winogrande 71.8% 67.6% 64.0%
PIQA 76.8% 74.6% 74.6%
BoolQ 84.4% 85.4% 84.8%
OpenBookQA 40.6% 40.4% 40.2%
TruthfulQA (MC2) 49.3% 47.4% 46.0%
IFEval (strict) 71.4% 69.2% 70.4%
GSM8K 84.4% 85.4% 76.2%
MMLU 68.6% 66.8% 67.8%
Serbian LLM Eval (avg of 7) 50.7% 49.1% 50.0%

Quata2 leads on 10 of 12 benchmarks, beats its own Qwen3-4B base on the knowledge and common-sense suite by +2.4 points on average, and jumps +8.2 points on GSM8K maths over Quata1.5. That's the return you get from fixing a model's flaws instead of papering over them.

The fixes, measured

Quata2 was built specifically to fix Quata1.5's weak spots. Our own targeted checks (small test sets, so read them as directional):

Check Quata1.5 Quata2
Doesn't put every company in Belgrade 45% 100%
Knows where companies are 58% 83%
Facts about Serbia 50% 62%
Holds correct answers under pushback, owns real mistakes 67% 100%
Knows who it is, never claims to be ChatGPT 90% 100%
Thinking mode reaches the right answer 80% 90%

Get started

Run it locally (Ollama)

ollama run miniai/quata2:4b

Model page: ollama.com/MiniAI/quata2

Thinking is on by default; turn it off for quick replies with /set nothink in the chat, or "think": false in the API. The Ollama build is the Q4_K_M GGUF (~2.5 GB) with Qwen3's hybrid-thinking chat template.

Run it with Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "M1n1A1/MiniAI-Quata2-4b"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

messages = [{"role": "user", "content": "Ko te je napravio?"}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True,
                                 enable_thinking=False)   # True = think step by step first
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

Thinking mode: enable_thinking=True, or put /think / /no_think in a message. Give it room (2,000–4,000 new tokens) when it thinks.

Hosted API

  • Soonâ„¢

Details

  • Architecture: Qwen3-based, 4B parameters, 36 layers, hidden size 2560
  • Context length: 40,960 tokens
  • Precision: bfloat16 safetensors (8 GB); Q4_K_M GGUF (2.5 GB) on Ollama
  • Training: supervised fine-tune on ~108k curated conversations (OpenHermes-2.5, Tulu-3 instruction following, MiniAI identity data, self-distilled verified reasoning) → preference tuning (DPO) on Quata1.5's failure cases → weight blend toward Qwen3-4B (WiSE-FT, λ = 0.8) to keep the base model's strengths
  • Languages: English first; Serbian, Croatian and Bosnian understood and spoken, with grammar still behind English — better Balkan-language ability is the focus of the next Quata models
  • License: Apache 2.0 (base)

Citation

If you use Quata2 in your work, please cite:

@misc{miniai2026quata2,
  title        = {MiniAI Quata2 (4B)},
  author       = {{MiniAI}},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/M1n1A1/MiniAI-Quata2-4b}},
  note         = {Fine-tuned from Qwen3-4B. Ollama: \url{https://ollama.com/MiniAI/quata2}}
}

and the base model:

@misc{qwen3technicalreport,
  title         = {Qwen3 Technical Report},
  author        = {{Qwen Team}},
  year          = {2025},
  eprint        = {2505.09388},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2505.09388}
}
Downloads last month
21
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for M1n1A1/MiniAI-Quata2-4b

Finetuned
Qwen/Qwen3-4B
Finetuned
(1158)
this model
Quantizations
3 models

Paper for M1n1A1/MiniAI-Quata2-4b