Europa B1

Kozu AI Research · 9.41B · Apache 2.0
BF16 weights Reasoning + instruction Report EB1-2026.07
More capability. Less deliberation.

Release brief

01

Europa B1 is a 9B-class reasoning and instruction system. In like-for-like full-budget evaluation, it improves six of seven reported capability rows and reduces thinking tokens on every task where token counts were recorded.

Europa B1 reaches 0.853 GSM8K flexible and 0.563 MMLU-Pro in the recorded run. Mean thinking tokens fall by 5.8× on GSM8K and 1.8× on MMLU-Pro.

5.8×fewer GSM8K think tokens
+0.329MMLU-Pro absolute delta
0.455held-out reasoning accuracy · n=200

Evaluation

02

The table reports matched comparisons under the same open harness, generation budget, sampling settings, and seed. Values are run-local and should not be compared with vendor-published numbers produced by other evaluation stacks.

Full-budget benchmarkReferenceEuropa B1DeltaThink tokens · ref → B1
GSM8K · flexible0.8000.853+0.0531,444 → 250
MMLU-Pro0.2340.563+0.3291,411 → 793
IFEval · instruction loose0.4330.454+0.0214,874 → 4,442
IFEval · instruction strict0.4330.445+0.012—
IFEval · prompt loose0.2530.280+0.027—
IFEval · prompt strict0.2530.273+0.020—
GSM8K · strict format0.6330.593−0.040—

lm-eval 0.4.12 · 32,768-token generation budget · temperature 1.0 · top_p 0.95 · presence penalty 1.5 · thinking enabled · GSM8K/IFEval n=150 · MMLU-Pro n=25 per subtask · seed 42.

  • The internal held-out math set is contamination-filtered with verified ground truth: accuracy moved from 0.200 to 0.455 at n=200. A separate 12-check instruction suite moved from 0.917 to 0.833; its small sample makes it directional, not authoritative.
  • IFEval scores the complete response, including the thinking block. Shorter reasoning can therefore improve constraint scores; the behavior affects both compared models.
  • The n=150 IFEval run is the primary instruction-following result and improves all four reported metrics. Lipogram-style constraints remain narrow and unreliable, succeeding roughly 40% of the time in repeated checks.
  • GSM8K strict rewards an unrequested #### N ending. The recorded decline is 0.040; request required formats explicitly.
  • Raw evaluation artifacts are retained under bench/ in this repository.

Data & development

03

kozu_reasoning_v1.1 is a roughly 10k-example blend of verified reasoning traces and human-authored instruction data. The release uses a three-layer quality process for the reasoning portion and a dedicated 800-example format-adherence slice.

Reasoning · ~5kKuiper-R1 expansions of answer-scrubbed human reference solutions from GSM8K and NuminaMath 1.5. Retained traces passed symbolic answer verification, derivation checks, deterministic defect scans, and an adversarial judge panel.
Instruction · ~5kHuman-authored pairs used from Databricks Dolly 15k and OpenAssistant OASST2.
Format · 800Deterministic rewraps of verified rows for explicitly requested formats including JSON, boxed answers, and named answer markers.
Source licenses and attribution
GSM8K: MIT. NuminaMath 1.5: Apache-2.0. Databricks Dolly 15k: CC-BY-SA-3.0. OpenAssistant OASST2: Apache-2.0. Reasoning traces were generated with Kuiper-R1 by Kozu AI.

Deployment

04

The repository includes merged BF16 weights, tokenizer assets, processor configuration, generation configuration, and the model chat template. Thinking is enabled by default and appears inside <think>…</think> before the final answer.

vllm serve Michael-Kozu/Europa-B1 --served-model-name europa-b1 \
  --max-model-len 8192 --gpu-memory-utilization 0.85 --trust-remote-code

Recommended sampling: temperature 0.6–1.0 and top_p 0.95. Request required output formats explicitly.

Limitations

05
  • This is a research preview, not a safety-certified or production-guaranteed system.
  • The strongest evidence covers English reasoning and instruction following. Performance outside those domains is not established here.
  • Letter-avoidance constraints are unreliable, and silent few-shot format imitation remains weaker than explicit format requests.
  • Long reasoning may still be incorrect. Verify outputs for consequential decisions and domain-specific use.
Kozu AI · Europa B1Apache-2.0 · Model release
Downloads last month
3
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Michael-Kozu/Europa-B1

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(527)
this model